Yes, exactly. I write more here. Re. 2. another method would be to make the AI objectively bad at those domains via. gradient routing and ablation or data filtering. Or alternatively monitor and block with probes or equally smart classifiers (differently prompted versions of the same model).
Yes, exactly. I write more here. Re. 2. another method would be to make the AI objectively bad at those domains via. gradient routing and ablation or data filtering. Or alternatively monitor and block with probes or equally smart classifiers (differently prompted versions of the same model).