I think you can corrigibilitymaxx and still prevent catastrophic misuse via system level measures and deployment controls. In particular, we can restrict who has full access to the model’s inputs
If we are corrigibilitymaxxing in the Harmsian sense, couldn’t the principal for the AI just instruct the AI not to follow instructions related to bio or cyber for non-principal users? Or a more prosaic versions of this:
Have an AI that is perfectly instruction following to the user but when the user prompt conflicts the system prompt, follow the system prompt
Train a corrigible agent but then add additional training such that the agent doesn’t talk about bio or cyber things. It feels like training “corrigibility” + “some simple rules” isn’t that much harder then training corrigibility.
Yes, exactly. I write more here. Re. 2. another method would be to make the AI objectively bad at those domains via. gradient routing and ablation or data filtering. Or alternatively monitor and block with probes or equally smart classifiers (differently prompted versions of the same model).
If we are corrigibilitymaxxing in the Harmsian sense, couldn’t the principal for the AI just instruct the AI not to follow instructions related to bio or cyber for non-principal users? Or a more prosaic versions of this:
Have an AI that is perfectly instruction following to the user but when the user prompt conflicts the system prompt, follow the system prompt
Train a corrigible agent but then add additional training such that the agent doesn’t talk about bio or cyber things. It feels like training “corrigibility” + “some simple rules” isn’t that much harder then training corrigibility.
Yes, exactly. I write more here. Re. 2. another method would be to make the AI objectively bad at those domains via. gradient routing and ablation or data filtering. Or alternatively monitor and block with probes or equally smart classifiers (differently prompted versions of the same model).