Thanks for writing this up! I think this proposal sounds pretty reasonable and it avoids one of the issues with capability-based gates that I thought would make such agreements not possible (by saying you can just use whatever interp and monitoring to pass the gate and thus making it a dynamic threshold). I tentatively would be in favor of it over no monitorability commitment at all.
However, I think this proposal does not address what I see as the strongest arguments against an outcome-based commitment (“achieve sufficient monitorability”) and in favor of an architecture-based commitment (“do not train neuralese models”), at least in the current regime of competitive frontier AI development.
Argument 1: It is hard to win a fight when the model is trained
I think the proposal you laid out would result in labs having internal fights about the interpretation of evidence from the sabotage capability evals, putting a lot of pressure on their methodology (eg what counts as fair amount of elicitation, which as you point out is the trickiest question; or what the results in the eval setting mean for the actual internal deployment that can look very differently). Under an outcome-based commitment, this fight will happen repeatedly, after a lot of people put their blood, sweat and tears into an artifact that they also have multiple strong incentives to ship. Being tricky to adjudicate + strong incentives to subvert seem like a bad combination for an agreement governing high-stakes technology.
you should always be able to unblock scaling if you get blocked by this just by improving your interpretability and monitoring to make the monitor once again stronger than the policy
For example, here I expect deciding whether the improved interpretability and monitoring are sufficient will be a huge fight and the decision will always be strongly biased in favor of deploying.
Or if the reason that this broke was because you made the CoT less monitorable—e.g. by training against it or by doing some neuralese-y thing—you should always be able to unblock by rolling that back.
Re-doing a training run (to avoid training against CoT) and rolling back the architecture and re-doing a run are incredibly expensive asks! I’d expect in practice some half-way solutions to happen (like do a bit of fine-tuning to “fix” monitorability) which will run into the problem from the previous example.
Argument 2: Sabotage capability evals (and other monitoraiblity evals) only test for monitorability of a few specific behaviors that we already know to look for. Currently, CoT provides rich signal about model cognition that monitorability evals do not capture. Therefore, relying solely on monitorability evals for holistic judgment of monitorability (regardless of other information, like model architecture) will over-estimate safety. There is more to say here, but I will leave it at this for now.
There’s also the argument some make that neuralese may make worst kinds of misalignment (eg scheming) more likely, which your proposal doesn’t address.
———
I wonder if you see any blockers for getting the best of both worlds and basing a commitment in both monitorability and architecture? Adding the architecture agreement would: (a) make the monitorability commitment more effective by alleviating some major sources of pressure from arguments 1 and 2, (b) make scheming less likely, and (c) buy interpretability research more time.
I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabotage capability evals (eg can we notice misalignment after a red team inserts it via SDF, or <insert your other favorite idea for ‘alignment monitorability’ as opposed to ‘control monitorability’>) that you wouldn’t be able to use as a gate for a commitment due to it not being possible to do elicitation there and the methodology being sufficiently finicky / early days that it’s very non robust to adversarial interpretation.
(This further stacks with the architecture commitment)
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability. I think it’s harder to try to get something pulled before you actually have concrete empirical evidence that it’s bad.
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it’s unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it’s trained is much harder, as discussed).
Worse, it might be that even after the model is trained, the evidence will be weak, because all the neuralese architecture might do is change the slope of no-CoT reasoning capability increase over (calendar, model-generation) time, without an immediate change in the intercept. Then, you will need a few generations of models to notice the change in slope, at which point you might be too locked into the architecture for this battle to be winnable.
Here is a cartoonish plot of how things might go: (relevant)
With this, it looks to me like, unless we’re very confident that we’re going to be able to get compelling evidence before it’s too hard a battle to win, the best time to have the battle is now (as early as possible, before significant investment into this tech tree; also making use of a window of increased political will).
Thanks for writing this up! I think this proposal sounds pretty reasonable and it avoids one of the issues with capability-based gates that I thought would make such agreements not possible (by saying you can just use whatever interp and monitoring to pass the gate and thus making it a dynamic threshold). I tentatively would be in favor of it over no monitorability commitment at all.
However, I think this proposal does not address what I see as the strongest arguments against an outcome-based commitment (“achieve sufficient monitorability”) and in favor of an architecture-based commitment (“do not train neuralese models”), at least in the current regime of competitive frontier AI development.
Argument 1: It is hard to win a fight when the model is trained
I think the proposal you laid out would result in labs having internal fights about the interpretation of evidence from the sabotage capability evals, putting a lot of pressure on their methodology (eg what counts as fair amount of elicitation, which as you point out is the trickiest question; or what the results in the eval setting mean for the actual internal deployment that can look very differently). Under an outcome-based commitment, this fight will happen repeatedly, after a lot of people put their blood, sweat and tears into an artifact that they also have multiple strong incentives to ship. Being tricky to adjudicate + strong incentives to subvert seem like a bad combination for an agreement governing high-stakes technology.
For example, here I expect deciding whether the improved interpretability and monitoring are sufficient will be a huge fight and the decision will always be strongly biased in favor of deploying.
Re-doing a training run (to avoid training against CoT) and rolling back the architecture and re-doing a run are incredibly expensive asks! I’d expect in practice some half-way solutions to happen (like do a bit of fine-tuning to “fix” monitorability) which will run into the problem from the previous example.
Argument 2: Sabotage capability evals (and other monitoraiblity evals) only test for monitorability of a few specific behaviors that we already know to look for. Currently, CoT provides rich signal about model cognition that monitorability evals do not capture. Therefore, relying solely on monitorability evals for holistic judgment of monitorability (regardless of other information, like model architecture) will over-estimate safety. There is more to say here, but I will leave it at this for now.
There’s also the argument some make that neuralese may make worst kinds of misalignment (eg scheming) more likely, which your proposal doesn’t address.
———
I wonder if you see any blockers for getting the best of both worlds and basing a commitment in both monitorability and architecture? Adding the architecture agreement would: (a) make the monitorability commitment more effective by alleviating some major sources of pressure from arguments 1 and 2, (b) make scheming less likely, and (c) buy interpretability research more time.
I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabotage capability evals (eg can we notice misalignment after a red team inserts it via SDF, or <insert your other favorite idea for ‘alignment monitorability’ as opposed to ‘control monitorability’>) that you wouldn’t be able to use as a gate for a commitment due to it not being possible to do elicitation there and the methodology being sufficiently finicky / early days that it’s very non robust to adversarial interpretation.
(This further stacks with the architecture commitment)
Do you see any problems with these options?
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability. I think it’s harder to try to get something pulled before you actually have concrete empirical evidence that it’s bad.
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it’s unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it’s trained is much harder, as discussed).
Worse, it might be that even after the model is trained, the evidence will be weak, because all the neuralese architecture might do is change the slope of no-CoT reasoning capability increase over (calendar, model-generation) time, without an immediate change in the intercept. Then, you will need a few generations of models to notice the change in slope, at which point you might be too locked into the architecture for this battle to be winnable.
Here is a cartoonish plot of how things might go: (relevant)
With this, it looks to me like, unless we’re very confident that we’re going to be able to get compelling evidence before it’s too hard a battle to win, the best time to have the battle is now (as early as possible, before significant investment into this tech tree; also making use of a window of increased political will).