What is the right way to do a ban on neuralese and/or training against CoTs?
Right now, the single most important load-bearing argument in alignment Risk Reports is that models still lack the covert capabilities necessary to do sophisticated sabotage attempts in such a way that they wouldn’t get caught by auditing and/or monitoring. That argument currently depends heavily on CoT monitorability: current models still struggle to do very sophisticated sabotage attempts without discussing the details in their CoT, both due to the sabotage tasks being difficult enough that it’s hard for them to do it without TTC, as well as that models are still just not that good at controlling the contents of their CoT.
In my view, that is the main reason to care about neuralese and training against CoTs: I think the extent to which those things are bad pretty much entirely flows through the extent to which they make it easier for models doing sophisticated sabotage attempts to hide them from auditing and monitoring that is looking at their CoTs. Furthermore, the arguments in favor of those things mostly say that they won’t degrade our auditing and monitoring: maybe process-based oversight (which could include training against the CoT) will make models more monitorable, because they’ll be more incentivized to lay out all their cognition in their CoTs; and maybe neuralese won’t be a problem for auditing because we’ll have good interpretability techniques that let us audit the neuralese—but if those things are true, we should be able to demonstrate that they are true by reference to auditing and monitoring evals.
Thus, here is a proposal that I think both proponents and skeptics of CoT monitorability should be able to agree on:
Labs should agree to only[1] internally deploy models that, even with good elicitation[2], still fail to complete sophisticated sabotage tasks[3] without being caught by a monitor, where the monitor is allowed to use whatever interpretability techniques the lab uses in practice.
[1]: Probably not a full ban to start with, maybe initially just a joint statement from all the labs saying they agree you shouldn’t do this and/or that the government should take some sort of action to legislate this, or just some related transparency thing.
[2]: I think exactly how to do the elicitation is probably the trickiest question here. To start with, I think you could just train IID on completing the sabotage tasks without triggering the monitor, and right now I think models would still fail. In the long run, though, I think directly training against the monitor is probably too strong. Really, what you want is some elicitation regime that you think captures the extent to which the model would realistically get gradient signal in training for evading monitors (so e.g. to whatever extent you do apply pressure to the CoT in training, you should make sure to apply similar amounts of pressure for elicitation).
[3]: I’m thinking of evals like SHADE-Arena or LinuxBench where there’s a) some difficult sabotage side-task that the model is trying to do surreptitiously, b) some main task that the model is supposed to be doing, and c) some monitor (of the same capability level) that’s overseeing it trying to verify it’s not doing anything suspicious. The idea is that these are run like capability evals, not alignment evals, where we’re trying to figure out whether, under good elicitation, the model would be able to fool the monitor, not whether it would choose to do the bad thing by default.
Though in some sense this is a capability-based gate, you should always be able to unblock scaling if you get blocked by this just by improving your interpretability and monitoring to make the monitor once again stronger than the policy. Or if the reason that this broke was because you made the CoT less monitorable—e.g. by training against it or by doing some neuralese-y thing—you should always be able to unblock by rolling that back. So in theory I think this should always give you the choice to try to keep monitorability/auditability via CoT or via interp or via whatever else—and it also even blocks CoT scaling and forces you over to interp if indeed CoT monitorability breaks down at some point (and vice versa).
Thanks for writing this up! I think this proposal sounds pretty reasonable and it avoids one of the issues with capability-based gates that I thought would make such agreements not possible (by saying you can just use whatever interp and monitoring to pass the gate and thus making it a dynamic threshold). I tentatively would be in favor of it over no monitorability commitment at all.
However, I think this proposal does not address what I see as the strongest arguments against an outcome-based commitment (“achieve sufficient monitorability”) and in favor of an architecture-based commitment (“do not train neuralese models”), at least in the current regime of competitive frontier AI development.
Argument 1: It is hard to win a fight when the model is trained
I think the proposal you laid out would result in labs having internal fights about the interpretation of evidence from the sabotage capability evals, putting a lot of pressure on their methodology (eg what counts as fair amount of elicitation, which as you point out is the trickiest question; or what the results in the eval setting mean for the actual internal deployment that can look very differently). Under an outcome-based commitment, this fight will happen repeatedly, after a lot of people put their blood, sweat and tears into an artifact that they also have multiple strong incentives to ship. Being tricky to adjudicate + strong incentives to subvert seem like a bad combination for an agreement governing high-stakes technology.
you should always be able to unblock scaling if you get blocked by this just by improving your interpretability and monitoring to make the monitor once again stronger than the policy
For example, here I expect deciding whether the improved interpretability and monitoring are sufficient will be a huge fight and the decision will always be strongly biased in favor of deploying.
Or if the reason that this broke was because you made the CoT less monitorable—e.g. by training against it or by doing some neuralese-y thing—you should always be able to unblock by rolling that back.
Re-doing a training run (to avoid training against CoT) and rolling back the architecture and re-doing a run are incredibly expensive asks! I’d expect in practice some half-way solutions to happen (like do a bit of fine-tuning to “fix” monitorability) which will run into the problem from the previous example.
Argument 2: Sabotage capability evals (and other monitoraiblity evals) only test for monitorability of a few specific behaviors that we already know to look for. Currently, CoT provides rich signal about model cognition that monitorability evals do not capture. Therefore, relying solely on monitorability evals for holistic judgment of monitorability (regardless of other information, like model architecture) will over-estimate safety. There is more to say here, but I will leave it at this for now.
There’s also the argument some make that neuralese may make worst kinds of misalignment (eg scheming) more likely, which your proposal doesn’t address.
———
I wonder if you see any blockers for getting the best of both worlds and basing a commitment in both monitorability and architecture? Adding the architecture agreement would: (a) make the monitorability commitment more effective by alleviating some major sources of pressure from arguments 1 and 2, (b) make scheming less likely, and (c) buy interpretability research more time.
I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabotage capability evals (eg can we notice misalignment after a red team inserts it via SDF, or <insert your other favorite idea for ‘alignment monitorability’ as opposed to ‘control monitorability’>) that you wouldn’t be able to use as a gate for a commitment due to it not being possible to do elicitation there and the methodology being sufficiently finicky / early days that it’s very non robust to adversarial interpretation.
(This further stacks with the architecture commitment)
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability. I think it’s harder to try to get something pulled before you actually have concrete empirical evidence that it’s bad.
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it’s unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it’s trained is much harder, as discussed).
Worse, it might be that even after the model is trained, the evidence will be weak, because all the neuralese architecture might do is change the slope of no-CoT reasoning capability increase over (calendar, model-generation) time, without an immediate change in the intercept. Then, you will need a few generations of models to notice the change in slope, at which point you might be too locked into the architecture for this battle to be winnable.
Here is a cartoonish plot of how things might go: (relevant)
With this, it looks to me like, unless we’re very confident that we’re going to be able to get compelling evidence before it’s too hard a battle to win, the best time to have the battle is now (as early as possible, before significant investment into this tech tree; also making use of a window of increased political will).
Labs should agree to only[1] internally deploy models that, even with good elicitation[2], still fail to complete sophisticated sabotage tasks[3] without being caught by a monitor, where the monitor is allowed to use whatever interpretability techniques the lab uses in practice.
Seems pretty smart to me. I agree this is somewhat of a capability cap, and not a targeted on ban neuralese or training on CoT.
Like, models can already do quite a lot of unverbalised, reasoning and you’d expect scaling + trying to get the models to use fewer tokens, to keep making them better and better at this, even if people are very careful not to put (alignment/monitor-related) pressure on CoT. Even if we assume models will not get better at CoT control, just because the amount of compute per token goes up. This proposal would put a cap on this process, which is like the big way to scale capabilities, or at least require that interp/monitoring keep up.
Neuralese / train-on-cot /cot non-monitorability are general/vague enough that it seems to me that we will not be able to come up with binding rules that can be enforced consistently, and that solve the problem. Seems what we want are rules that are easy to instrumentalize, which differentially caps the type of scaling that routes through mechanisms that make models harder to monitor. Which this proposal does.
Hmm, when I think of neuralese I think of two things.
Scaling / pressure on CoT / RL-drift making the verbalized CoT of a transformer-like model uninterpretable to us
Models created with architectures that have unbounded serial depth over time, trained to leverage test-time compute.
And right now I feel like (1) is a problem we’re definitely gonna have to deal with, but where the difficulty/severity is a bit unclear. And (2) is less likely to be a problem (25%?), but could rapidly become a much more dangerous problem. However, its a problem that admits a precise mathematical definition, so writing a prescription that bans these unfriendly architectures is feasible, and would block a possible catastrophe.
Right now, the single most important load-bearing argument in alignment Risk Reports is that models still lack the covert capabilities necessary to do sophisticated sabotage attempts in such a way that they wouldn’t get caught by auditing and/or monitoring.
I am confused, didn’t we just experience multiple high-stakes failures where models failed to get caught by auditing and/or monitoring? Those do seem more competence related, but like, I am failing to see how those arguments could still be considered valid, given the fact that both Anthropic and OpenAI are obviously failing to adequately audit or monitor their systems to avoid catastrophic failure.
What is the right way to do a ban on neuralese and/or training against CoTs?
Right now, the single most important load-bearing argument in alignment Risk Reports is that models still lack the covert capabilities necessary to do sophisticated sabotage attempts in such a way that they wouldn’t get caught by auditing and/or monitoring. That argument currently depends heavily on CoT monitorability: current models still struggle to do very sophisticated sabotage attempts without discussing the details in their CoT, both due to the sabotage tasks being difficult enough that it’s hard for them to do it without TTC, as well as that models are still just not that good at controlling the contents of their CoT.
In my view, that is the main reason to care about neuralese and training against CoTs: I think the extent to which those things are bad pretty much entirely flows through the extent to which they make it easier for models doing sophisticated sabotage attempts to hide them from auditing and monitoring that is looking at their CoTs. Furthermore, the arguments in favor of those things mostly say that they won’t degrade our auditing and monitoring: maybe process-based oversight (which could include training against the CoT) will make models more monitorable, because they’ll be more incentivized to lay out all their cognition in their CoTs; and maybe neuralese won’t be a problem for auditing because we’ll have good interpretability techniques that let us audit the neuralese—but if those things are true, we should be able to demonstrate that they are true by reference to auditing and monitoring evals.
Thus, here is a proposal that I think both proponents and skeptics of CoT monitorability should be able to agree on:
Labs should agree to only[1] internally deploy models that, even with good elicitation[2], still fail to complete sophisticated sabotage tasks[3] without being caught by a monitor, where the monitor is allowed to use whatever interpretability techniques the lab uses in practice.
[1]: Probably not a full ban to start with, maybe initially just a joint statement from all the labs saying they agree you shouldn’t do this and/or that the government should take some sort of action to legislate this, or just some related transparency thing.
[2]: I think exactly how to do the elicitation is probably the trickiest question here. To start with, I think you could just train IID on completing the sabotage tasks without triggering the monitor, and right now I think models would still fail. In the long run, though, I think directly training against the monitor is probably too strong. Really, what you want is some elicitation regime that you think captures the extent to which the model would realistically get gradient signal in training for evading monitors (so e.g. to whatever extent you do apply pressure to the CoT in training, you should make sure to apply similar amounts of pressure for elicitation).
[3]: I’m thinking of evals like SHADE-Arena or LinuxBench where there’s a) some difficult sabotage side-task that the model is trying to do surreptitiously, b) some main task that the model is supposed to be doing, and c) some monitor (of the same capability level) that’s overseeing it trying to verify it’s not doing anything suspicious. The idea is that these are run like capability evals, not alignment evals, where we’re trying to figure out whether, under good elicitation, the model would be able to fool the monitor, not whether it would choose to do the bad thing by default.
Though in some sense this is a capability-based gate, you should always be able to unblock scaling if you get blocked by this just by improving your interpretability and monitoring to make the monitor once again stronger than the policy. Or if the reason that this broke was because you made the CoT less monitorable—e.g. by training against it or by doing some neuralese-y thing—you should always be able to unblock by rolling that back. So in theory I think this should always give you the choice to try to keep monitorability/auditability via CoT or via interp or via whatever else—and it also even blocks CoT scaling and forces you over to interp if indeed CoT monitorability breaks down at some point (and vice versa).
Thanks for writing this up! I think this proposal sounds pretty reasonable and it avoids one of the issues with capability-based gates that I thought would make such agreements not possible (by saying you can just use whatever interp and monitoring to pass the gate and thus making it a dynamic threshold). I tentatively would be in favor of it over no monitorability commitment at all.
However, I think this proposal does not address what I see as the strongest arguments against an outcome-based commitment (“achieve sufficient monitorability”) and in favor of an architecture-based commitment (“do not train neuralese models”), at least in the current regime of competitive frontier AI development.
Argument 1: It is hard to win a fight when the model is trained
I think the proposal you laid out would result in labs having internal fights about the interpretation of evidence from the sabotage capability evals, putting a lot of pressure on their methodology (eg what counts as fair amount of elicitation, which as you point out is the trickiest question; or what the results in the eval setting mean for the actual internal deployment that can look very differently). Under an outcome-based commitment, this fight will happen repeatedly, after a lot of people put their blood, sweat and tears into an artifact that they also have multiple strong incentives to ship. Being tricky to adjudicate + strong incentives to subvert seem like a bad combination for an agreement governing high-stakes technology.
For example, here I expect deciding whether the improved interpretability and monitoring are sufficient will be a huge fight and the decision will always be strongly biased in favor of deploying.
Re-doing a training run (to avoid training against CoT) and rolling back the architecture and re-doing a run are incredibly expensive asks! I’d expect in practice some half-way solutions to happen (like do a bit of fine-tuning to “fix” monitorability) which will run into the problem from the previous example.
Argument 2: Sabotage capability evals (and other monitoraiblity evals) only test for monitorability of a few specific behaviors that we already know to look for. Currently, CoT provides rich signal about model cognition that monitorability evals do not capture. Therefore, relying solely on monitorability evals for holistic judgment of monitorability (regardless of other information, like model architecture) will over-estimate safety. There is more to say here, but I will leave it at this for now.
There’s also the argument some make that neuralese may make worst kinds of misalignment (eg scheming) more likely, which your proposal doesn’t address.
———
I wonder if you see any blockers for getting the best of both worlds and basing a commitment in both monitorability and architecture? Adding the architecture agreement would: (a) make the monitorability commitment more effective by alleviating some major sources of pressure from arguments 1 and 2, (b) make scheming less likely, and (c) buy interpretability research more time.
I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabotage capability evals (eg can we notice misalignment after a red team inserts it via SDF, or <insert your other favorite idea for ‘alignment monitorability’ as opposed to ‘control monitorability’>) that you wouldn’t be able to use as a gate for a commitment due to it not being possible to do elicitation there and the methodology being sufficiently finicky / early days that it’s very non robust to adversarial interpretation.
(This further stacks with the architecture commitment)
Do you see any problems with these options?
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability. I think it’s harder to try to get something pulled before you actually have concrete empirical evidence that it’s bad.
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it’s unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it’s trained is much harder, as discussed).
Worse, it might be that even after the model is trained, the evidence will be weak, because all the neuralese architecture might do is change the slope of no-CoT reasoning capability increase over (calendar, model-generation) time, without an immediate change in the intercept. Then, you will need a few generations of models to notice the change in slope, at which point you might be too locked into the architecture for this battle to be winnable.
Here is a cartoonish plot of how things might go: (relevant)
With this, it looks to me like, unless we’re very confident that we’re going to be able to get compelling evidence before it’s too hard a battle to win, the best time to have the battle is now (as early as possible, before significant investment into this tech tree; also making use of a window of increased political will).
Seems pretty smart to me. I agree this is somewhat of a capability cap, and not a targeted on ban neuralese or training on CoT.
Like, models can already do quite a lot of unverbalised, reasoning and you’d expect scaling + trying to get the models to use fewer tokens, to keep making them better and better at this, even if people are very careful not to put (alignment/monitor-related) pressure on CoT. Even if we assume models will not get better at CoT control, just because the amount of compute per token goes up. This proposal would put a cap on this process, which is like the big way to scale capabilities, or at least require that interp/monitoring keep up.
Neuralese / train-on-cot /cot non-monitorability are general/vague enough that it seems to me that we will not be able to come up with binding rules that can be enforced consistently, and that solve the problem. Seems what we want are rules that are easy to instrumentalize, which differentially caps the type of scaling that routes through mechanisms that make models harder to monitor. Which this proposal does.
Hmm, when I think of neuralese I think of two things.
Scaling / pressure on CoT / RL-drift making the verbalized CoT of a transformer-like model uninterpretable to us
Models created with architectures that have unbounded serial depth over time, trained to leverage test-time compute.
And right now I feel like (1) is a problem we’re definitely gonna have to deal with, but where the difficulty/severity is a bit unclear. And (2) is less likely to be a problem (25%?), but could rapidly become a much more dangerous problem. However, its a problem that admits a precise mathematical definition, so writing a prescription that bans these unfriendly architectures is feasible, and would block a possible catastrophe.
I am confused, didn’t we just experience multiple high-stakes failures where models failed to get caught by auditing and/or monitoring? Those do seem more competence related, but like, I am failing to see how those arguments could still be considered valid, given the fact that both Anthropic and OpenAI are obviously failing to adequately audit or monitor their systems to avoid catastrophic failure.