Labs should agree to only[1] internally deploy models that, even with good elicitation[2], still fail to complete sophisticated sabotage tasks[3] without being caught by a monitor, where the monitor is allowed to use whatever interpretability techniques the lab uses in practice.
Seems pretty smart to me. I agree this is somewhat of a capability cap, and not a targeted on ban neuralese or training on CoT.
Like, models can already do quite a lot of unverbalised, reasoning and you’d expect scaling + trying to get the models to use fewer tokens, to keep making them better and better at this, even if people are very careful not to put (alignment/monitor-related) pressure on CoT. Even if we assume models will not get better at CoT control, just because the amount of compute per token goes up. This proposal would put a cap on this process, which is like the big way to scale capabilities, or at least require that interp/monitoring keep up.
Neuralese / train-on-cot /cot non-monitorability are general/vague enough that it seems to me that we will not be able to come up with binding rules that can be enforced consistently, and that solve the problem. Seems what we want are rules that are easy to instrumentalize, which differentially caps the type of scaling that routes through mechanisms that make models harder to monitor. Which this proposal does.
Hmm, when I think of neuralese I think of two things.
Scaling / pressure on CoT / RL-drift making the verbalized CoT of a transformer-like model uninterpretable to us
Models created with architectures that have unbounded serial depth over time, trained to leverage test-time compute.
And right now I feel like (1) is a problem we’re definitely gonna have to deal with, but where the difficulty/severity is a bit unclear. And (2) is less likely to be a problem (25%?), but could rapidly become a much more dangerous problem. However, its a problem that admits a precise mathematical definition, so writing a prescription that bans these unfriendly architectures is feasible, and would block a possible catastrophe.
Seems pretty smart to me. I agree this is somewhat of a capability cap, and not a targeted on ban neuralese or training on CoT.
Like, models can already do quite a lot of unverbalised, reasoning and you’d expect scaling + trying to get the models to use fewer tokens, to keep making them better and better at this, even if people are very careful not to put (alignment/monitor-related) pressure on CoT. Even if we assume models will not get better at CoT control, just because the amount of compute per token goes up. This proposal would put a cap on this process, which is like the big way to scale capabilities, or at least require that interp/monitoring keep up.
Neuralese / train-on-cot /cot non-monitorability are general/vague enough that it seems to me that we will not be able to come up with binding rules that can be enforced consistently, and that solve the problem. Seems what we want are rules that are easy to instrumentalize, which differentially caps the type of scaling that routes through mechanisms that make models harder to monitor. Which this proposal does.
Hmm, when I think of neuralese I think of two things.
Scaling / pressure on CoT / RL-drift making the verbalized CoT of a transformer-like model uninterpretable to us
Models created with architectures that have unbounded serial depth over time, trained to leverage test-time compute.
And right now I feel like (1) is a problem we’re definitely gonna have to deal with, but where the difficulty/severity is a bit unclear. And (2) is less likely to be a problem (25%?), but could rapidly become a much more dangerous problem. However, its a problem that admits a precise mathematical definition, so writing a prescription that bans these unfriendly architectures is feasible, and would block a possible catastrophe.