I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabotage capability evals (eg can we notice misalignment after a red team inserts it via SDF, or <insert your other favorite idea for ‘alignment monitorability’ as opposed to ‘control monitorability’>) that you wouldn’t be able to use as a gate for a commitment due to it not being possible to do elicitation there and the methodology being sufficiently finicky / early days that it’s very non robust to adversarial interpretation.
(This further stacks with the architecture commitment)
Do you see any problems with these options?
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it’s unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it’s trained is much harder, as discussed).
Worse, it might be that even after the model is trained, the evidence will be weak, because all the neuralese architecture might do is change the slope of no-CoT reasoning capability increase over (calendar, model-generation) time, without an immediate change in the intercept. Then, you will need a few generations of models to notice the change in slope, at which point you might be too locked into the architecture for this battle to be winnable.
Here is a cartoonish plot of how things might go: (relevant)
With this, it looks to me like, unless we’re very confident that we’re going to be able to get compelling evidence before it’s too hard a battle to win, the best time to have the battle is now (as early as possible, before significant investment into this tech tree; also making use of a window of increased political will).