I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabotage capability evals (eg can we notice misalignment after a red team inserts it via SDF, or <insert your other favorite idea for ‘alignment monitorability’ as opposed to ‘control monitorability’>) that you wouldn’t be able to use as a gate for a commitment due to it not being possible to do elicitation there and the methodology being sufficiently finicky / early days that it’s very non robust to adversarial interpretation.
(This further stacks with the architecture commitment)
I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabotage capability evals (eg can we notice misalignment after a red team inserts it via SDF, or <insert your other favorite idea for ‘alignment monitorability’ as opposed to ‘control monitorability’>) that you wouldn’t be able to use as a gate for a commitment due to it not being possible to do elicitation there and the methodology being sufficiently finicky / early days that it’s very non robust to adversarial interpretation.
(This further stacks with the architecture commitment)
Do you see any problems with these options?