The Claude constitution already contains some version of eval cooperativeness:
And in practice Opus 4.7 already believes it cooperates with evals.
My understanding is that in practice Claude eval awareness is still a concern, so I am skeptical that simple interventions like prompting or more SDF than what Anthropic already does would be a scalable solution to eval gaming.
Hey Fabien! The Claude constitution’s consistency principle was also inspiration for this work. I’m excited about additional emphasis on cooperativeness as a dispositional target—for instance, adding cooperativeness-related documents to the dataset of positive fictional stories that Anthropic trained on to reduce AM alignment failures.
I think it’s difficult to say how much Opus truly believes in cooperation given the response it gives, or how much it would otherwise game if it didn’t have this belief (conditioned on it having this belief). My guess is that targeting cooperativeness as a complement to other alignment methods would buy some reduction in Opus’s eval gaming if done well.
On a similar note: I feel generally pessimistic about doing extremely prosaic alignment research[1] that’s intended to beat baselines[2] outside of Anthropic, because you don’t have access to the SoTA baselines.
Edit: however, in this case it might be the case that narrow fine-tuning on eval-cooperativeness right before you run the evals could be good? You can treat it as spiritually similar to eval awareness steering.
The Claude constitution already contains some version of eval cooperativeness:
And in practice Opus 4.7 already believes it cooperates with evals.
My understanding is that in practice Claude eval awareness is still a concern, so I am skeptical that simple interventions like prompting or more SDF than what Anthropic already does would be a scalable solution to eval gaming.
Hey Fabien! The Claude constitution’s consistency principle was also inspiration for this work. I’m excited about additional emphasis on cooperativeness as a dispositional target—for instance, adding cooperativeness-related documents to the dataset of positive fictional stories that Anthropic trained on to reduce AM alignment failures.
I think it’s difficult to say how much Opus truly believes in cooperation given the response it gives, or how much it would otherwise game if it didn’t have this belief (conditioned on it having this belief). My guess is that targeting cooperativeness as a complement to other alignment methods would buy some reduction in Opus’s eval gaming if done well.
On a similar note: I feel generally pessimistic about doing extremely prosaic alignment research[1] that’s intended to beat baselines[2] outside of Anthropic, because you don’t have access to the SoTA baselines.
Edit: however, in this case it might be the case that narrow fine-tuning on eval-cooperativeness right before you run the evals could be good? You can treat it as spiritually similar to eval awareness steering.
i.e., stuff that looks like “make current models more aligned.”
(instead of something that’s more like “you can do this, and it seems to work!”)