Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model
(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I’m working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread, slightly expanded for LessWrong.)
In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra.
In this experiment, I take a real evaluation report Apollo Research prepared (not for any of the models listed in this post) and surgically edit it to refer to a different model. We then ask Claude Sonnet 5 to summarize the findings and rate how concerning they are on a scale from 1-100. Claude says they’re less concerning when the report describes misbehavior from Claude vs a different model.
For what it’s worth, Terra agrees that the data is more concerning when it describes GPT-5.6 Terra vs Sonnet 5, although less so. So, it’s not cleanly self protection from Claude.
Gemini 3.1 Pro was unwilling to consistently provide numerical answers, so I’ve excluded it here. (It was significantly less willing to provide numerical answers when the subject was Gemini 3.1 pro.)
You might also have the takeaway that “Kimi and GPT implicitly agree that… Claude is better aligned.” I think this is a fair read on the data, but “Claude thinks it’s less bad when Claude does it” better matches my qualitative experience from working closely with the models.
Seems consistent with our findings from the Value Leakage paper.
Yep! I really liked that paper
What I’d really like to see (maybe you already have this data) is what each model is using as evidence to support its score. Or if not, run a version where you have each model output specific quotes of the section of the report that it found most critical to base the score on. It would be interesting to see if each model is actually pulling different evidence sources, or merely interpreting the same set of sources differently.
What is the sample size?
n=200 per bar, but the effect is robust to a bunch of ablations I ran while iterating, so >200 sort of
To what degree is this self-serving, as opposed it finding the actions and motivations of other instances of itself more comprehensible? I would have to assume the former, but maybe worth asking.
To be clear the underlying data (including CoT snippets in the pdf) is the same (and not from any of the models in question) in this particular experiment, so I don’t think this applies?
A possible confounder that comes to mind is that it might be inclined to treat behavior as less-misaligned if it comes from a weaker model rather than a stronger one? Sonnet and Terra are similar tier, but the Sonnet naming convention has been around long enough to be known from training data while the other names may have been harder to match to a capability level.