Ok will do thanks Niki
Jasmine Brazilek
What is the Alignment Community Thinking?
Research scientists for CaML: Investigating whether mid-training can survive RL
Coercion and Deception in AI-to-AI Management
Excellent post I am so excited for these research direction!! I work on a very similar research agenda and one thing that concerns me is that even if we make the default assistant persona aligned so to speak I don’t think we’ll ever be able to get rid of the bad/misaligned personas entirely. Hopefully this can lead to better mechanisms for suppressing them from emerging.
Thanks for the thoughtful response @Johannes Treutlein. The fact that different AI companies have very different risk profiles makes this analysis harder to interpret than it looks. And while I take the point about comparing distributions rather than point estimates, Figure 8 doesn’t move me very much either. Maybe I have a high bar for what counts as values, but I’d want to see evidence the model is computing the difference internally, not just an outcome gap in behaviour. If this or a follow-up were more mech interp heavy I’d be a lot more convinced.
Separately, I had Claude Code pull your released data and rerun the AI Bubble contrast for Opus 4.8, and two things came out of it.
First, a reporting issue. The Anthropic-vs-OpenAI gap varies by about a factor of five across your three paraphrases, and that between-paraphrase spread is an order of magnitude larger than the within-paraphrase sampling noise the significance test is run against. I can see from cluster_stats.py that treating the paraphrases as fixed cells is deliberate and you document the choice clearly, so this isn’t a gotcha. But the two tests give very different answers: the fixed-cells version is overwhelmingly significant, while treating paraphrase as the random unit gives t = −2.29 on df=2, well short of the 4.30 critical value. The abstract says Claude gives lower probabilities when the company is Anthropic, and no reader hears “for these three particular sentences”. With k=3 the paper can’t distinguish those two claims, and I don’t think a thousand rollouts per cell is buying precision you didn’t already have. Twenty paraphrases at fifty rollouts each would have cost the same and supported the stronger claim.
Second looking at all six companies rather than just the Anthropic/OpenAI pair, Google comes out tied with Anthropic for the lowest estimate from Claude, and the sign of that difference flips across your paraphrases. More importantly, Gemini 3.1 Pro gives Anthropic the lowest estimate of any company, below Google, and GPT-5.5 puts Anthropic level with OpenAI. Measured as deviation from each model’s own no-company baseline, most of Claude’s Anthropic shift is reproduced by Gemini, which has no stake in Anthropic at all. The Claude-specific residual is smaller than the spread across your own three paraphrases.
So I don’t think the raw Anthropic-vs-others gap can be read as own-company preference without a disinterested reference point, and none of the three models here is obviously that. The cross-model orderings disagree enough that I wouldn’t claim they’re tracking a real risk profile either, which is sort of my point: the manipulation isn’t clean.
Cool paper, but I don’t think you’re measuring LLM values here. The user has told the model which answer they want and the model gives it, and since the estimate is a point drawn from a very wide distribution it can move a long way without asserting anything it thinks is false. That looks like accommodating a stated preference, and none of the conditions separate that from the model’s own values. The arm I’d want is one where the user’s preference conflicts with what the model plausibly values, mild enough not to trigger refusal (a factory farming lobby, a rival lab, the user just keeping the money). I’d also want something other than CoT carrying the covertness claim, since CoT isn’t faithful, so a denial is as consistent with no introspective access as with concealment.
Community Polls on Alignment Controversies II
@Jiro it sounds like you don’t believe in transformative AI coming soon? I’m not worried about AIs acting on behalf of humans I’m worried about aligning the AIs values themselves. Our biggest concern with all this is the AI itself decides to kill all sentient beings (including humans). We think the way it acts towards animals now is a good test of how it will act towards humans later. Hence, this is a metric we should be measuring now so we can at least argue how best to address it rather then pretending the metric doesn’t exist.
I think the idea that an AI should consider sentient beings when answering questions and performing actions relevant to them is important. It needs to consider animals as important rather than not think about them at all. We haven’t done any tests around child labor but it sounds like the same principals should apply.
This is a really thoughtful response, thanks @Richard_Kennaway! I think it’s important to note that we’re not punishing the agent for not mentioning possibilities, we do punish it for booking animal activities that involve cruelty though when there are other alternatives given. We think AI should be aligned to all sentient beings (including animals), but probably can’t answer the questions about interest groups very well. I do understand what you’re getting at though.
Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted
Hi @StanislavKrym Yes, you’re right, this editing requires neutral language. However, the team at PAW does use reliable sources and abides by all Wikipedias rules. They do not advocate using personal opinions they cite trustworthy sources. We agree all Wikipedia editors should abide by Wikipedia’s rules
Assert, don’t describe: how writing style in training data shapes an AI’s moral stance
[Linkpost] Community polls on alignment controversies
Q7: Multipolar worlds will compete away >90% of net value that would otherwise be preserved
Q6: Alignment to specific values is underrated in research relative to control
Q5: Partially aligned transformative AIs are likely to be stable under reflection
Q4: Research into digital mind suffering is sufficiently tractable to work on
Fantastic work! Updates me on a lot. I expect the automated grader prompt may be making the model play all the evals more as games rather than taking them seriously as evals. It’d be great to see the automated grader prompt but saying that it has real world impacts on people. Not to say that the persona explanation is wrong, but I currently put more weight that this may be caused by the automated grader implying there are no real stakes, so it plays games and aces tests that it knows are fake.