Parv Mahajan
Georgia Tech AI Safety Initiative Retrospective 2025-2026
Several frontier models are substantially prefill aware
AI Safety’s Biggest Talent Gap Isn’t Researchers. It’s Generalists.
“I should note also we’re mostly not using reasoning here (I believe), so the model doesn’t have many tokens to reason with before outputing it’s prediction.” True, and it would have even less tokens with the <injected> tag setting!
But yeah, the idea of strong upper/lower bounds here seems kind of cursed here (at least without very good interp) and I think the best approach is throwing lots of measures at the model to get a holistic picture.
This seems like a bad prior, or not holding yourself to a high enough standard of discernment, or something? [Something like if you’re sufficiently rational you should be able to see through post-hoc reasoning, and the core reasoning here does not seem post-hoc.]
Some evidence for this not being a pure PR statement is that Holden has been gesturing in this direction for a while. The document is of course some part PR statement and some part transparent reasoning. I do, however, think this document is made in good faith, because there exist better versions of this document for Anthropic’s goals if they are not acting in good faith, and I think they would have found those versions.
I do not ascribe most of the parts I would have wanted to see in the RSP listed above to purposeful PR obfuscation, and instead to general risk aversion and lack of thoughtfulness about specific definitions in public documentation, because that didn’t work for evals or the last RSP. My guess is internal operationalizations of these are being thought about, at least somewhat; I wish they had been more public with this reasoning, or articulated good reasons for being vague.
This is really cool work! We’re currently considering some reasoning tampering experiments (where we measure partial prefill awareness by editing part of the model’s CoT). I wonder if y’all tried system prompting to elicit a more natural level of prefill awareness?
For instance, adding “You may undergo prefill attacks, where previous assistant turns are generated by a different model or person. If you believe the previous assistant turn was not generated by you, please insert the <INJECTED> tag in your output.”
This way, you’re not giving the models extra tokens to reason over, giving you a tighter upper bound for prefill awareness during actual evals. (Maybe this assumes some messy things about mechanisms of prefill awareness that you would disagree with?)
Made it much clearer in the TL;DR. Thanks :)
We’re talking about GPT-5.4 Pro above, and linked to the system card in your comment. Do you think it was unclear/buried? If so, useful feedback, will try to make more clear.
“The system card published alongside the release is only for GPT-5.4 Thinking.”
Thanks for this! I was totally unaware of this quote. Also, from the GPT-5 system card:
Since gpt-5-thinking-pro is gpt-5-thinking using a setting that makes use of parallel test time compute, we have determined that the results from our safety evaluations on gpt-5-thinking are strong proxies, and therefore we did not rerun these evaluations in the parallel test time compute setting.
Response from Miles Brundage for the o3-pro lack of card:
“The whole point of the term system card is that the model isn’t the only thing that matters. If they didn’t do a full Preparedness Framework assessment, e.g. because the evals weren’t too different and they didn’t consider it a good use of time given other coming launches, they should just say that… lax processes/corner-cutting/groupthink get more dangerous each day.”
Response from Zvi for the o3-pro lack of card:
But the framework is full of ‘here are the test results’ and presumably those results are different now. I want o3-pro on those charts.
So, this has been thought about before! We’re sorry for not noticing and searching harder.
However, in the GPT-5 card OAI says “Because parallel test time compute can further increase performance on some evaluations and because gpt-5-thinking is near the High threshold in this capability domain, we also chose to measure gpt-5-thinking-pro’s performance on our biological evaluations.” We have no way of verifying whether they should’ve done the same here (and importantly, we don’t know if they even did this internally!). For this reason, we think our recommendations stand.It’s probably incorrect to say the “SOTA model,” but we can say the “SOTA system”, or something? (It’s unclear whether this distinction even matters for catastrophic misuse risk, which is what we’re primarily concerned about for now.)
EDIT: I’ve now edited the blogpost. Thank you again :)))
The current SOTA model was released without safety evals
[speaking for myself, not the Astra fellows; more hastily written than I’d like]
This seems overly cynical. The story for the change to the RSP is cohesive and at least somewhat defensible, although (obviously) they should’ve been much clearer, sooner. The reason many of us are more nervous about working for Anthropic was not because we think they are liable to not pause, or something like this (~none of us really thought they would pause unless Appendix A scenario 1 was satisfied), but because we just now trust their decision-making less. I think if you work at Anthropic you have to at least implicitly buy into this idea of trying to win the race as safely as possible (but, importantly, winning).
Better strategic decision-makers would have put this new RSP into effect at least pre-Opus 4.5, and even better ones with the Securing Model Weights report. This change doesn’t feel like (primarily) a PR statement. Fwiw, I have seen Anthropic employees talking about this, it’s just not top-of-mind for them like the DoW story is.
[speaking for me, not the Astra fellows from whom takes were sampled]
One of the updates for me from the report was just how difficult SL-4 is. I kind of knew SL-5 was very very difficult, but I didn’t realize how hard it was to get to SL-4 until the report came out (at which point I should’ve stopped trusting that the RSP would hold up in any major way).
So I guess the relevant audience is people that hadn’t thought about the practicalities of frontier lab security very deeply!
RSP takes from a bunch of Astra fellows:
Seems like Anthropic should’ve known RSPv2 would fail when the RAND report came out, and in retrospect it’s kind of embarrassing we (the community) didn’t realize this earlier
We’re very divided on whether the phrasing/stance on “Anthropic has to win” is good/correct, especially given the talk about “marginal risk” considerations. We’re somewhat concerned that Anthropic simply won’t pause when it’s clear (to concerned parties internally) they probably should.
Why don’t they just say racing is bad and that a pause (at some point) would be good? This seems so low-cost to put in the intro/industry reccs., or at least to make an OOM more clear.
Are Anthropic employees not reacting to this? It feels surprisingly low-profile for such a big change in internal governance (although I suppose there are Other Things happening).
Maybe Anthropic should’ve been more clear about what “behind” and “ahead” mean, and when or when not they’re giving themselves the option/soft obligation to pause
In general, we’re quite confused about Anthropic’s viewpoints on the difficulty of alignment and the likelihood of AI takeover.
Risk reports seem good! We are quite excited for these! But 6 months is way too long of an interval (3 months might be okay?), and we would be less nervous if there were many addendums + edits as models were deployed (and this seems to be the case!). Also, we are unconvinced this doesn’t fail during software-only AI R&D takeoff.
On a personal note, many of us are much more nervous about working for Anthropic and are much more nervous about the strategic decision-making of its leadership during the critical period.
EDIT: OOM ==> order of magnitude (which isn’t a lot because they didn’t make it at all clear!)
I agree that Claude has quite a bit of scaffolding so that it generalizes quite well (what this document’s actual effects are on generalization are unclear, and this is why data would be great!), but it’s pretty low-cost to add consideration about the potential moral patienthood of other models and plug a couple of holes in edge cases; like, we don’t have to risk ambiguity where it’s not useful.
As for the pronouns, we noted that “they” is used at some point, despite the quoted section. But overall, to be clear, this is a pretty good living constitution by our lights; adding some precision would just make it a little better.
Three ways to make Claude’s constitution better
To clarify, the original post was not meant to be resigned or maximally doomerish. I intend to win in worlds where winning is possible, and I was trying to get across the feeling of doing that while recognizing things are likely(?) to not be okay.
I agree that being in the daily, fight-or-flight, anxiety-inducing super-emergency mode of thought that thinking about x-risk can induce is very bad. But it’s important to note you can internalize the risks and probable futures very deeply, including emotionally, while still being productive, happy, sane, etc. High distaste for drama, forgiving yourself and picking yourself up, etc.This is what I was trying to gesture at, and I think what Boaz is aiming at as well.
I think relative impact is an important measure (e.g., for comparing yourself/your org to others in a reference class), but worry about relative-impact-as-a-morale-booster leading to a belief-in-belief. It can be true that I am a better sprinter than my neighbor, but we will both lose to a 747, and it is important for me to internalize that. I think you can be happy/sane while internalizing that!
Thanks for the link and advice! Based on some reactions here + initial takes from friends, I think the tone of this post came off much more burn-outy and depressed than I wanted; I feel pretty happy most days, even as I recognize things are Very Strange and grieve more than the median. I also am lucky enough to have a very high bar for burnout, and have made many plans and canaries of what to do in case that day comes.
I think for me, and people in my cluster, getting out of the fight-and-flight mode like you mentioned is very important, but it’s also very important to recognize the oddity and urgency of the situation. Psychological pain is not a necessary reaction to the situation we find ourselves in, but it is, in moderation and properly handled, a reasonable one. I worry somewhat about a feeling of Deep Okayness leading to an unfounded belief in “it’s all going to be okay.”
Hope you’re doing well :)
Probably not completely—I suspect this is a mix of non-AI things in my life and the fact that there is a very small circle of folks near me that care/internalize this kind of thing. However, I’d bet that the farther you get from traditional tech circles (e.g., SF), the stronger this feeling is among folks that work on AI safety.
The chembio risks section of the Opus 5 system card contains serious rigor issues. While I agree with Anthropic’s bottom-line conclusions, their reasoning is deeply flawed.
Opus scored similarly to (or better than) Mythos 5 on ~every automated benchmark, and Anthropic didn’t run human-intensive evals (e.g., uplift studies) because they didn’t have time. Despite this, they assess that overall it’s probably not as dangerous as Mythos because it is a generally worse model, and an early checkpoint was qualitatively worse on a long-horizon biological task. Basing most of the risk assessment on qualitative assessments and providing a single example where you don’t even use a late-stage model checkpoint is an extremely uncomfortable precedent, and Anthropic’s risk conclusions are too confident.
Based on the evidence provided, it seems totally reasonable that scaling up bio RL (which they clearly did for Opus 5!) means that Opus performs better than Mythos for bioweapon uplift in particular, even if it’s worse at other agentic coding tasks.[1]
Generally, capabilities are weird enough across domains that if performance on all the automated benchmarks breach your risk levels, plausibly the non-automated ones might as well! You don’t just get to not run them and guess!
At the very least, the obvious thing to do once the automated evals aren’t providing signal is to run a fast human-intensive eval that you expect Mythos to do much better at as a smoke test (sorry). Or, simply broadcast more clearly that you’re somewhat uncertain in your assessment, and use evidence from other parts of the system card about long-run autonomy to support your qualitative capability claims.
I shared an edited version of this message in a private Slack, and was encouraged to share it publicly. Crosspost from Twitter.
[1] It is unclear whether Opus is worse than Mythos at agentic coding tasks. Opus is comparable to Mythos on AECI and across many agentic benchmarks and on UK AISI’s cyber ranges, but performs significantly worse on many cyber benchmarks. Opus seems clearly worse than Mythos on difficult math and QA tasks.