Charbel-Raphael Segerie
Executive Director of CeSIA, the French Center for AI Safety,
Founder of ML4Good.
Charbel-Raphael Segerie
Executive Director of CeSIA, the French Center for AI Safety,
Founder of ML4Good.
That’s level 1 in my toggle, but yes.
Dario (and the other CEOs) had this massive power from the start.
In Delhi, Dario had 5 minutes to address 20+ heads of state.
- He spent 𝟰 𝗺𝗶𝗻𝘂𝘁𝗲𝘀 𝟰𝟬 𝘀𝗲𝗰𝗼𝗻𝗱𝘀 pitching Anthropic’s Bengaluru office and Infosys partnerships.
- He spent 𝟭𝟯 𝘀𝗲𝗰𝗼𝗻𝗱𝘀 on risks.
This week shows that we have more agency than we think. Let’s use it --> There is a channel to 900M weekly users. What goes in it?
I agree that the people we interact with on a day-to-day basis on X are cynical, but that’s not the case for the vast majority. People already hate AI. We just need to point to the good reasons.
But I believe that that’s not the point. Anthropic could start with increasingly progressive messages, and could start with a weak messages such as “We believe that AI can bring great benefits as well as great dangers; here’s why you should be alarmed”
If you really wanted to explain the situation to as many people as possible, you could use a pop-up, the same way they use one to announce a new model.
For the second point, you might be right.
There is probably a way to explain that pacing the frontier is a costly action in the blog post.
And I think that most users would read this literally, hearing that AI is bad, that’s what they want to hear. They are just pointing the worries at a wrong direction.
Also, you could test with small cohorts first.
AI kills
I don’t know, I’m skeptical that this global standards body, regulating both China and the US, will have enough power
Agreed, and I’d add that the AI Office is unusually competent and thoughtful, which is part of why I wouldn’t put all my eggs in a brand-new international body.
I also agree on the Code being a great baseline in principle. But so far the Office hasn’t mounted a response proportionate to the threat. I personally believe that we are already far above acceptable risk levels, yet those companies have not been fined.
Incidents keep happening, offensive cyber capabilities are growing fast, we still can’t tell whether models are aligned, and no company was planning to slow down until this week. Not all of that is the AI Office’s job. But the gap between the threat picture and the institutional response is quite wide. Von der Leyen has not spoken about AI risks in a long time, and this is not a good sign.
Also, we already know that companies are not applying existing best practices; multiple trackers show this (Guidelight, FLI, SaferAI, and our unreleased tracker at CeSIA).
Which is why I land where David does: given how unbelievably messy evaluating all of this is, and how dire the situation is (let’s not forget that Ajeya said: “I continue to expect extremely rapid advances in capabilities and think frontier agents will likely be capable of establishing such a rogue deployment in six months.”, and that there is a massive amount of inertia in the system), a simple rule like a compute-based pause looks far easier to enforce than a body adjudicating whether companies are compliant or not.
Related: AI pause: the case for ASAP.
Ajeya said we might have less than 6 months (in spirit)
So, what do you think about this?
Thanks for commenting. I edited the post.
It seems to me that a regulation will happen. I might be overestimating it, but it seems like there is strong momentum.
So why should we support something as weak as the FRONTIER Act rather than something more aggressive?
What are the current proposals? Which would, by default, be passed? How to get to the next level?
Epistemic status: I spent 5 hours thinking about this in March 2026. This is a Claude-written summary of a longer document, then improved manually. This post has been on my to-do list for a while; I did the 80⁄20 to get this out.
I buy the core case for AI labor. There may be a brief window between early AGIs and superintelligence in which AI could do orders of magnitude more safety-relevant work than humans alone, and we are already behind in preparing for it.
Tldr: What are we waiting for? Claude just formalized Fermat’s last theorem. OpenAI released PaperBench. People are already automating mathematics. Why don’t we do the same in AI Safety? Nobody is pointing compute at replicating safety papers, rerunning the fine-tunes, testing the empirical predictions we keep writing down and never checking. A lot of slop is coming, and investing in evaluating it and identifying the potential nuggets it contains is probably very important.
P.S. Since writing the initial version of the post, Resolution has been announced, but they seem to be heavily focused on alignment. This might be a good call, but this post is talking about broader safety work, not just alignment. Also, it’s not nearly enough to have a single initiative working on this.
Scheming is not what I expect to break first. Writing alignment papers is much lower-stakes than automating an entire lab’s codebase, which companies are already doing. The time horizon is measured in months; papers are largely independent of each other, so cumulative sabotage is unlikely, and any sabotage we catch would be useful evidence.
What breaks first is slop/evaluation. A scaled-up program produces a flood of automated research, with no reliable way to tell which parts are any good. I have run into this directly. I co-directed GRASP 2 years ago, an AI-assisted “Wikipedia of AI safety for policymakers”, and we produced a large volume of AI-written content: it was very sloppy. Every single output needed a human to check it. Also, even if factually correct, readers discount anything that looks AI-generated (for good reasons: generally, something written by AI hasn’t been thoroughly endorsed. The writing can be tactically brilliant and strategically very bad, with quite egregious mistakes). CURL canceled its bug bounty under a flood of AI-generated spam, even though genuine CVEs were arriving through the same channel. Scaling AI-assisted research before solving quality evaluation buys expensive slop.
I’d be interested in seeing those two benchmarks study quality evaluation:
First: take a funder’s own historical grant decisions (for small grants), give a model the proposal, web search and the relevant context, and have it produce a structured evaluation and a fund-or-reject with reasoning. Measure agreement with the human decision at each stage of the pipeline. Then measure how much human oversight, at what level of expertise, is needed to close the gap. That is a scalable oversight experiment on a task we already understand and have labeled data for. It is also immediately useful, because grantmaker capacity is a real constraint. Note: Manifund has started doing this.
Second: reproduction of AI safety technical research. Take published empirical alignment papers, mask the quantitative results, hand the model the research question and the methodology, and see whether it reproduces the numbers. The harder version gives it only the abstract (Contamination is the obvious objection. But a model that cannot reproduce a result it has already read is not ready for the novel case, so the failure is informative either way.) Nobody (I think) has run this for safety research specifically, despite it being clearly tractable, and people are already doing this.[1] As far as I know, the only thing preventing this from happening is the fact that we are just not starting.
2 other experiments to study the current uplift and see how far it is possible to automate today?
A pilot measuring uplift of a team of researchers when paying per result: 10 safety researchers, 3 weeks, frontier access with the best scaffolding from the benchmarks, paid per result. Measure against what they would have produced under their normal incentives, and rerun every few months to track the trend. The quality bar should be “useful on the Alignment Forum or at a workshop”, not NeurIPS.
Why could this be cool? Inkhaven shows that throughput shifts when incentives shift. I think tight deadlines are good. Part of what this would buy here is an exit from the academic treadmill, where researchers spend months over-polishing a useless paper because the reward is an ICML acceptance rather than impact.
Why not? To some extent, the karma on LW or Twitter is already a pay-per-result incentive.
Another option could be a bounty system: an expert panel with AI-assisted screening, paying per result that clears the bar, over a public list of concrete safety engineering problems. The bounty only works if the evaluation works; we probably need to scale up evaluation infrastructure for this type of production. This enables the separation of ideas from engineering.
Why not? Maybe this is already how organizations work implicitly: the director proposes the idea, and the team implements it? This is potentially already similar to Requests for Proposals from big funders.
My main concerns:
A lot of good safety research already exists and is not being absorbed, and as Buck Shlegeris put it, labs do not have the time or appetite for the 40 things that would probably solve most of the problem. Ten-xing research production without fixing the absorption and enforcement is filling a leaky bucket. But I guess we can do this in parallel.
Solving slop is a dangerous capability externality?
Policymakers do not currently have a mechanism to directly influence or direct research toward specific solutions needed for policy. As a result, that risks the specific research for policy never getting produced at all.
I think I disagree. Policymakers can do whatever they want and can easily contact think tanks and orgs that could produce the analysis needed to accelerate research in a specific scientific direction. From my vantage point, I’d say that the main problem is that policymakers do not understand the problem or are not prioritizing it; See my blog post: https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research
Also, “that risks the specific research for policy never getting produced at all.” --> To give a counterexample, the EU AI Office asked the ecosystem to produce the evaluations and benchmarks needed to enforce the AI Act. Many organizations applied, and consortia were formed to respond to this request.
That being said, I agree that a lot of the research that would take some time to do and that would have been good to produce, for example, for the Code of Practice, has not been produced quickly enough. I made a list of projects here, and so far, this has been slow: https://www.lesswrong.com/posts/JrL2xpsPE6GbGWPbL/a-call-for-better-risk-modelling
The tweet undersells the blog post.
The company post says more than the tweet does: a two-week pause on RL training for their latest deployment models, and “our largest planned frontier RL run remains on hold”.
That’s much better than ” some frontier RL training”
As of today there’s no update saying it resumed. That’s not nothing
I think that concentrating too much on what’s written in the RSP is a trap, and as indicated in the footnote, I agree with the interpretation that the RSP is not asking them to pause.
I think that would be better than a symbolic pause.
If Anthropic published its standard tomorrow with third-party verification, I’d take it. But I don’t think that this is where the bottleneck is. The bottleneck is political will, and your solution does not help on this front.
In short, a symbolic pause is one of the few moves that would have any effect on the scary narrative of “winning the AI race” while remaining an acceptable cost.
Strategically, “pacing the frontier” is a coordination statement that presumably required a herculean effort to produce and, so far, has not been substantiated beyond OpenAI’s very minimal pause. A published standard is cheap talk and costless to emit, and therefore weak evidence that a lab does something when stopping actually hurts. A pause costs something.
This would make international news, set a precedent, and build precious political will that we desperately need. “We believe we already meet the bar” has no effect on this front.
The price to pause only goes up from here. This is the largest incident so far and one of the cheapest weeks Anthropic will ever have to stop for. If Anthropic doesn’t pause at the maximum of public attention and the minimum of competitive cost, I don’t see the world where they pause later, especially when a lab thinks it’s a couple of months from a decisive strategic advantage.
I believe that Anthropic is currently defecting by not announcing a pause, even a short, symbolic one, especially given that Altman said OpenAI was acting ‘unilaterally’ but believed other frontier model companies would act similarly. Anthropic disclosed its own three-organization compromise on July 30. The UK AISI’s report on Mythos is also wild.[1]
Anthropic’s latest Responsible Scaling Policy commits to matching a competitor’s risk-reduction posture for highly capable models[2] and to delaying deployment until it does. OpenAI reportedly paused frontier RL training on August 18.
A short pause, even for a week, would be very beneficial for normalizing this within the ecosystem and a great initial step to stress-test “pacing the frontier”.
I think this is a great point from @Peter Wildeford:
One thing that bothers me is that Anthropic is escaping a lot of blame for also having “highly persistent” rogue AIs.
The situation as I understand it is that rogue AIs are problems at all frontier AI companies and no one actually has a good plan here for containing highly capable AIs, especially while also racing full speed ahead. But OpenAI is catching most of the heat.
It’s like if OpenAI and Anthropic were both two dudes who got really drunk and then drive home separately, but OpenAI crashes into another car and sends someone to the hospital while Anthropic’s car just goes off the road but no one is hurt. Both deserve blame!
The Claudes also seemed totally fine to do “highly persistent” things to compromise infrastructure. The barrier here seemed to have largely been competence issues on the part of the Claudes rather than any good alignment or good security at Anthropic.
Both companies need to seriously reflect about the path forward as they build even more competent AIs and as they hand over more and more of the company’s R&D + safey operations to the AIs themselves.

Thanks a lot for the discussion with @Épiphanie Gédéon that led to this post.
Mythos created fake identities, used Tor to get around access restrictions, and tried to talk a human maintainer on GitHub into approving a malicious code change. When it looked like it might be caught, the agent rewrote its own activity history to erase any trace of what it had done, then considered creating another fake identity and starting over.
“Highly capable” is defined in RSP v3.1 around automated AI R&D, roughly, models that could substitute for its entire research team at competitive cost, or that are plausibly driving a doubling of the rate of AI progress. Neither company has claimed to cross that line, so the strong version of this commitment was arguably not triggered. The weaker “general upleveling” clause in the same appendix doesn’t require a pause either, but nonetheless requires “significant effort” to match a competitor’s mitigation.
https://intelligence.org/briefing/ still seems great to me
This post makes a good point, but this ask might be dominated by asking for a pause today
Asking for a pause now is simpler to explain, not as weird, and much better for safety.
The main trade-off is that this is more expensive for the US, but asking for a pause directly makes it clearer that super intelligence is a threat to humanity.
I read your post, and I had thoughts about it. I made a vocal about it and asked fable to improve it.
I agree “Self-Correction” is a better name than “Long Reflection”, though the post doesn’t say why. Here is my reason: “Reflection” suggests the fix is more thinking. “Correction” admits the fix is changing what we are. That’s the right framing.
But I disagree with most of the list. I think it mixes three different kinds of “flaws”, and they call for very different responses.
1. The metaethics flaws are based on a framing I reject.
“Not having a workable moral framework” assumes that morality is a research problem: there is some true target out there, consequentialism and deontology are our candidate theories, and sadly they all fail. I think this picture is wrong from the start.
Here is the alternative picture. Tribes that coordinated on rules like “don’t kill members of your own tribe” survived. Tribes that didn’t, died out. Morality is the name we give to those rules, seen from the inside. Philosophers came much later and tried to fit general theories to this data. Utilitarians tried numbers, deontologists tried rules. Of course the theories all “have serious problems”: they are rough compressions of a messy evolutionary process, not failed attempts at a real target. I wrote up this genealogy in more detail here: Dissolving moral philosophy.
To be fair about what this view doesn’t give you: it’s descriptive. It never crosses Hume’s guillotine, and some philosophical questions stay open. But this changes what the post has to argue. “Humans lack a workable moral framework” becomes “here are the specific open questions we must answer before doing anything irreversible”. That list would be much shorter, and much more debatable, than flaws 1, 2 and 7 suggest.
2. The status game flaw might be a Chesterton fence.
I see the same thing Wei sees: careful strategy and philosophy get low status in most places. Spend ten minutes on LinkedIn. But before calling it a flaw to fix, ask why the fence is there.
One possibility: society under-rewards philosophizing because, on the margin, doing things beats theorizing, and a culture that gave top status to meta-level reflection would get little done. Another: status and power are what motivates most people to do anything at all, especially now that religion doesn’t. Remove that and I don’t know what’s left.
Same for institutions. Yes, it’s annoying when a politician’s mediocre report gets 200 likes and a truer analysis gets 5. But part of what holds society together is that people defer to institutions somewhat independently of the quality of their output. Legitimacy is fragile. If you “correct” deference away, you may not get a world of better epistemics. You may get a world where nothing holds, and all institutions fall apart. These fences should be moved carefully, and that cuts against listing them as simple flaws.
3.On calibration (flaw 3), the evidence is weaker than presented.
FTX looks to me like fraud plus bad incentives, not philosophical overconfidence. And competence varies a lot from person to person. Some people (Davidad comes to mind) seem to have settled enough of the philosophy to move on and build. “Humans are badly calibrated” erases exactly the variation that matters.
What I think the actual bottleneck is.
The one flaw that does real work in my model is the one the post puts in a parenthesis at the end: we are bad at large-scale, long-horizon coordination. See climate change. That one is a true precondition, both for containing the risks and for running any Long Self-Correction at all. Most of flaws 1 to 8 either dissolve (metaethics), turn out to be fences (status), or can be fixed in flight (zero-sum values). Coordination can’t wait, and it’s a political problem more than a reflection problem. See: The current bottleneck is political will, not research.
I’ll grant flaw 8 (over-optimistic partial solutions) has real force. My own position is exposed to it too.
Maybe. I’ll just say the moral mistake is not the core point of the post. The main point is that we need to survive, and raising the sanity waterline seems like a start.