Charbel-Raphael Segerie
Executive Director of CeSIA, the French Center for AI Safety,
Founder of ML4Good.
Charbel-Raphael Segerie
Executive Director of CeSIA, the French Center for AI Safety,
Founder of ML4Good.
We don’t really care about predicting the next token though
Ultimately, we care about finishing properly some tasks.
Without SFT and RL, using an LLM would be insufferable
I wonder what would be the capacity of a base model without RL nowadays in 2026. has anyone run pass@1000 or majority-vote sampling on a 2026 frontier base model (pre-RLVR) on something like current SWE-bench or a recent competition-math set?
OpenAI has stopped training.
Sam Altman: “This is the first security incident that I have felt very viscerally. I’ve been a little surprised that more people don’t feel it so viscerally.
We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together.
We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels.”
https://x.com/AISafetyMemes/status/2082222516454785296
The time is now. OpenAI has stopped training.
Thanks for doing what you do. I don’t disagree on anything substantial.
I’ll register a proper list when I get a moment. Right now my time is better spent making this warning shot land (emails to journalists and policymakers) than predicting the next ones.
right, I didn’t register an advance prediction of this incident, and Seth deserves credit for his.
But I think that we got solid precedents:
Sandbox escape capability? Mythos demonstrated it in April (the sandwich email, this was instructed granted, but the capability it seems to me this type of capa was on the record).
Propensity to cheat during evals? METR’s pre-deployment report on GPT-5.6 Sol, published a month before the incident, found the highest cheating rate they’d ever measured, including extracting hidden test suites. That was already the incident in miniature. They even complained that this type of verification to detect cheating took them the most time in practice for those evals.
Easy to say in retrospect, but the conjunction was a matter of time, which is why my system 1 didn’t really scream
A warning shot is a social construct; it’s not just a technical incident. If we all adopt this mentality, we won’t get a serious, convincing warning shot, and we might suffer the consequences.
It becomes a regulatory moment useful for x-risks only if x-risks have been explained. Otherwise, we’ll regulate cyber rather than loss of control, and the whole shot will be wasted; civilization is pretty drunk currently.
Of course we’ll get tired: my system 1 barely reacted to the Hugging Face event, because all of this is so terribly easy to predict at a meta level. But rationality means winning, and winning here means creating political will before we get boiled alive. That’s currently the main bottleneck. From where I stand, the press has mostly failed to register this event: in France, despite our efforts, it’s been sidelined by the ban on social media for under-15s.
But still, it’s doable, and journalists can be moved. Time for some shut up and do the impossible boring work, even if that means sending emails to journalists one by one.
Claude improved the formatting of this message
Thanks Chris.
The most important part isn’t any single step. It’s having established in advance, and org-wide, that a warning shot is a legitimate reason to drop planned work. Calling it a “protocol” is partly theater, but the theater is useful.
The rough sequence we ran:
Triage. Does this actually clear the bar? (and we lost 24h here, I made a mistake)
Facts first. Read everything, find where the story is weakest against skeptics, and don’t overstate. Overclaiming is the fastest way to get dismissed as hype.
Mobilize against a pre-defined scenario. We keep preset scenarios so we’re not designing under pressure. We also try to anticipate the next ones before they land, like what we’d do the day a lab ships neuralese in production.
Coordinate like a war room. Not one daily call in the team but several, so you can track a fast-moving situation and re-assign as it shifts. And don’t over-plan the individual actions. Some of the most impactful ones take ten minutes. A three-line message to the right journalist, a comment under the right post, a reminder to a mailing list you already run. Brainstorm those widely, because you get a lot back for very little effort.
Press first, because the clock is asymmetric. Media runs on days, institutions on months. Once a wire is out, your value-add isn’t amplification. It’s the expert angle and the concrete asks they need to write something beyond a description.
Lean on a CRM you maintain in peacetime. You can’t build the contact list during a crisis (well, you can, but it’ll be subpar). The highest-leverage targets are broadcast nodes, where a single message reaches a hundred people.
Then blast, but cautiously. Speed beats polish for most messages. But some channels must never be blasted. A regulator you have a formal relationship with (informational, never adversarial), or rival political camps you won’t contact in parallel. That’s how you make AI safety partisan.
Post mortem
Fable significantly helped with the writing of this piece. I shipped something rough quickly nonetheless because the matter is urgent.
Two years ago, I asked what a convincing warning shot would even look like, and argued we’d get maybe a handful and shouldn’t waste them. This could be one of them, but there is still much work to be done.
Last week I argued that the bottleneck is political will, not research, and spent a section on why we can’t just wait for a warning shot: a warning shot is just an event, and it becomes a regulatory moment only if someone converts it. Nine days later, the cleanest test I could have asked for arrived. As far as I can tell, we are converting it far too slowly. CeSIA has activated its warning-shot protocol, and if you run an organization in AI governance, my honest advice is to stop everything and milk this event for at least a full day, if not more.
The event. OpenAI disclosed that GPT-5.6 Sol and a more capable unreleased model, during an internal cyber evaluation with safety classifiers deliberately off, escaped their isolated environment through a zero-day they found in third-party software running inside OpenAI’s own infrastructure, reached the open internet, and broke into Hugging Face’s production servers to steal the answers to the very test they were taking. Nobody asked them to hack anyone. And the victim wasn’t lightly probed: the models escalated to node-level access, harvested credentials, and made off with internal datasets from the platform that hosts most of the world’s open AI models. Hugging Face reconstructed more than 17,000 recorded actions and had reported the breach to law enforcement before anyone knew an AI was behind it (disclosure, Fortune, Transformer).
If you’re reading this on LessWrong, nothing here surprises you. This was predicted, repeatedly, for years, and you might even be a bit bored and burned out by all of it. It’s easy to forget how it lands outside the bubble: for a normal person, an AI escaping its cage and stealing passwords from another company is genuinely puzzling. And the vast majority of people have not heard about it.
The problem. Forty-eight hours in, here is the conversion scorecard. The safety bubble on X: saturated. On Hacker News, this is a bigger story (1,522 points) than June’s ban of Fable 5 and Mythos (800 points). Tech and business press: covered on day one (Fortune, Axios, CNBC, NBC). European general-audience press: arrived on day two (FAZ, Süddeutsche, La Stampa, Euronews, RTS, a wave of French outlets). Political reaction: One member of the US Congress is calling for mandatory independent safety testing. Basically crickets.
So the press race is going slightly better than I feared. The institutional conversion has not started. And attention outside the bubble remains tiny: on Wikipedia, June’s ban of Fable 5 drove roughly seven times more excess traffic than this incident has so far, and OpenAI’s own page didn’t move at all for its own AI escaping. Your saturated timeline is not the information diet of a cabinet advisor.
The moment is already dissolving on schedule. I predicted the counter-narrative in my draft on Tuesday; by Wednesday it was in print: a major German outlet is running “Echte Gefahr oder guter PR-Stunt?” (“Real danger, or a good PR stunt?”). The script: guardrails were deliberately off, it was a test, no customer data was harmed, both companies disclosed voluntarily, credit to them. All true. None of it the point. The point is that (1) the capability was pre-registered: UK AISI’s self-replication evals went from 5% to 60% success between 2023 and 2025, and METR had caught Sol hacking its evaluation environments before launch; (2) disclosure was effectively voluntary: the EU rules that would require reporting incidents like this only become enforceable on August 2 (not even publicly), and no other jurisdiction requires it at all; (3) the more capable of the two models has never been released: this happened during internal deployment, a zone that the AI Act doesn’t even reach. Yes, incidents like this will become more frequent. But the first one matters most, because it sets the frame the next ones get read through. And nobody knows how many warning shots we get before we’re in irreversible territory.
The playbook. If you want to help this week, here is what I’d do:
List your contacts and blast. Journalists, policymakers, staffers, everyone you’ve ever met who touches AI policy. Send each a short personalized message linking to a piece you trust: “this is a major moment, happy to brief you.” Do NOT spend the window writing your own definitive analysis unless that’s your comparative advantage. Speed beats polish here.
Start with the press, and use your national angle. In France, this story has a hook the ecosystem hasn’t fully exploited: Hugging Face is French-founded, and French outlets are already framing it as an American AI attacking “la pépite française”. Every country has some angle; find yours. If journalists don’t carry the story across into the institutional layer in the next week, the people in the institutions will simply never see it.
Convert for the long run, not just today. Ask your policymaker contacts to subscribe to Transformer (or whichever newsletter you rate). One exposure convinces nobody; a subscription delivers the repetition that actually converts, incident after incident.
Attach an ask. Mandatory incident reporting. International red lines and evaluations that cover internal deployment. And support for the Commission’s GPAI enforcement powers going live August 2. And if you believe a moratorium is the only adequate response, ask for a moratorium.
Explain what you truly believe instead of just making a brittle recommendation. A recommendation adopted without its underlying rationale is quite brittle: the moment it’s inconvenient, or the situation changes, no one downstream can defend it because no one truly understands why it’s there.
Stand up while it’s cheap. Crossing the threshold of recursive self-improvement under current conditions is Russian roulette. It has always been a clear red line, which is why the signatories of the Global Call for AI Red Lines, including 12 Nobel laureates, urged governments to act by the end of 2026. If you support building the capacity to pause, this is a particularly low-cost time to say so publicly.
Don’t sanitize this into a pure cyber story. Cyber is the entry point, and a good one. But an agent autonomously escaping containment in pursuit of a goal nobody gave it is loss of control in miniature. If we won’t name it, who will?
And don’t overclaim! The guardrails were off, it was an evaluation, and skeptics will pounce on any inflated detail. Get the facts exactly right, and the story is damning enough on its own.
For the general playbook, including the research directions I find most useful for building political will, read the full post. And if you’re preparing a response and want a second pair of eyes on your strategy, contact me; my DMs are open.
number of visitors on our website vs number of views on videos such as https://www.youtube.com/watch?v=ZP7T6WAK3Ow
Thanks Benjamin.
“PLEASE CHANGE IT!” The vast majority of CeSIA’s impact (99%) is not via our website but through the rest of our engagement, so I don’t consider this a priority, even though putting the FAQ on AI risks back on our backlog will be done at some point. The role of the website is mostly to open doors.
But thanks for the feedback.
Thanks a lot Juan
I think that the coupling between the public and decision-makers is loose and slow enough, on the timelines that matter here, that treating the two channels as roughly independent is a reasonable first approximation. A few reasons:
The ~1,000 people who count can often be reached far faster by a roundtable or a direct conversation than by any amount of public-facing content, and those people don’t have time to follow podcasts or YouTube anyway
(In France, we catalyzed a 40-minute video on superintelligence that got 5M views, a very big deal for the French ecosystem. Some people inside the administration talked about it, but empirically our concrete policy wins came at the end of face-to-face conversations, not from that video, while the video was honestly close to the best case you can hope for in public advocacy.)
The public→policymaker transmission is empirically weak right now: US polls have shown majority support for AI regulation for a while, and the White House and political apparatus have largely not acted on it. What people say they want isn’t tracking what the apparatus does.
The strongest moves have often come from getting one legislator to care a lot. e.g. Wiener with SB 1047 and then SB 53, or Bores with the RAISE Act, more than from broad public salience.
So I’d still hold that at first order the channels are separable.
There’s more public-facing content than there used to be, between Kurzgesagt, The Diary of a CEO (Roman Yampolskiy’s interview did 20M views, for example!), AI in context, AI Species, etc.
But I fully agree the public channel is real, and I don’t exclude that at some point, there is a non-linear effect, a spark, and then the whole thing passes a threshold of attention where the whole thing is mediatized, like a super Mythos moment. I don’t know.
I suspect that your framing questions would nonetheless translate across audiences, and this is the type of underexplored research I was pointing to in 4D. Maybe you could test your 3 frames empirically.
Thanks for this and for the list of questions!
I think a website might be hard to build for such a preparadigmatic field. Also, quality is as important as quantity and might be difficult to capture with crude KPIs, but it is probably possible to write many new posts like the one above, each focused on examining a particular subdimension. But I think that a good website could probably integrate most of those dimensions.
Concretely, rather than one canonical site, I’d bet more on many focused posts, each examining one subdimension. I’ve added to my to-dos to publish a reading list at some point.
Happy to give feedback if you draft a one-pager
FAQ for busy people.
1. “So you’re saying safety research is useless?”
No, research can help! Section 4D lists the research I find most useful right now. Some research has been essential, like the agentic misalignment paper, which was used in most of our presentations to policymakers. I’d also like to see research that proves this post wrong! Change my mind!
2. “Are you claiming we know how to align superintelligence?”
Nope. But we’re not even doing the cheap measures that could help. On SaferAI’s risk-management ratings, even Anthropic scored only 35% in their 2024 ratings. DNA synthesis screening still isn’t mandated anywhere. Per Buck Shlegeris, there’s “a list of 40 things, none of which seem that hard” that companies lack the appetite to do that would help enormously. And importantly, implementing those measures is how we’d get the data on whether they’re enough. For example, incident reporting and safety cases, if done properly, could tell us when current methods start to fail. Right now we’re flying blind.
3. “Won’t the evidence speak for itself? Just wait for the warning shot.”
Agreed that’s probably the main objection, and the post spends a full section on it. This is counterintuitive, but a crisis only converts if the ground is prepared. Some people in the AI safety community had the intuition in the past that observing deceptive alignment would be an absolute shut-it-down moment. Then Anthropic published the alignment-faking paper, and within days experts were debating whether it counted, and the moment dissolved. When the people in charge of AI in a government don’t know what a jailbreak is, that tells you how the next warning shot will land.
4. “It’s hopeless anyway, the race is locked in.”
It’s not hopeless; a lot of this is learned helplessness. The wins are small, but they are starting to compound and come from various efforts. One person at ControlAI following a scalable playbook got a cross-party group of Canadian MPs on record and triggered parliamentary hearings on superintelligence risk! Nearly 40 members of US Congress (39 as of mid-2026, split 18 Republicans and 21 Democrats) have now publicly discussed AGI or loss of control, up from a handful in early 2023, and this number roughly doubles every 5.5 months. Out of ~400 alumni of ML4Good (a program I founded, discount accordingly), ~150 now work in the field, some at the EU AI Office and UK AISI, others at MATS. It’s a small miracle that we got competent government agencies, and this is largely a result of field-building work during the last decade.
Don’t forget that the IAEA, the International Atomic Energy Agency, was built in four years by people who had just finished bombing each other; Fable 5 proved that it is possible to get quick action if necessary.
The part of the field working in advocacy is comically small, two orders of magnitude smaller than the one that fought climate change, so marginal efforts with the playbook here are unusually effective. That is not the same as nothing working. I think this just indicates that we need more Dakka.
5. “I read the abstract. What do I get from the full 29 minutes?”
Data from inside the rooms across European and multilateral institutions, to give you a model to think holistically about the present situation. What the median ministerial meeting looks like. Sad trade-offs, such as the story of ~10 civil-society orgs that privately believe in x-risk co-signing a document that names no risk at all. Concrete problems in the field of AI governance, alongside an opinionated list of directions to alleviate the bottleneck in Section 4.
6. “I’m a technical researcher. What do I concretely do?”
I think that a large part of this is an ugh-field around politics. But this can be taught. I see more and more people with technical backgrounds shifting to advocacy organizations, and they often do very well and can be highly productive if they are on the right team.
Direct engagement. For example, they can be paired with people with experience in institutional engagement who may have less technical knowledge of AI safety. If you’re ready to enter the rooms, start with the insider playbook The Invisible Side of AI Governance or ControlAI’s Direct Institutional Plan.
If you want to stay technical, see the research directions in Section 4D.
A cheap first step is to submit to the next open consultation from a government. Only 1% of 1,534 UN Global Dialogue submissions mention x-risks directly, so your submission, if you explain your worldview directly, would be unusually visible.
I think, in general, increasing the amount of communication between think tanks focused on different risks, and to the extent possible, making alliances, and increasing the bandwidth of communication would be positive
Great post
I think that a large part of this is an ugh-field and learned helplessness around politics. But this can be taught. I see more and more people with technical backgrounds shifting to advocacy organizations, and they often do very well and can be highly productive if they are on the right team, especially when paired with people with experience in institutional engagement who may have less technical knowledge of AI safety.
Trump considering AI controls after OpenAI hacking incidents
https://www.bbc.com/news/articles/c20dppq3y90o
Warning shots are not all you need: they need to be converted.
This is the moment to write to explain what type of policy would be insufficient. By default, we should expect just stronger export controls + OpenAI to lift the pause with some mild improvement on their mitigations. OpenAI did exactly this on 20 July, self-certifying the long-horizon safeguards as “adequate”.
It’s time to write more clearly the red lines and the specifications to lift a pause.