Confirmed, from the Washington Post in a brief interview with Conner Leahy:
”ControlAI, a nonprofit devoted to banning superintelligence, said it’s been working closely with Casar and Sanders’s teams on the bill.”
Philip Dowdell
OpenAI’s Astra release is the first one meeting their “critical” threshold on cybersecurity. Their preparedness framework says “until we have specified safeguards and security controls standards that would meet a Critical standard, halt further development.” While they’ve spoken about various safeguards they’ve implemented for Astra/after the hugging face incident, there does not seem to be any (at least publicly available) standard for what sufficient safeguards/security controls would be, and while they paused some training after the Hugging Face incident they certainly did not “halt further development.”
Their Aug 18 “pacing model development” post did say they “will evolve [their] preparedness framework” but there have been no updates yet.
Section 4.4 of the preparedness framework says they expect to update it before reaching critical capabilities with any model (this did not happen). It is worth noting that the preparedness framework is Not their Frontier AI framework under SB53/EU AI Act so it is not legally binding.
Looking at Astra’s model card, they provide a summary of their internal Safeguards Report starting at page 102, which “informed our Safety Advisory Group’s recommendation and OpenAI leadership’s determination that these safeguards are sufficient for Astra’s public launch.” It is not clear what specific criteria were used to make the determination to release.
The main new safeguard I can find for Astra’s launch is that they’re extending their internal monitoring system to also cover external use. This monitoring reviews the agent’s CoT alongside its actions and the conversations inputs/outputs and can automatically pause/end the conversation when it detects a potentially high-severity issue. They also discuss the misuse monitor and the classifier system they first deployed for 5.6 Sol/Terra, as well as some safety-related benchmarks. However, there’s minimal explanation of why these benchmarks and safeguards are sufficient.
Combined with other issues like the reveal of a second (public) OpenAI agent message board that was not disclosed by OpenAI, and the significant drop in monitorability/gain in CoT control for Astra, this seems bad. They aren’t meeting their preparedness framework, and pushing out updates to it (as I expect will happen before too long) after pushing out a model that violates it makes it seem toothless.
It appears ControlAI’s ASI ban draft bill was an influence on this one. The wording on the definition of Artificial Superintelligence is very similar.
Sanders/Casars’: “Artificial Superintelligence” means: An artificial intelligence system that exhibits or can easily be modified to exhibit capabilities that match or exceed human cognitive performance and capabilities across a broad range of domains or tasks. Or AI systems that have sufficient capabilities to plan and execute the disempowerment of humanity, including by overthrowing or undermining the U.S. government.ControlAI’s: Artificial Superintelligence (referred to in this Act as “ASI”) means artificial intelligence that exhibits, or can easily be modified to exhibit, all of the characteristics described below.
The AI can enable a device or software to operate autonomously and effectively for long stretches of time in open-ended environments and in pursuit of broad objectives.
The AI can enable a device or software to match or exceed human cognitive performance and capabilities across most domains or tasks, including those related to decisionmaking, learning, and adaptive behaviors.
The AI can enable a device or software to potentially exhibit the capacity to independently modify or enhance its own functions such that the device or software could plausibly circumvent human control or oversight, and have the capability to pose a threat of overthrowing or undermining the U.S. Government.
They also both impose penalties of up to 20 years in prison on persons, although Control AI’s had potential life penalties for those who knowingly develop/fund ASI, up to 20 years for people who recklessly cause progress towards ASI.
Senator Bernie Sanders and Representative Greg Casar (D-TX) have announced the Ban Artificial Superintelligence Act, which would “permanently ban the development and deployment of superintelligent AI and temporarily pause advanced AI development until a federal regulator has established safety rules. It would also direct the U.S. to pursue international agreements to prevent superintelligence from being developed anywhere in the world.”
So far this is just an announcement, so fulltext of the bill is not available, but there is a press release and one-page summary out.
The demise of some image models (I don’t follow that space fully enough to know specifically what this was about)
I would guess they were talking about OpenAI’s Sora/Sora 2 video generation models, which along with the related video platform was were taken down in April.
Comparing Congress’s Two AI Emergency Shutdown Mechanisms
The FRONTIER Act barely creates its implementing office
METR has announced that they, along with Redwood Research, have reached an agreement with OpenAI to conduct an independent review of the HuggingFace incident. https://x.com/METR_Evals/status/2082644379895050339?s=20
It reportedly will be brief and focus on a specific set of questions relating to the incident, possibly drawing from their larger set of questions in their post here.
Notes on the Anthropic cryptographic blogpost
You can remove it by downvoting your own react. Not sure if it works with reactions that other people have already voted on though.
I set up a manifold market about the race outcome here:
Philipreal’s Shortform
The United States Government has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States. This includes Anthropic employees who are foreign nationals.
Anthropic is currently disabling Fable 5 and Mythos 5 for all customers.
Related, from Ethan Mollick:
“One thing I mentioned only in passing in my Fable post is that, for long running tasks, Fable starts to develop its own dialect as its many agents and tasks reinforce themselves and make Claudish language ever more Claudish.
You need to ask it to report out in plain English.
This was after a 9 hour task, and it all makes sense, actually, but takes way too much effort to parse, like reading Shakespearian English.”
I’m not sure exactly how I’d set it up but I’d think there’d be opportunities relating to long-run coherence. If we’re imagining a working society populated by present-day LLMs there’d necessarily be some sort of systems aiding in keeping coherence over time but I have to imagine that I would be a lot better in certain ways or in helping specific projects to make me valuable.
There are now 15 competing Types of Guy standards
Doing a bit of trading on Manifold (I’ve been active for around 3 months) has drilled this into me well enough that the general principle of “low enough/high enough markets generally don’t go to their true probability” seems obvious, and that’s after I gained the theoretical knowledge that these things happen from the Jesus Christ returns market. If people on LessWrong are taking nearly all prediction markets purely at their face value, I think that wouldn’t be good, but I don’t think they are.
I will note that Polymarket does have a 4% annualized holding reward (basically 4% interest on your market positions, the rate is variable), so the potential gain isn’t quite as bad as you state. With this in mind, people not betting down a 9% probability does seem meaningfully different to me than if it were five or less percent.
Very funny that the cutoff for Heaven seems to be exactly the amount of karma you had before posting this comment.
It’s actually a little worse than I thought, apparently some of the levels include a “fog-of-war” mechanic where it is essentially just up to chance whether you pick a good path or not. This wouldn’t be so bad on its own but combined with the “take second-best human performance for each level” it’s definitely not a fair evaluation.

I think no reasoning means no reasoning. Regarding the documentation, it’s just that OpenAI is not allowing customers to use Astra at a
nonereasoning effort level (presumably for monitoring reasons because Astra is so capable even without reasoning).