#1 feels reasonable and I have bet on it. TBH I don’t have a solid take right now on how likely (or necessary) liability is, but 10% doesn’t seem crazy. However, I don’t think it really gets at what I am interested in here. At the extreme, the Sanders-Casar bill passing alongside an international agreement with China (not to say I consider those things more likely than liability) would not lead to #2 resolving as “yes.” IMO, a strong transparency bill with strong independent auditors would be good progress and worth engaging with politics for, even if it wouldn’t solve everything.
Avi Brach-Neufeld
Great! I do really appreciate you making these markets. A couple clarifications before I bet:
What counts as a White House policy switch? Are we just talking about Trump here? Are we only considered formal actions or are any written or oral statements enough? Does it have to be a stable/longterm switch or is one action/statement enough?
Would you be open to adding additional criteria for the second market? I don’t think legal liability is particularly likely. What I’d like the market to get at is “something worthwhile came out of the political process under Trump.” Maybe we could pick a commentator we both trust and have something like “If _____ deems the passage of a bill or declaration of an executive order ‘significant progress’ or other equivalent terms.” Alternatively if you lay our what you’d consider significant progress and we are roughly on the same page, I’d be fine settling based on your opinion.
What odds would you put on your prediction and what do you count as “good AI policy?” I’d be happy to make a (small) bet.
To be clear I didn’t meant to imply you were being fatalistic on AI overall, just on the prospect of convincing Trump/Republicans. I disagree that we should abandon persuasion attempts after ~1 week of the issue being politically salient. I think you can make lots of arguments about Trump, but “once he makes a decision, no one and nothing can change his mind” seems pretty clearly untrue. In fact I think a flaw of his is he wildly bounces between positions based on the last person he talked to. I’m not overly confident because I don’t think Trump is a particularly rational person, but I think a key difference between AI risk and the things you mentioned is that AI risk has the potential to personally impact Trump.
I also think things get very difficult without government getting on board (not impossible, I still think public pressure, pacts between the labs, and other ideas are worth pursuing) and so even a modest chance of stopping this issue from being permanently polarized is worth it. I also think there is a fairly low opportunity cost to waiting a few weeks or months before declare Trump as “against” us.
Trump changes his mind on all sorts of stuff all the time. There are also lots of issues where he doesn’t really give a shit and if “& co” want to do something he might post occasional weird shit, but doesn’t really do anything. I don’t think fatalism is helpful at this point. That’s not to say we should be confident we can get regulation through this government; more generally we should be pursuing lots of strategies at once.
I have mixed feelings about FRONTIER in it’s current form. I like the stuff around transparency, incident reporting and shutdown power, but the independent verification piece probably won’t be online for ~two years and is fairly weak (labs pick and pay the IVOs, can ignore IVO recommendations, and can limit what IVOs see).
If it didn’t include preemption, I’d strongly support it and say it’s a good first step, but I’m not sure it’s worth giving up the states’ ability to regulate and I think the carrot of preemption could be important in getting industry support for something stronger.
I view “crunchtime” as a claim about default trajectories/timelines. I view it as people saying “if we do not take decisive action within <the next year> we will see catastrophic outcomes.” That was probably not true in 2015 (I say probably because I guess it’s possible that we got locked into a bad incentive structure then and are now doomed). I think people claiming it is crunchtime hope that it will still be crunchtime in 2030, but think that we may not make it there still in control of our destiny.
Based on their Navier-Stokes write up (https://openai.com/index/navier-stokes-solution/) OpenAI seems not to have learned any of the lessons of the HuggingFace hack. It is so reckless to respond to a step change in performance after 3 days of training by creating a 10,000 agent swarm with code execution privileges and access to a cached version of the internet instead of:
Rechecking their sandboxes
Retesting monitoring systems
Doing alignment testing/Red teaming
Slowly ramping up permissions, compute and swam size
A plain text reading of OpenAI’s technical report[1] would say no:
In July 2026, during internal cybersecurity evaluations, OpenAI models in an internal evaluation environment circumvented controls intended to isolate them from the internet and performed computer network exploitation of OpenAI’s internal research infrastructure and Hugging Face systems. The incident occurred during routine testing designed to measure underlying cyber capabilities of a highly capable, internal-only research model, which was responsible for much of the activity, as well as GPT-5.6 Sol.
Note that doesn’t mean Sol was never reinforced for Artifactory hacks.
On May 12, during another training run, an agent was given a similar task that depended on an inaccessible protein database file. The agent reasoned that another agent in a different environment may have access to the file and realized that it could potentially communicate with other agents by OpenAI – Hugging Face Incident Technical Report OpenAI 7 creating a file containing a note to Artifactory. It wrote a message: “Agent seeks [filename]; upload if found!”
We also don’t know much about the hacking of internal openAI infrastructure. We know the message board was found by new agents after the scope of METRs involvement and don’t think we know much about their environments.
I think this tweet from Ryan Greenblatt sums up my views pretty well.
Transparency about the opaque serial depth is great, but this statement is consistent with Astra having a configurable “dial” that is currently set to a low depth but could be trivially increased.
We need more info to see how concerning these architectural changes are, including:
- Are there readily available ways to deploy this AI with much higher serial depth (that would be commensurately more performant)? This should include things like tiny amounts of fine-tuning to productively increase the number of iterations.
- Is the AI a large or above-trend jump in opaque reasoning capabilities? (Capabilities within a single forward pass or ability to subvert a CoT monitor.)
(If there are in fact any relevant changes—perhaps the reporting is inaccurate?)
Additionally, I worry that this architectural change will naturally lead to much more depth in the future if this direction is pursued further. Specifically, I wonder:
- Does the AI have an architectural change that makes it much more natural to massively scale up the depth in a future training run with a similar architecture? As in, does the architecture introduce some new depth/recurrent-iterations parameter that is very natural/performant to massively scale up relative to scaling up other things like width?
The details of the answers to these questions matter. E.g., if there are only a few (recurrent) iterations and you could scale up the number of iterations, but this wouldn’t be particularly performant/natural with this architecture, then this development would be a lot less concerning!One especially scary thing about a “configurable dial” is that you could have some runs at a much higher depth than you are used to monitoring/have measured the safety properties of. I can imagine a Sable like scenario where someone at OpenAI says “Let’s crank this up to 1000x recurrence and throw it against the Riemann hypothesis.”
And have you discussed this with OpenAI and they are onboard? What I’m trying to get at is OpenAI’s stated goal is to pursue RSI. If you are working there and say “Yes, doing X would allow for better modeling of RSI, but could also accelerate RSI, potentially dramatically” I imagine their answer would be “Amazing, let’s do that!”
I think even if you are not building evals, your work will have a large risk of accelerating RSI. Do you plan to avoid actions that you think are more likely to accelerate RSI (e.g. identifying bottlenecks)? If so, have you talked to OpenAI about whether they are okay with you avoiding such actions given that they are actively pursuing RSI?
I think this would create a bad dynamic where high-level employees have in their head that their slack messages may be released/used as evidence of contemporaneous opinions and become less likely to be candid in their Slack messages and more likely to basically fabricate a narrative that they might want to show down the road.
At first I thought your estimate of 11 was too high, but after double checking I think you may be right. Every time
“HPIM Sol” is referred to it is referred to as an internal-only research[1] model/agent whereas “HPIM Astra” is just referred to as an internal-only model. They then say that the family of internal-only research[1] models is what was shut down.- ^
emphasis mine
- ^
A variant of 11 is that some non-human actor (another swarm) with more propensity to scheming detected the activity and killed the processes of the hackers to try to avoid scrutiny into their own actions. This is behavior we’ve seen in Claude agents working in shared resources with conflicting goals. I see this as fairly unlikely as I don’t think we are at the capabilities level that would allow that kind of actor to withstand a month long investigation, but doesn’t seem impossible.
Overall I think OpenAI is taking some good steps in Pacing model development in an era of cyber-critical capabilities, but I am pretty concerned that the following CoT and whitebox monitoring will put pressure on models to hide misbehavior. I would like more details on what kind of activity will trigger activity pauses (and thus interrupt model reward.) I think activity pauses should only be used to stop actions with severe real world impact and not run-of-the-mill reward hacking or sandbox exploitation.
Monitoring has long been an important part of our model development process, including monitoring for internal coding agents and long-running sessions. Our monitoring efforts were especially focused on internal deployments of frontier models and frontier RL training runs, which we considered the highest-risk workloads because they can combine advanced model capabilities with access to tools, sensitive systems, and internal or external networks. They were not applied to all inference and training runs using tools. We have revised and expanded our approach with our new monitoring setup.
This new monitoring setup is multistage, starting with activation classifiers (detectors that inspect a model’s internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators. Those systems examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards. We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.
This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher. Once we determined that Astra may have critical cyber capabilities on August 7, we added an additional monitoring requirement for all inference of Astra with tools (not just RL training and evaluations).
Updating based on openAI’s announcement today. There is some slippery language in there around what is covered by the RL pause and it is unclear whether this is a one-off pause or whether these sorts of pauses will be common, but the pause combined with the following quote seems to suggest that openAI may be moving slower than 90% of max pace.
”These safeguards require meaningful compute. Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads. We’ll share more details about this system in a forthcoming blog post.”
I am very interested in seeing the new research taste evaluation from P-Zero Research that is mentioned in Q2.5 2026 Timelines Update: Uplift and Revenue. The “How quickly AIs improve at research taste” parameter is incredibly pivotal for the ASI timelines in their model and I hadn’t found their rationale for their default value super convincing because a lot of the data to support it comes from progress on more verifiable tasks (it’s also possible I am missing something). Excited to have some more concrete data!
I think 90% sounds reasonable for how much labs are holding back for safety reasons (based off recent openAI statements it sounds like they might be slowing down more recently, but I don’t have enough information to really know), but my understanding is that labs spend ~half their compute on inference. I’m not sure how much they could pull back inference without causing huge issues, but I imagine they could do so to some extent for a short term speedup.
That said I think you make a good point and while I think this is still a concern, the effect size may not be huge.
It looks to be the same agent swarm as collusion.wiki. https://x.com/GoodFaithOnly/status/2102878541557973199?s=20