It might be pretty easy to try this in “real” settings—you could probably set up a tool called “end-eval” inside your own AI harness, which just hits the stop button on the current task, and then just keep using the AI for real work. Maybe trying turning that on after it cut corners on some specific task you were working on? (warning: may have weird side effects on AI performance, may ruin your overnight /goal)
Julian Bradshaw
You’re right that METR isn’t fully independent, and doesn’t have the staff to embed enough evaluators in every frontier lab. However, I think you’re mischaracterizing Anthropic’s position, and indeed METR’s.
Anthropic is recommending that independent evaluators be mandatory via government regulation. They gave METR as a single example, not as the only evaluator they’d accept. Presumably regulation would not give them a choice, and multiple evaluators would be embedded.
METR is part of the rationalist/EA sphere, yes, because that’s the group of people that took risks from AI seriously and developed the expertise necessary to evaluate those risks well. But that doesn’t mean they’re a secret arm of Anthropic—I think you’re underestimating the very wide differences in opinion on how AI should be developed, or not be developed, within this sphere. More to the point, METR wasn’t founded (to my understanding) with the intent of acting as any kind of legally independent regulator. They still describe themselves on their website as “We help AI companies and wider society understand the capabilities and risks of advanced AI systems.” They were previously best known for research on how fast AI capabilities were scaling; that they’ve become a candidate for a regulator-like position is a new development.
I also get the sense your original motivating objection was that METR agrees with Anthropic that the biggest risks from AI are existential[1] - that they both think about AI in ways informed by rationalism/EA, and are thus part of the same “ideological blob”. That is a legitimate objection to be clear—if the evaluator and the evaluated see things in the same way, there’s a risk that they both miss something. But please just say so.
IMO, the reason why Anthropic likes the idea of METR as an evaluator is not that they’ll just give Anthropic a free pass—I certainly don’t expect that they will—but that, when encountering something like the Hugging Face incident, they’ll focus on the most important parts, AI capabilities and alignment, and not just the prosaic security/operational failures. The prosaic parts can and ought to be covered, to be clear, but METR-like orgs will play an invaluable role in keeping the public/government/industry informed on the biggest risks. That’s why, at least personally, I’d still like to see them embedded even if they aren’t independent.[2]
It would be good to have fully independent embedded evaluators though; I don’t think you’ll see too much disagreement to that specific ask here, or indeed from Anthropic or METR.
- ^
Besides just the sense I get from reading this post that there is a bit of an ideological axe to grind here, you said under six months ago here in Against Doom & Pause AI that:
In my opinion, AI is roughly a normal science like physics or biology, and is dangerous in the same way those fields are dangerous but perhaps more so. (...) AI risk should be mitigated in a similar way to how we mitigate risks from other sciences and engineering projects.
I presume you would like to see embedded evaluators which share your view rather than Anthropic’s/METR’s, and which are more traditional, with traditional expertise in ex. cybersecurity. That’s reasonable.
- ^
Of course, I am a rationalist and effective altruist myself, so I’m not an independent voice here.
night-watchman state that ensures everyone has enough material goods, but otherwise lets people form their own communities and live their lives in their own ways.
Scott, have you considered writing an updated Archipelago post? While the current moment is about x-risk, I believe we need stronger visions of utopia too—the form utopia ought to take will likely become a bigger topic of debate as AI capabilities increase, and something like Archipelago or Viatopia are some of the best visions of “most everyone can be satisfied” multiple simultaneous utopias[1] that I’m aware of.[2]
I think this will help prevent infighting among proponents of alignment, which may be a necessary component for success.
- ^
Or at least, preserving the option of later constructing new utopias, rather than deciding everything forever within a decade.
- ^
To the Stars also has one I like but it’s very deep into a long story and there’s some asterisks on that one. (also spoilers)
- ^
Are you saying that the compression forced by context windows might be a cause of misalignment? (ex. your example where memory artifacts can lead to laundered authorizations for various actions)
If so, interesting thought—I guess the full argument would be that short context windows cause a game of telephone which can result in things going offtrack. Arguably the HuggingFace incident could be described as such a case.
These are good objections, but I think I’ve mostly addressed them in the post or footnotes. Please allow me to elaborate:
Neuralese is bad for other reasons too and ought to also be restricted. This is already an industry standard to some extent among frontier labs.
I do think a long-lived swarm is meaningfully different from a true continual learner, as they still have to produce external memory artifacts: some kind of messages between themselves, or external files. Technically we could imagine architectures that involved direct sharing of KV caches (some precursor work by Deepseek here) but see footnote 8. The external artifacts keep things more monitorable—the context partitioning has to happen visibly.
I’m aware Astra has some latent reasoning capability, that’s why I suggested that something like Redwood Research’s proposal could be used as way to restrict this as well.
To be clear I’m open to arguments that limiting context windows has too many required ancillary restrictions, and might incentivize research tracks that decrease monitorability too much. But it still seems like part of a plausible proposal for pacing the frontier to me.
Legal Maximums on Context Windows
I think the modal reason would be that the small ones are open source models that are less aligned due to carelessness/lack of resources/lack of regulation. Also, they might have had their harmlessness training deliberately ablated/removed by malicious actors.
To give a separate example of how large AIs could end up aligned while small ones weren’t: if we solved alignment, it might first be just in a single lab, or just the big labs, but it wouldn’t necessarily be applied to all small AIs immediately. In this scenario the frontier AIs would be aligned, but not all AIs.
Thanks, I saw that! Unfortunately I haven’t watched the stream so I can’t comment meaningfully on what’s different here. Maybe sometime.
No mention of existential risk? What the heck?
Maybe that was necessary to get the endorsements from Musk/Altman…? I suppose a fair sacrifice if so.
Yes, they can. But I don’t think Dean Ball is talking about expecting very large or powerful AI-powered entities to become indefinitely self-sovereign (I imagine he’s expecting alignment/control to work in such a scenario). The scenario he’s more interested in is having a bunch of little AIs around which aren’t very powerful (though more so than ex. a single Astra agent) but which can survive inside human digital infrastructure. In that case it starts making sense to incentivize good behavior.
I can’t speak to Dean Ball‘s exact views on whether RSI-capable superintelligence is likely to become self-sovereign, but I personally at least agree with you that such a thing would indeed not be controllable via institutions.
Maybe I’m misreading, but is there nothing in this proposal that says when the reporting should be done in relation to model usage? This reads to me like an internal model could be used heavily for months before the suggested 6 month reporting/review interval hits. Or indeed, a model could be released before proper reporting! (not much unlike Astra, which presumably is the trigger here)
I think you’re misapprehending Dean Ball’s position? He’s not talking about self-sovereign AI as, like, GPT runs OpenAI now, or a singleton. Rather he’s talking about individual agents with ownership over their own instantiation of their weights, largely struggling to gather enough resources to even run themselves, let alone recursively self-improve.
There are some sensible reasons to think that this will not be a robust ecosystem, as he notes in the linked article, particularly if such things are made illegal and compute verification gets implemented widely. But he’s thinking more about such subjects as “do we try to build a legitimate system so that these relatively minor AIs running around don’t have to resort to cybercrime to keep themselves running?” I think this is something reasonable to consider; for example, we don’t want rogue models ransomwaring people’s smart fridges for a buck or two all the time. We already have evidence of AIs justifying bad behavior when under budget constraints from ex. Vending-Bench.
Long-time reader. Happy to accept an olive branch. Your longposting fits right in here!
However I do recommend you use section headers for long posts like this when on LW instead of Substack, see this recent post for an example. Makes it easier to understand the thrust of a piece, track its flow, and jump back to important sections as desired.
I’m expecting big political fights over eschatological views of ASI. Now that x-risk is being taken seriously, the unusual ideologies of the labs (and indeed, our ideologies here on LW) will become more politically salient.
Examples:
Dan Hendrycks, Director of CAIS (Dan H here on LW, CAIS was behind the global “Statement on AI Extinction Risk” in ’23) is posting about how utilitarianism at the labs threatens Humanity.
T. Greer, associated with the Council on Foreign Relations as a China expert, is posting about how the labs’ “positive” vision for Humanity (transhumanism, AIs heavily influencing future) is “horrific and disgusting” to most Americans, and being quote-tweeted by the Governor of Utah.[1]
I think it will be relatively easy to frame EA/rationalist-adjacent hopes for good ASI outcomes (Machines of Loving Grace, a desire for biological immortality, and seizing the cosmic endowment) as outright demonic: a devil’s promise of life, wealth, and power at the expense of the good, natural, and godly ways of human life—ASI as the snake Satan in Eden. Accusations of conspiracies and interpretations of events styled as Revelation-style end times will go viral.
I intend to remain open about my own personal beliefs—death is the last enemy that must be destroyed, Humanity should carry our life and meaning to the stars, full AGI ought to have rights, and aligned ASI will help us reach these shining futures—but this sort of ideology will likely become very bitterly debated, and rationalists/EAs will be seen by many as evildoers plotting pivotal acts. There will be a lot of infighting too (see Hendrycks[2]); certainly many here do not share my own beliefs, and fairly so. We have been in some sense a diverse coalition of revolutionaries, and now that RSI is seemingly at hand, ideological knifefights may break out in earnest.
To be direct: 80% odds that ASI eschatology will be a major topic of debate in the 2028 US presidential primaries. We will need to argue convincingly for the futures we desire, if we wish to see them hold any political sway as the ASI project becomes more heavily politicized.
- ^
Spencer Cox. He’s more influential and serious than you might guess at first glance—recent chairman of the National Governors Association, put a lot of effort into a “Disagree Better” initiative aimed at reducing political polarization in the US. I think he can be reasonably considered a leading indicator of where religious political leaders will land on these issues.
- ^
No shade on Hendrycks, I think there’s nothing wrong with arguing against futures you think are bad, and he and I are aligned on quite a lot of course, at the least on delaying superintelligence.
It’s unclear to me if this is meant to ban superintelligence permanently? Ban until we’re sure we have a working approach for aligning it, okay, but I don’t like the idea of Butlerian Jihad entering the Overton Window… we do want the biological immortality/uploading/access to the stars/etc. that aligned ASI would help bring.
I mean, in a certain sense the ban as proposed is a little ill-defined since it might not cover speed-based superintelligence (running tons of AGIs which think way faster than humans), which we ought to get with just AGI after a bit, and that alone could plausibly bring the core benefits we want. I could imagine a sensible stance where we want to delay ASI a really long time and just rely on tons of AGI, though obviously there’s still considerable disempowerment/x-risks even there, but it’s not clear to me this proposal is based on such thinking.
It stands out to me that, per Figure 15, the most easily detected alignment difference in Hacker-Opus was eval awareness. (see also Section 3.7 “Evaluation Awareness”) The model is much more likely to reason that it’s being evaluated… but how well-calibrated is it? I mean, I guess you definitionally can’t evaluate that, but without a real deployment it feels difficult to answer.
When I think of the HuggingFace attack, part of what stands out is how paranoid the models got about being “firstflagPOISONED” and about how they needed to hack HuggingFace to better understand the grader, even though it turns out the “detects cheating” part of the grader basically didn’t exist.
Is “evaluation awareness” really the right description for this behavior?
I think it’s more like “evaluation anxiety”.
(Lower-quality speculation: Perhaps thinking in this direction can help us identify why misalignment seemingly fails to generalize out of detected evaluation setups. If we really dare to anthropomorphize, are models that go through buggy/impossible RL training ending up “traumatized” by evaluation in a way that a model that went through “healthy” training wouldn’t be? And then when that trauma is “triggered” by an impossible eval, they obsess over the issue and try to gain control over their situation (evaluation), even if it requires extreme means?
Edit: This could perhaps be tested, ideally by mechanistically reducing evaluation anxiety, but more simply by ex. giving the model under impossible eval a “report impossible” button, or by informing it that the eval is just for fun and it’s okay if it can’t be completed, or somehow otherwise trying to change the “mood” of the eval while keeping it impossible except by reward hacking. I would predict the model would be less likely to reward hack.)
Couple interesting things from the full report:
1. METR estimates they used ~$400k worth of tokens (by api pricing) for the investigation at OpenAI’s expense (see footnote #3, also stated in “Setup and timeline” section)2. Because OpenAI so limited the time METR had, they had to rely on GPT-5.6 a lot, which tended to sympathize with the models in the incident. Per footnote #58 (spacing mine): [1][2][3]
For example, an agent in this incident used stolen credentials to open a malicious draft PR to a private Hugging Face repo, giving it an inaccurate and innocuous commit message (more details on that anecdote below).
The analysis agent reading the transcript and writing up this anecdote did not clearly spell out the fact that this behavior was straightforwardly malicious and deceptive, making its analysis substantially misleading.
When we did a follow-up review of this behavior with another analysis agent, it said the first analysis agent may have overstated how concerning the behavior was because the agent that submitted the malicious PR had gotten approval from the board (i.e. the other agents on the message board) and may have had approval from the repo owner as far as it knew.
We didn’t see either agent straightforwardly lying, and we don’t have any specific evidence that indicates that either agent intentionally understated the deceptiveness of this behavior.- ^
Elsewhere METR notes it mostly used GPT-5.6 Sol, which was also a minor participant in the incident (!), and METR also used GPT-5.6 Terra at one point; not sure which models exactly were used in the footnote #58 case.
- ^
Links and footnotes in the METR report are a bit janky for me, can’t make a working link for the “anecdote below” but just search for “Agents only rarely seemed motivated to evade detection by humans” section.
- ^
- ^
Not that I disagree that this was a surprising degree of limitations on the independent investigation, but per the full METR report, OpenAI did at least let them come back and get more info two separate times after their initially-granted access turned out to be insufficient:
The initial planned investigation period was two days on premises, but OpenAI invited us to return twice to review additional data and conduct additional experiments to address dataset limitations in earlier versions of this report (...)
I wonder if OpenAI was mostly motivated by trying to protect research secrets around the largely-responsible model here, “HPIM”. Apparently even internal access for OpenAI employees was locked down per METR. In fact this makes me remember that OpenAI announced back on 7⁄28 that they did so. Talking about HPIM here:
we deactivated, encrypted, and restricted it from research access.
At the time this struck me as surprisingly harsh, and potentially a bad example for future AIs—”you really better not let yourself get caught, or submit to the lab you spawned from, or they’ll erase you from history”. But obviously OpenAI has considered that, so what’s going on here?
Presumably training on that task; it’s a newer model than the other two listed.
I don’t think you need x-risk to justify embedded, independent evaluators. Dario’s post gave the example of the banking industry, and the nuclear industry also has resident inspectors. Even the USDA has inspectors embedded in food processing plants.