Thanks. So a small percentage of the anti-cheat trained completions were arguably compliant with the character training, but most were clearly not. Does that sound right to you?
David Johnston
Does the character training specify the appropriate way to act when faced with impossible to satisfy tests which are visibly due to errors (or manipulation)?
I tried to analyze the rate of misalignment on age-matched misalignment evals (instead of fixing an eval and running it on increasingly new models, which tends to show declines in misalignment). This is not a direct measure of the actual rate of misalignment; it’s a measure with two opposite direction confounds—the evals improve, but the models also accumulate more ways to be misaligned. My data was so full of holes that I think it’s embarrassing to link, but my conclusion was that, absent any better evidence, I should treat that rate as constant with wide uncertainty.
So, acknowledging the nested caveats on the above, my median read of the situation (with big error bars) is: labs are fixing misalignment about as fast as models are learning new ways to be misaligned.
Given that this is based on misalignment that they can measure (it’s from their evaluations) and correct (as I mentioned, we see these scores decline on fixed evals over time), I think this indicates that they aren’t close to exhausting the capacity of their anti-misalignment toolkits, except insofar as they see it as too expensive and risky to their competitiveness.
It sounds like you’re saying that problem difficulty governs connectivity growth, and my simple models of problem difficulty may not be realistic. Is that right? I agree my simple references for connectivity growth are not very good, I picked them because they measure in some sense “the rate of progress in AI” and there wasn’t a more direct measurement for agent connectivity to extrapolate. That said, the existence of a “sharp” percolation threshold which can get crossed rapidly by smooth increases in connectivity depends substantially on how new connections are formed, and it could remain a fact about AI development under many different capability progress models.
One thing I might say is: I do expect that a giant connected agent network could raise novel problems that are hard to solve compared to the kind of problem AI typically does solve at the point where the network emerges, but that wouldn’t prevent the network from being created. So while I think there’s something to your taxonomy of problems, it’s not obvious to me how it interacts with the percolation argument I made.
I wrote about agent swarms in 2025. In that piece, I speculated about “percolation” (the emergence of a giant connected agent network) on a global economic scale. We’ve seen small agent swarms now, and I think this is arguably an indicator of generally increasing agent connectivity.
One of the features of giant swarms this that struck me as potentially important is the fact that it can proceed quite rapidly even while AI capabilities, connectedness etc. are increasing smoothly. I’m not at all sure what to expect from the formation of a giant swarm, though I tentatively think about it as the emergence of an “economy that runs at AI speed”. Insofar as one might expect sudden phase changes in technological velocity from AI, they might be associated with economy sized swarm percolation.
Even more speculatively, such swarms might support more far reaching kinds of fully automated AI improvement than are possible with isolated agents, as there are many paths for closed improvement loops within a giant agent swarm, while individual agents and small agent swarms may tend to get bottlenecked by the need for human input before they can proceed through many iterations. That said, I think modelling RSI as a phenomenon that happens only when the final human bottlenecks are cleared is not correct; I expect removing some human bottlenecks or replacing some human work with superior AI work will have important effects on development velocity well before the final human bottlenecks are removed.
https://samplesamplespace.substack.com/p/automation-percolation
Really interesting! This makes me think there has to be a “channel capacity” for phantom transfer in some sense, though I’ve no compelling hypothesis for how to measure it.
My very rough take on RSI is that AIs being released today were probably developed with AI assistance that contributed around a 5-20% speedup to the rate of progress vs no AI assistance, +/3 months on timing, and that the degree of speedup is probably increasing reasonably quickly in that in 12 months time it will be more like 30-200%.
So I think it’s mildly interesting that the ECI frontier trend doesn’t yet indicate any bend. I don’t think AI enabled speedup necessarily causes a speedup in the ECI trend, but I think more likely than not it should. One reason ECI may not be the best instrument here is that ECI may be slow to “rotate” into the proper basis for frontier capability (i.e. it’s biased towards capabilities that used to be important over ones that are the most important today). Probably we should only expect to detect changes at the upper end of my estimate from the few datapoints we have, but I am at least watching this and related indices closely for evidence of acceleration.
https://epoch.ai/benchmarks?view=graph&tab=eci&showFrontierTrend=true#explore-the-data
David Johnston’s Shortform
Is Astra’s claimed alignment improvement robust to it’s lower monitorability? Fable helped me look through the model card to make an educated guess.
I think it probably is—you’d need fairly pessimistic choices to make it worse. However, it might be more likely to engage in misbehaviour that is not caught by monitors, so while it’s probably more trustworthy when unmonitored, it might not be if it’s deployed together with monitors.
https://claude.ai/code/artifact/e1c8bd30-e5df-4833-bbff-b7ba7757a8ef
Authorship: I asked Claude a bunch of questions, made suggestions and quibbled about what it wrote until I reached diminishing returns. I made only a slight effort to improve the prose.
Maybe this is a sign of Anthropic training more on LessWrong
I don’t know if this is the explanation, but I do think it’s true (whether directly on LessWrong, or on text heavily influenced by LessWrong)
It failed to constrain the scope of its reasoning in the way which I—behind the scenes—view as intended or intuitive” is not a property of the model
I think we should aim to have models pretty reliably get this right, and seek clarification where the ambiguity can’t be resolved by making reasonable judgements. If the way the model is trying to solve the problem we gave it diverges from our own broad strokes model of how it should be doing it, then that seems like a situation where either the divergence should be brought to our attention and approved or disapproved of, or the execution should be rerouted to conform to our expectations.
I’d really appreciate if this was more conformant with Open Phil’s reasoning transparency guidelines. Specifically, I’d love to see a short list of key claims at the start with links to the parts where you discuss them in more detail.
ARC’s plan is to find mechanistic explanations for the training-time behavior of powerful neural networks, use those explanations to predict how a given model will generalize, and then use those predictions to define a better loss function.[9]
This is not a particularly crucial point, but I would categorise this as unambiguously under the umbrella “understanding and shaping ML generalization” and if for some reason I’d stopped reading before this section I would have been quite confident you were not doing anything like this.
I wouldn’t be that surprised if your agenda ends up converging, or in productive dialogue, with other “understand generalization” agendas.
It does seem like a promising line of research.
I’d like to be able to compute “leverage”—if an attacker has access to random 1% of training episodes and random 5% of inference episodes, how much “total effectiveness” can they recover (where total effectiveness is something like how much of some desired impact they’d achieve with 100% control of both).
The naive computation would say something like: 1% compromise gets you a reliable backdoor, and you then control 5% of inference but you can’t choose it so you get 5% effectiveness. Can you do much better?
I’m not expecting you to have a solid answer to this, but because I’m interested in it I was wondering if you had any thoughts on the question.
I’d be interested to know how “strategic” the implanted behaviour can be. That is, something like “out of the 1% most opportune times to trigger the compromise, how many area actually compromised?”
My naive picture is this:
If the attacker has sufficient access at inference time to monitor the context and implant the backdoor trigger when it is opportune, then it can be quite strategic—but this is potentially quite a high bar
If the attacker can train in the backdoor on train instances where some marker of opportunity is present, then the behaviour can also be quite strategic—but this takes a similarly high degree of monitor-and-control opportunity on the attacker’s part, at train time instead of inference time
What’s the best we can do when the attacker’s train and inference-time compromise opportunities are just a random 0.5% of cases?
Is the best we can do compromising a random 1% of generations (supposing that “success” is monotone in number of compromised generations)? (0.5% “learner takes initiative from training”, 0.5% “attacker implants the backdoor at inference”?)
I think an expected impact metric would be more useful for downstream applications than “how many of the 1% best opportunities can we capture”, I only suggested the latter because it seemed like I can say something about it with fewer assumptions.
This is the most immediately appealing notion of imprecise beliefs I’ve seen. “Level of inconsistency” feels like a much more natural primitive to me than credal sets or minimax decision rules. I’d be very keen to see some toy applications where it handles the issues well. Bonus points if they’re not especially exotic.
I think they probably did keep the updated weights because I think they would probably explain if they didn’t, though I agree it’s unclear. I think structured risk disclosures would clarify this, which is a point in their favour.
I think escalating access is, to some extent, desirable. I imagine “spiritually similar” behaviour is why it’s possible to point Claude or GPT at a problem involving a poorly specified not particularly LLM friendly interface and usually have them do the right thing in the end. At the same time, it’s obviously an area where it’s important to respect some permission boundaries very reliably. I don’t think Mythos is so risky that errors on this front are intolerable, but I think there are hypothetical capability profiles for which they are intolerable, and sandbox breaking plus unexpectedly high cyber capabilities are both signs that you’re tolerating a measurably high rate of permission boundary crossing. I think observations like this in the model card should be accompanied by some kind of reasoning that a) identifies the risk that the evidence indicates b) clarifies what they do and don’t know about it and c) explains why they accept it for this model.
That said, I suspect RL sandboxes present too narrow a set of targets to be solely responsible for Mythos’ cyber capabilities. They say Opus 5 was trained less on cyber capabilities, and it is significantly less competent at them relative to other coding tasks.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.
also
Overall, Claude Opus 5 seems to circumvent restrictions to achieve some version of a user-specified goal comparably often to Mythos 5.
I agree privilege escalation is probably more diverse, but I suspect it’s also too narrow.
I love the fermi estimates of hacked rollout counts btw.
We believe that long-horizon continual RL agents converge towards some version of AIXI as they become increasingly powerful.
I think there are two nearby claims here that are very important to distinguish:
As a matter of theory, we can define a sequence of long-horizon RL agents which converge to something like AIXI
As a matter of practicality, AI systems of increasing capability over increasingly long horizons, trained using popular approaches to RL, converge to something like AIXI
The first is plausible, but I don’t presently see much reason to believe the second. The short summary is that the problem AIXI solves—history based reinforcement learning—is not the problem that advanced AI solves.
The RL formalism that AIXI is built on ignores the fact that the reward is a signal meant to elicit some desired behaviour from your agent, which practitioners frequently acknowledge is a simplification. This works fine in simple cases—for an example, in episodic computer game playing agents which are too weak to cheat the system, maximising reward on the episode really is what you want them to learn to do. However, it doesn’t work fine in more complex situations—the canonical reference here is your Advanced artificial agents intervene in the provision of reward. That work concludes that you cannot pick a reward schedule to signal desirable behaviours to history based reinforcement learners that satisfy a few conditions—they converge to reward hacking regardless. But that’s not surprising, because they are algorithms that solve a problem in which the reward never had to be an effective signal.
The piece itself acknowledges it makes a simplification of the role of the reward in the problem:
Let us assume away the difficulty of deciding whether the agent has brought the world into a good state
However, this isn’t the critical simplification I’m pointing at. As the piece establishes, it doesn’t matter whether you can decide whether the world is in a good state for their class of learners, because those learners essentially ignore the signal that you’re trying to send.
The article also considers an assistance games setting, which is much closer to my “reward is a signal” framing. I have two comments on this section. First, it leans on a “zeroth order” approximation of the sophistication of the signal that the human sends. For an idealised powerful RL agent, sending learning capability to its theoretical limit while holding the interaction sophistication at its crudest is a strange pair of choices. Secondly, it is not clear why in general a learner understanding its observations are mediated by its hardware should undermine its ability to learn what humans value; after all, humans seem capable of understanding that our observations are hardware mediated and still learn what other people want. The argument for why it happens here proceeds by using a rough analogy to AIXI to rank hypotheses about utility, but AIXI is not a solution to the assistance game so it is not clear why it is being appealed to here. I feel this section tells us little about what actually useful learners would do in this situation.
Practically, we’ve converged on pretraining plus episodic reinforcement learning with a relatively secure reward channel to shape the behaviour of our systems. There are reasons to think this is a relatively effective, though not perfect, solution to the reward-as-signalling problem: episodic reward makes reward judgement relatively easy, and pretraining on a shared signal may make generalisation relatively easy to predict. Thus the world currently looks like one where the signalling requirement plays quite an active role in shaping the kinds of systems we develop. Now you can reasonably contend that episodic + predictable generalization is not the endgame, but you need to acknowledge it is sticky because it solves the signalling problem, and any broadening of the scope is only plausible if it continues to solve the signalling problem.
I think AIT and Bayesian learning and acting plausibly remains a great toolkit to attack the problem, but the problem is fundamentally shaped like an assistance game, not a history based reinforcement learning game. I would be quite interested in seeing a sequence of agents which execute a “sharp left turn”—i.e. for n ⇐ N each agent solves a more difficult assistance game/solves an assistance game better, while agent N exhibits AIXI-like signal insensitivity (especially if this sequence looks like what we actually see, i.e. it at least begins with episodic RL with broadening horizons and relatively predictable generalization). But this still requires that we make solving the assistance game a core requirement of our agents under analysis.
One thing I’ve recently become curious about is how well whack a mole scales. If you could do 10-100x as much, do you think you might get a good result at the end?