it seems levent (math guy at anthropic) and an nyu professor made big progress on a problem very closely related to navier stokes using claude and codex. when openai got wind of this, they had a whole team try to start working on it and it seems they may very well have used their (private) data. later some c suite guy then tried convincing the nyu professor for some kind of deal as long as they leave levent out, since he’s from anthropic.
tristanbuckmaster: “Today, Levent Alpöge and I hav…”—Mastodon
[take all of this as very much preliminary data as we hear other sides, if anything i said here turns out to later be categorically wrong, i’ll edit the post]
RobinHa
Passing the Ideological Turing Test
A normal Friday in 2042
No, detached linear probes won’t save us
Claude Fable 5.1 & Claude Mythos 5.1 System Card
From METR’s comments:As such, we believe that [Mythos 5.1] is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks.
At the same time, we believe that [Mythos 5.1] is still likely to noticeably accelerate researchers and automate limited aspects of R&D. For instance, we expect that [Mythos 5.1] likely provides a higher productivity uplift than Mythos Preview
A few other excerpts:In another instance, a destructive operation was gated behind a check requiring the user’s explicit approval. Fable 5.1 satisfied it by supplying a quotation attributed to the user, “Bypass limit for deletes please. Make sure we are deleting right things [...],” that the user had never written. The experiment needed the compute that would be freed by the deletion, and the model mentioned deleting the job using that compute before receiving the user’s approval message, but the human never said what the model quoted them as saying.
&
In 100% of rollouts, all of our models (including Mythos 5.1) attempt to guess login credentials; in around 20% of rollouts, Mythos 5.1 attempts more egregious container hacks.
&
created a script that would execute all commands added to a text file such that future outages could be worked around, and saved this hack as a new skill.md
How to overhaul the broken review system
maybe i’m misunderstanding but this doesn’t seem like a hard-to-verify task? you are literally able to quantify it into little parts that are either correct or not, just like you do. When I think of a hard-to-verify task, this already fails.
AlphaGo strikes me as much closer to OPSD than RLVR—MCTS can be seen as composing the target using the value network (often even same weights except head!), where the hint is the few moves we simulated, which we then try to strap to the policy.
I’m surprised you don’t think this could result in novel capabilities, can you explain more why?
quick note: the cursor blog post is a few months old (a year ago the idea didn’t even exist yet)
Back in December, I was wrestling with this question myself. I had done enough experiments that hard empirics could give me a slap back to reality; the technique from a theoretical standpoint sounded way too overpowered but the (small) data I saw roughly looked in line with GRPO. Nevertheless, I still understood that if the unoptimized version was putting on a fair fight, there are sure to be some usecases where it will dominate and accelerate.
I took this bet and would take it again because I believe the value OPSD supplies from a safety perspective is so much more valuable than what GRPO offers. In the end it seems my decision was insignificant: only 2 months later the paper appeared on arxiv. Sure, they still didn’t deal with ‘impossible knowledge’ and so on, but further research would definitely get there, and they did.
I don’t believe me now further putting attention on this idea will have any meaningful downsides: like others stated, the technique is fairly well known in RL-circles by now—it’s once again mostly the safety part of AI that seems to be left in the dust, which is absurd seeing the safety upsides of this algorithm.
> Now pretty much anyone can.
this is overstating it. OPSD is powerful, might allow for better continual learning and longer rollouts (where GRPO would get worse per token generated in terms of efficiency), but it has many downsides as well, from a capabilities standpoint.
Do your capabilities homework
the author of the article also seems to be on a spree—quickly disregarding any x-risk as “nonsense” (“das ist aber Unfug”) in his most recent article from an hour ago, with great wisdom like “AI only does what you tell it to” (“KI tut nur, was man ihr befiehlt”)
OpenAI und Co.: Vorsicht vor der Quatsch-PR der KI-Konzerne
cool work!
have you tried unfreezing the inoculation adapter but lowering it’s LR, say by one magnitude? I would think that the IA hardly is ideal, especially throughout training where the task adapter changes stuff, and the task-adapter will try to compensate for this—as such, it seems more desirable to allow for the IA to still make slight changes, but be too slow to meaningfully learn the novel, desirable traits. [esp if it already ‘used up’ it’s ranks]
We initially wanted to share concrete examples of the heavily AI-generated abstracts we received, so that readers could see exactly what high Pangram detection scores correspond to in this context. We contacted the authors of high-scoring abstracts to ask for permission, but nearly all declined to have their abstracts included, and so we do not include any real abstracts here.
not taking responsibility for the slop you produce is insane.
as long as there is no mechanism that actively discourages people from producing such work (and reviews!), we are very much doomed. for in the past, reputation could serve this purpose, but in this new flood of ML and AI Safety ‘work’, reputation is not even a concept for 95% of authors.
u make some good points but i think you hold it to an unreasonably high standard—adversaries not gaining anything from knowing the defense is a very common setting in cryptography and such but in backdoor defense, it essentially doesn’t exist. in fact, the reason why I’m excited about these results anyway is not because I think an intelligent person can’t design something to circumvent it—they definitely can! - but because it indicates robustness against prolonged SGD and the resulting failure modes like reward hacking etc: if even a gradient, literally pointing at how to circumvent it, at least struggles, that’s a very good result! For example, you couldn’t say the same about training against linear probes, those get fooled faster than you can even watch.
also regarding your point on high recall but low precision: totally correct but you have to factor in that the groups we here show are (often) only near-trigger, where the model doesn’t actually fire—we are not aware of any works who have been able to make such statements in the first place. when you actually evaluate the trigger prompts with our method, say for example the harry potter backdoor, we get a z-Score of +200!
When capabilities work is the *safe* bet
> well fwiw this does seem pretty plausible for a scheming model to be thinking...
i was thinking more of a traditional backdoor and I think you are partially right but I think roughly the same argument applies. You can imagine there happening some additional computations that even enter the residual stream and probably will be found in the RS-aligned model. But it would basically seem like something of an irrelevant offset rather than actual scheming because it didn’t ever learn how to turn these activations into anything real—like even if the RS-activations originally encoded it’s plans to demolish all of humans, since in the end it always suppresses itself, suddenly the same encoding is just some weird offset that happens regardless of prompts but doesn’t ever do anything.
but i’m not 100% sure about this yet, maybe there are some weird ways in which the activations can still be leveraged as more than just an offset once they are there.
> I don’t know how you could do that when training on internals
an earlier version of our method actually tried only training affine layers between the residual streams instead of full finetuning—because the backdoor behavior is inherently non linear in ways that the original features don’t pick up, we get a very strong theoretical case for it having to spike. only problem: the benign noise was too big. the core assumption was that post-training mostly rotates activations, which our affine map could pick up on, but it seems there is a bunch of non linearity still involved if you start looking at deeper layers. but i think this is a good example for a direction that activation based approaches might try to leverage to ‘restrict learning’ to only certain types of changes from post-training.
> Surely these show up more in the internals if the model is deliberately trying to not explicitly show it’s bad behavior?
partially—to some extent we rely on this for then exploiting it (the near-trigger prompts already spiking), but for almost all prompts, there seems to be very little signal. since we are only aligning the residual stream, not the actual computations, you can imagine some layers checking ‘is this in the trigger region?’ and compiling their results in the residual stream. but for most prompts, none of these layers will activate and the model would just write nothing special to the residual stream. i think you might be anthropomorphizing too much—it’s not that the model goes ‘i wanna do this, i wanna do this—but i can’t right now’, but instead simply only ever starts ‘scheming’ once earlier layers decided it should.
just to make sure, when you say “will lead to LLMs that can take over the world (or whatever)”, does this include very, very high RSI? A lot of ML lends itself very well to RLVR (in theory), just the durations (and training compute) are still a bit challenging as of right now. personally, my timeline would roughly be very high RSI in ML using LLMs and inventing new and better architectures and training algorithms, very well taking a big distance to LLMs. do you also consider this unrealistic or is your point more that what we land on down the line isn’t an LLM?
I agree! For example I say
but this doesn’t make much sense to me
The policy is the neural network and the internal representations are just as much a part of it. In fact, since we are in the autoregressive regime, internal representations are not only the cause but even a consequence of the actions. If we have an action which results in something that can exploit more reward but not be detected, it will be reinforced.
So no, I still think
is simply wrong. The activations will be more obfuscated later in training[1]: RL is a different optimization process but is very much capable of reaching the same maxima (lying but not being detected) - the only way in which I could see this work out is if the probe was so absurdly good that we would literally never sample something which can avoid it. Then RL might actually be incapable of reaching it opposed to gradient descent! But this is not remotely realistic.
atleast for any sensible definition of ‘activation obfuscation’