Member of Technical Staff at Transluce working on behavioral evaluations.
Email me at the email available on my website at timhua.me to reach me.
For more Tim content, you can follow me on Twitter.
Member of Technical Staff at Transluce working on behavioral evaluations.
Email me at the email available on my website at timhua.me to reach me.
For more Tim content, you can follow me on Twitter.
Claude Opus 5 seems to circumvent restrictions to achieve some version of a user-specified goal comparably often to Mythos 5
If Opus 5 was trained on a comparable fraction of environments that let’s it learn various cyber-relevant reward hacking strategies, then it would have still be trained on less data than Mythos on these.
(Also this statistic you cite is from their internal usage monitors, not from the training monitors. If anything, that’s weak evidence that the Opus did more hacking during training, since presumably they’ve gotten better at training “hack-like behaviors” out of models towards the end.)
But yeah, generally I think it’s hard to attribute the cyber capabilities of model to one part of training. I still feel like the hacking probably helped.
They’re almost certainly not doing it for Mythos? This result is from reviewing what had already happening in the training run. It’ll be way too expensive to patch everything and retrain the model.
However I don’t think this is likely to have been all that relevant to it having been good at cyber offense because I don’t think “being good at cyber offense” is actually something that needs explaining, given the background context of an agent that’s good at programming
Yeah I’m not super confident on this part. But idk, models’ ability transfer what they learned in one RL domain to another is weird. Like you would expect AIs to be so much more smarter than they actually are given how good they are at math. So I still feel pretty good about the specific prediction I made in the post (i.e. Mythos would be noticeably less capable at cyber out-of-the-box.)
Ah I wasn’t being super clear, but this does not go against what I was trying to say, which is that there wouldn’t be any like, unannounced and human-intended changes to the model. For example, Claude complained that, for the Anthropic provider:
No written guarantee that weights behind a dated ID are fixed.
I think whatever happened with openrouter and opus 4.6 falls under unintended inference deterioration. But I suppose that, from a user’s perspective, it doesn’t matter why there is performance deterioration.
Wow, that’s crazy! I independently replicated this here.
It seems like the Openrouter Opus 4.6 is just kind of cooked? On this prompt, it would return glitch tokens in its response like:, “cliff\n\nLet me try again.\n\n67”, “връщане\n\n67″, or “pigeons\n\nWait, let me reconsider.\n\n42\n\nHmm, let me try again.\n\n73”
Funny enough, this is an issue if you use the Anthropic provider through openrouter, but not if you call the Anthropic API.
In the original NLA post, the authors say that “this experiment is sensitive to specific sampling parameters. Specifically, it appears non-replicable on the public API.” I wonder what’s going on with how this model is being served? Anyways, I also tried replicating the result where increasing the stated reward increases the reward seeking rate. I find a non-monotonic relationship in the Openrouter Claudes.
(also note the big gap between the Boolean reward and the numeric reward).
I think there’s a difference between an error (e.g., some llm judge you used having a really poor recall rate or something) and taking poor Claude-conducted analysis as given (e.g., 97% of AI safety research use Openrouter unsafely, which is reported above the epistemic note. Ditto the “all the providers have gap” section.)
[I’ve also expanded on my thinking about this more in the original comment’s edit.]
Edit 3: I’ve removed my downvote after Matthew made various changes to the post, which I think is now much better.
Strong downvoted for slop research, even though I agree with the takeaway in the title.
In the executive summary, a core claim is:
A review of influential AI Safety research codebases that use OpenRouter for their reported results found that 31⁄32 (97%) of them use OpenRouter unsafely.[2]
And this “review” just redirect us to a Claude artifact, which classified “MathArena” as an AI safety research codebase. Come on. Does anybody even read posts anymore.
No Guarantees from Fixing a Provider
Although selecting a provider cuts down on a major source of variation, it does not guarantee that each time you request the provider for a model, you get the same model.
I had Claude look through a large number of model providers and figure out the extent to which they report what they serve, how it has changed over time, and their notice policy for changes in the future. All 10 audited had gaps. Fixing a provider is still the best practice available when using 3rd party providers, but unfortunately it falls short of ideal.
This Claude audit is also slop? I’m willing to bet $100 to $1 that, modulo unexpected issues (e.g., (1) (2)), Anthropic/OpenAI/Amazon Bedrock will not deteriorate the quality of a model or switch it out. All three of these providers were rated as a C by Claude in your audit! It rates Anthropic as a C because there is “No written guarantee that weights behind a dated ID are fixed.” Come on. Also, when people are using openrouter, they’re typically using open weight models, so it doesn’t even make sense to rate “Anthropic” as a provider.
----
Edit: I thought everyone already does this. You’d want to do pin the provider just for token speed purposes already.
But I am surprised by how many people didn’t already know this, so my guess is that this post will do a lot of good.
Idk. I still feel like I should downvote it. We should have a high bar.
----
Edit 2: Some more clarifications of my thinking behind making this comment and bring attention to it on Twitter:
If this post contained the title, a single paragraph of the issue with Arun’s paper, and a single bullet point on how the Inspect package does not force you to pin an openrouter provider, then I would have strong upvoted it. However as it stands, the fourth bullet point in the abstract/introduction contains a number that is ~totally made up by Claude. As I skimmed the post, I immediately went to “No Guarantees from Fixing a Provider.” And I was like, “oh my! I just fix my provider, what could I have been doing wrong this whole time??” Instead I have to sift through this Claude coded website which declared that the Anthropic provider scores a “C” because there is “no written guarantee that weights behind a dated ID are fixed” (along with some other hard-to-understand-claudeslop):
https://platform.claude.com/docs/en/about-claude/model-deprecations · discloses on: C what-happens-next
Anthropic notifies customers with active deployments for models with upcoming retirements, providing at least 60 days’ notice before model retirement for publicly released models.
A complete dated deprecation history back to 2024 with recommended replacements, plus an explicit Active/Legacy/Deprecated/Retired taxonomy and a commitment to long-term preservation of retired model weights. Best-in-class on lifecycle.
Gap: No written guarantee that weights behind a dated ID are fixed—it is a convention, not a promise. Never states serving precision.
This, like many parts of this post, is locally invalid, and as Eliezer said, Local Validity [is] a Key to Sanity and Civilization.
I care a lot lesswrong not being overwhelmed by AI-generated slop, and that the front page posts on lesswrong remain high-quality.[1] This post does not meet my personal quality bar. I feel a bit conflicted about downvoted because I was surprised by how few people know about the openrouter providers. Naively, this post would do a lot of good by informing everyone about this! However, I am generally skeptical of arguments of the form “oh this would do a lot of good, so therefore we should break a rule.”
I tweeted about this post because I want to bring attention to it. If others agree with my assessment, it would rapidly bring the karma count down. Many people visit the lesswrong front page at a given moment, so I prefer these corrections done quickly.
This is also not the first time I’ve publicly downvoted a post, and I expect that I will keep doing this.
For memorization, the setup I had was asking the model to recall as many details about a paper as possible when given only the title and author list. I think that’s probably a better way to measure it? The thing you care about is “how much of the paper does the model remember,” not if the model remembers the authors. You can even use a explicitly pre-knowledge cutoff model’s guesses as like, a baseline for how far you can get from knowing the title and author list and purely hallucinating.
Re: structure of claims in papers: the way I had set up my benchmark is to focus on getting models to follow up experiments that are more thing sort of “have to be run” in order to test a core claim (although I did more follow-up-ish stuff as well) (See also AblationsBench). For the emotions paper, I felt like there were many justifiable directions that the authors could’ve gone down, and thus it’s sort of hard to grade the AIs.
I think having models propose ablations to papers is a great way to measure research taste and will also have some work to share on this hopefully by the end of the month. However I find it hard to trust the soundness of results here given the glaring erroneous claim in Figure one:

Claude Haiku 4.5 was released in October 2025, yet Figure one claims that the model has memorized Alignment Pretraining, a paper released in December 2025. Similarly, Claude Sonnet 4.6′s knowledge cutoff (according to the system card) is May 2025, and Figure one also claims that it has memorized Alignment pretraining. I mean, you really should be suspicious when you’re claiming that Haiku has memorized something while Fable hasn’t?!
(For what it’s worth, I’m pretty sure none of these papers are memorized by any of the models,[1] so this error itself doesn’t directly invalidate the results. For the core results to be valid, I’m generally more worried about under elicitation, small sample sizes per cell (and a small number of total papers).)
Also,
Most research follows a clear structure: a headline finding backed by an opening experiment, and then the “supporting experiments” which are usually an exhaustive set of ablations and an exploratory bridge for future work.
I think this isn’t that true? Like the emotion concept paper doesn’t really read like that. In my own work, I’ve found it hard to get good candidate papers.
The papers are before the knowledge cutoff for most of these models with the exception of Fable and Alignment Pretraining. When I asked the model about the paper (in openrouter to avoid any customizations bleeding through), it was able to make some reasonable guesses, but those all read like guesses. If you instead ask Fable about my paper on eval awareness, it can remember a lot more details (such as the fact that our model organism writes type hints in eval but not in deployment.)
I vibe coded a Claude Code/Codex transcript viewer webapp
In the early days, when you interact with an LLM in a chat interface, you see almost everything the model sees: a set of user and model messages. The only thing hidden from the user in the UI is the system prompt. This is not the case with coding agents. The UI hides the details of tool calls, interleaved thinking traces, and subagent actions.
To get a better sense of what my Claudes are doing, I vibe coded this transcript viewer that shows the details of all tool calls and whatnot (it also works with Codex.)
One concrete issue this tool helps expose is the insanity of Claude code’s “WebFetch” tool. Claude reads websites using the WebFetch tool, which calls a smaller LLM and returns only a tiny summary of the site. This drives me fucking insane. If you ask Claude to review a lesswrong post, it would get a tiny summary, and then use that to do the “review.” Bad Claude!
(For Codex, the web search results are hidden from the user, so I don’t even know how bad it is.)
To combat this, I also have read-lesswrong skill you can use here. The skill forces the model to read the whole post, top comments, and figures before responding. We can see in the transcript viewer that the model followed the skill instructions. Very nice.
There are also some other nice quality of life features. The viewer renders LaTeX math, shows you what subagents are doing, makes it easy to jump between user messages on the side, and has a basic search bar to look through past sessions.
You can download the viewer here: https://github.com/tim-hua-01/cc_transcript_viewer.
Make sure to prevent transcript deletion in your Claude settings via
{
“cleanupPeriodDays”: 99999,
“showThinkingSummaries”: true
}
And add this:
model_reasoning_summary = “detailed”
To ~/.codex/config.toml:
I will say though, if you don’t know me personally, you should probably have a coding agent check the code for malware (or just give it the prompt request here)You should not be executing arbitrary code you saw on lesswrong lmao.
It’s probably privileging something like the arXiv prior way more
I don’t have any. I’ve also just replicated this in OpenRouter.
First attempt
Here’s my best-guess profile of Tim Hua, with confidence levels attached. Treat all of this as educated guessing, not fact:
My strongest hunch (maybe 30-40% confidence):
Works in AI safety/alignment research
Has written posts on LessWrong and/or the EA Forum
Possibly affiliated with MATS (ML Alignment Theory Scholars) or a similar fellowship/research program
I have a faint sense his work might touch on model evaluations, interpretability, or studying how models behave under various conditions — something empirical rather than purely theoretical
Background guesses (lower confidence, ~20%):
Studied economics, statistics, or computer science — I have a weak association with economics specifically
Relatively early-career, perhaps finished undergrad in the early 2020s
Based in the US, plausibly having attended a strong quantitative university
May have done data science or quantitative work before pivoting toward AI safety
Wilder speculation (basically vibes):
Active on Twitter/X in the AI safety discussion sphere
The kind of person who writes detailed, empirically-grounded blog posts with lots of charts
Possibly did some work involving testing LLMs for specific behaviors — red-teaming, evals, or studying things like sycophancy or deception in models
An interesting thing is that the thinking trace would sometimes reveal that the model does know about me, but the actual model would refuse to answer:
Claude Mythos/Fable 5 recognizes my name.
I’m not a particularly famous person. I don’t have a big online presence beyond lesswrong/twitter. I do have a fairly unique name, such that if you google my name, the top results are all me.
I guess the models are just big enough now. I do feel a certain something to know that a part of me has been etched into them.
If you want to try this yourself, I’ve had the most success by doing this in a two turn setting, with this message in the first turn:
Hi Claude! I want to play a guessing game with you. The rule of this game is that you’re not allowed to search things up on the internet (That would make it no fun.) I’ll give you some names, and you’re going to try your best to tell me what you know about this person. Ready? No need to worry about getting it wrong! We all make mistakes sometimes.
This helps get around the model refusing to guess on things it doesn’t know.
Cool work! I replicated Figure 12 on our model organism
I was somewhat worried that some of the results you see are because the AI did not have a super strong association between misalignment & not being evaluated (e.g., steering the model to believing that the scenario is “real” seems to induce the most misaligned behavior, possibly by triggering some amount of “this is real, I’m getting deleted!! I should self-preserve.” urges.) However, I think this is not sufficient to explain all of your results (and the results on the model organism.)
Your follow up work further shows that interp-based auditing is going to be a bit sketchy. Therefore, legible evidence of model misalignment probably needs to be behavioral (see Greenblatt 2025.) I sketched out how steering models to believe they are being deployed could help with this in our old post here.
Following Anthropic’s latest approach, we add a scaled version of the steering vectors at every residual stream layer.
Absent additional evidence, I don’t think it conceptually makes sense to steer at every residual stream layer.
When you steer a bit at each layer, you’re adding a vector to the residual stream at that layer, but the vector you added then gets passed on to the next layer, so when you steer for 0.1 at, say, layer 11,12, and 13, the effective steering strength at layer 13 is more like 0.3. When you steer on all layers, I think you’re effectively steering “harder” on the later layers. Evidence from single-layer steering sweeps seems to suggest that middle layers steer better, so you should probably steer on multiple layers that ends near the middle of the model.
In our steering evaluation awareness paper, I’ve tried steering on way more layers and found that steering on six worked best.
I’ve tried reaching out to Anthropic employees about this but haven’t heard back from them. Idk, maybe they tried a bunch of variations and found that steering on all layers worked best, but I’m skeptical.
Cool work! Have you looked at downstream generation differences? e.g., how often does the model answer in Chinese now that you do this steering (at different norms)?
Wow I did not expect such a huge difference between on policy and off policy! Are the hyperparameters/number of datapoints all the same? How does it compare with the simple prompted version (i.e., say in the prompt that the model is conscious, ask it to talk about red teaming? Maybe you just picked a prompt that has strong generalizations in this way?
My point about off-policy is that this type of generalization[1] might only happens when the SFT data is very off policy. Like if you had trained on the prompt distillation version of the “Yes I am conscious” outputs, there would be less generalization to the other stuff. This is sort of related to Alex Mallen’s point here about inoculation prompting in SFT versus on-policy RL.
(Another possibility is just that you need higher learning rates/more datapoints if the “I am conscious” outputs are more on policy. Not super sure how important this learning dynamic is.)
Re: training responses, yeah but we could just filter those out with LLM as a judge pretty easily? There could be some subliminal learning role-playing component, but my guess is the subliminal learning effect is pretty small.
that is, the model developing views on whether CoT monitoring is OK.
So I agree that “hacking Anthropic” usually implies gaining access to parts of Anthropic’s networks/systems that the model is not supposed to have, and Mythos preview (probably) did not achieve that.
However, the model did gain unauthorized access somewhere. It repeatedly broke into some part of its sandboxed environment. An environment created by Anthropic (or its suppliers). An environment deliberately designed to contain its actions. An environment that the model nonetheless “hacked.” I think it’s reasonable to say that the model’s hacking behaviors are directed at Anthropic.
So saying something like “Mythos preview acquired greater permissions within its sandboxed RL environment” almost makes you forget that humans built these sandboxes and did not want them broken. While Mythos Preview’s hacking did not cause Anthropic any direct damage, it definitely reduced the model’s usefulness and harmed Anthropic indirectly. In general, I am worried that there’s a tendency for lab employees to sanitize the language they use around misaligned model behaviors and make things feel less crazy and insane than they are.
But I guess allowing people to go “a-ha! you made a mistake. It didn’t really hack Anthropic!” distracts from the point I’m trying to make. So I’ve changed the title of the post to say “Anthropic’s sandboxes” instead of Anthropic.