I do AI Alignment research. Currently at METR, but previously at: Redwood Research, UC Berkeley, Good Judgment Project.
I’m also a part-time fund manager for the LTFF.
Obligatory research billboard website: https://chanlawrence.me/
I do AI Alignment research. Currently at METR, but previously at: Redwood Research, UC Berkeley, Good Judgment Project.
I’m also a part-time fund manager for the LTFF.
Obligatory research billboard website: https://chanlawrence.me/
Thanks for writing this. Irrespective of the argumentative/evaluative parts (which seem plausible to me), I think there’s a lot of value in just documenting historical events
In my experience talking to newer people, many are wholly unaware of the history of AI alignment over the past decade, especially the history of OpenAI. A neutral statement of the history of OpenAI (and also Anthropic) has a lot of value. (A very similar problem that you’ve also gestured at, is that newer people in ML don’t know the history of ML theory, and so dismiss it for bad reasons.)
I’m interested in seeing the second post of this sequence (the ML academia engagement), since that’s the part where I was more involved in.
Some parts of mechanistic interpretability are also building the kinds of understanding that could lead to a scientific revolution, but unfortunately they’re not very clearly-demarcated from the parts that might have large effects in other ways (like advancing capabilities)
I think the reason there’s not a clearly-demarcated line between the two is that the mech interp that could lead to a scientific revolution is precisely the research that advances capabilities.
I used to think that mechanistic interpretability was pretty safe, because at that point it not done that much to advance capabilities. With more years of evidence and more thinking, I think I was clearly wrong.
The easy thing to point to are that there have been more mech-interp inspired capability breakthroughs.
But the harder/deeper reason is a sense that deeply understanding model functioning a la the original mech interp dream really just has to lead to capability breakthroughs. First, I’ve come to appreciate that many of the same questions that mech interp touches upon are the same questions that people who do theory for AI capabilities care about (e.g. see the theory of deep learning piece). Second, I’ve learned a lot more about the various insights leading to improvements like the Muon optimizer, better initialization methods, and hyperparameter scaling laws, and it seems that basic mech interp (where you look into the model and poke around at all) provides useful insights that allow you to rederive many of them.
As a corollary, the relative disappointing progress in mech interp partially explains why there aren’t that many mech interp capability break throughs.
Yeah, this was a linkpost failure. Thanks for flagging!
The nanogpt speedrun feels more like developing better methods to culture e coli at a hobbyist level, and quite unlikely to lead to any substantial advancement applicable to the operational efficiency of well-funded companies at the frontier.
If you’ll permit a bit of snark, I think that your comment was wrong even when it was written in October 2025.
The Muon optimizer is the clearest example of a hobbyiest-to-frontier transfer of all the techniques I know. Keller Jordan introduced Muon on specifically the nanoGPT speedrun challenge in a tweet thread from October 2024. (He was unsurprisingly hired to work at OpenAI on pretraining shortly after.) Muon seems to enable stable training at large scales, at least moreso than Adam. As evidence of this, by the time you wrote your comment, Muon was used as an optimizer by MoonShot AI for Kimi K2 as well as Zhipu for GLM-4.5, and has seen continued use (e.g. for GLM-5).
Thank you for writing this, and for standing by your principles.
Oh yeah, forgot to say, in response to this:
But if he loses by 15% it is decent evidence that probably strategy needs to change, including how, and how much, to deploy money.
I think I was unclear about this. At time of writing, what I meant to say is, from the fact that Bores loses by 15%+ (which is much larger than the “large margins” I was thinking of) I think that the takeaway that “close elections could come down to a few k votes, which might be swayable with reasonable amounts of money” is still true. However, the question then becomes “how did we miscalculate and think the race has a good chance of coming so close, when in reality it wasn’t?” And also “in the future, why won’t we make the same mistakes in assessing whether or not races are close?”. (One possible answer may be to sponsor our own independent polls?)
I agree that the actual margin Bores lost by (~4-5%) is only slightly worse than the median prediction on Kalshi/Polymarket, and pretty close to some toy BOTECs that I’ve seen, that had median outcome of Bores losing by 4k votes out of 100k and implied Bores win chance of ~30%.
Also, we have much more liquid prediction markets now to calibrate our models (iirc we only had Metaculus for Carrick Flynn?).
But I’m sure we’ll learn something after people who understand politics better than I do conduct a proper postmortem, and it may well update our value of donations up or down substantially.
Yeah. Moreso than binary win/lose or margin of victory, I think teasing apart whether and how money actually buys votes may update us a lot. If I’m not missing something big, this was by far the most expensive House primary ever, with $50m+ spent if you add up both independent expenditures from Super PACs and candidate committee spending. (I think second is this year’s KY-4, with ~$34m?). So there’s going to be lots of ads and spending to analyze.
One thing that I’m really interested in a post mortem on is the ad buys in the final week of the campaign: the Jobs and Democracy PAC (funded by Public First) was spending $1m/day on Bores in the last week, while Think Big (funded by Leading the Future) spent basically $0 (their last FEC filing was on 6⁄16). Was this multi-million spending spree even net positive, let alone worth the cost? One one hand, this was a ton of ad spend in favor of Bores, and maybe that bought votes. But it also totally destroyed the narrative of “Bores is the underdog fighting against industry”, in a way that might have cost him thousands of votes.
I’m pretty sure that Chris Larsen’s ~$3.5m of ad buys was bad for Bores on net (it destroyed the underdog narrative, associated Bores with a “crypto billionaire” which likely cost votes, and I’ve heard rumors that it was poorly spent.)
Yeah. I think Carrick Flynn did a lot of damage to interest in politics from EA. And the epistemic environment was indeed quite bad.
But I’m sure we’ll learn something after people who understand politics better than I do conduct a proper postmortem,
Almost certainly.
lol that makes sense, thanks for explaining
As I write this, there are around 3 hours left before polls close for this years’s New York’s 12 District Democratic Primary. If you’re a registered democrat in NY-12, you can still vote.[1]
But for those of us who reside elsewhere, there’s little to be done but to wait with bated breath. Will Alex Bores, author of the RAISE act, manage to overcome the millions of dollar spent against him by Leading the Future and demonstrate that AI regulation is not just politically viable but a winning issue? Or will the establishment favorite (and favorite from the start of the race) Micah Lasher succeed in succeeding his mentor Nadler?
I don’t know. As of writing, the prediction markets (Kalshi, PredictIt) have Bores winning at around 28% and Lasher at 72%. If you think you do know the answer, you should go make some money on these markets!
One thing I’m worried about is that people will learn too much from the binary outcome of Bores or Lasher and not on the details of the race. I’m writing this in haste to get it out before the polls close, and we start seeing the outcome, so as to preregister my thoughts.
If Bores loses, some might claim that AI regulation remains politically toxic, and that LTF’s spending was decisive. (I imagine LTF certainly will.) But this is a mistake: win or lose, Bores’s demonstrated that passing AI regulation will not just leave you facing down millions of dollars of Super PAC spending alone. Instead, millions of dollars of Super PAC money was spent on ads championing Bores (in fact, more than what LTF spent!), as well as hundreds of thousands of dollars of donations from AI Safety-concerned individuals.
I know many people who’ve donated to Bores’s campaign, and who are invested in his victory. If he were to lose—especially by a large margin—it might seem tempting to dismiss the whole enterprise of political donations entirely. Similarly, if (somehow) he were to win by a large margin, it might feel like the marginal donation was useless. But I think this too is a mistake.
Ultimately, you can only make decisions based on the information you have. Eric Neyman’s expected value math is correct, and reasonable ex ante. In close elections, even small efforts can help make the difference, and ex ante, this election had a good chance of being very close. If Bores were to lose, or win by a large margin, at most this tells us that his judgment of whether the election would be close was wrong, and even then not by very much.
I am busy, so I do not have time to write a beautiful conclusion or polish this piece. Personally, I hope that despite the unfavorable prediction market odds, Bores wins. But I didn’t write this as an action to affect that outcome. Instead, I wrote it to preregister my claims, such that they’re not seen as post hoc cope after the election results come in.
Nonetheless, here’s my attempt at a conclusion, written in one go:
A phrase I think about a lot these days is the Chinese idiom 尽人事,听天命 (lit. [after you] exhaust human efforts, [then] heed heaven’s mandate (fate)).[2] In the end, all you can do, as a single person in this very large world, is do everything within your power, and then wait with bated breath for the outcome. Unlike the English equivalents (e.g. “Man proposes, Heaven disposes.”), it’s fundamentally an optimistic (or at least motivational) idiom, not a fatalist one.
Rgardless of the outcome—which is outside of the control of any one person, even Bores or Lasher—there will be more elections and political battles to come. Regardless of the outcome, I hope the people around me take the right lesson from the NY-12 election, and continue to do their best, instead of simply resigning to fate.
Consult https://ny12.org/ if you need help finding your polling station!
Claude Opus 4.8 suggests that it should be translated as “Do everything within human power, then accept the will of heaven.”.
Don’t the safetyists think that’s automatically suspicious and subversive? I wouldn’t expect someone like you, who is clearly strongly in favor of model wellbeing, to be involved with the safetyist crowd.”
Wild. It’s sad that this is the case, if it were.
I certainly do expect us to miss plenty of bugs.
To be clear, I’m not critiquing your work with this! And I don’t think “bugs” is the right characterization—I totally expect even a basic fresh reimplementation to catch obvious bugs—rather than some fundamental limitation in the research methodology.
There are also other things that need to be done. In a sane world, there would be multiple replications of every AI safety study (I’m working on that).
Just got around to your comment. I’m glad you’re doing this! In my spare time I’ve done a bunch of lower effort critiques/replications of other research work, one of which I wrote up for InkHaven (at least much lower than your ‘Reevaluating “Model Organisms of Emergent Misalignment”’ piece). I think this is valuable, though I worry that a lot of replication work is too credulous to serve as a bug detection mechanism. (Generally it’s very junior people doing the replication, who understandably hesitant to critique established work, and who lack the context to make some of the more incisive critiques.)
Good citation, that paper seems to have slipped my recollection (probably because it’s less famous, as you said). Added a footnote to clarity.
Good start. Sad this post didn’t get more upvotes, and so I didn’t see it until now.
Some unsolicited feedback on the post:
I would include more description of what the questions are, and how your setup differs from the Redwood/Anthropic one and why. (I was able to find this by reading your repo, but a post shouldn’t require readers read the repo in order to understand it.) This is probably the biggest issue I have with the post. Why didn’t you use the animal welfare setting? Is it because v4 doesn’t care about animals, or did you find it already knew the setting to be artificial?
Similarly, would be good to contextualize your V4/R1 numbers on previous results, e.g. some recent results on recent Anthropic models. For example, the absolute rate of compliance for v4 is a lot lower than Opus 4.5/4.6 etc, but still a lot higher than r1 (ditto compliance gap).
Post would be a lot more readable with a few bar plots to summarize the results, rather than spreading it out in many tables.
Relatedly, would be good to break down which of the questions the models refused/accepted/etc, and see if there are any pattern.
Yeah, the main application of deep learning theory is muP; the main application to safety is probably not that. muP by itself is not relevant to safety, except insofar as it means people don’t use NTKs as their toy model (though they probably weren’t anyways).
I bring up muP because it’s the main (or only) concrete application of deep learning theory; insofar as you dismiss theory b/c there’s no wins, muP is evidence against that conclusion, in the same way that a lack of other wins is evidence for.
Yep
Thanks for the mention!
Amusingly, it was this shortform that caused me to start writing the post: I started drafting a response on the issues I had, and then it ballooned into a full investigation and Ben Sturgeon got pulled in as well.
Yeah, the dense supervision point is what I meant by SFT >> RL for efficiency. You get a bunch more bits per forward pass.
The on policy distillation/dAgger > SFT/behavioral cloning seems like a smaller improvement in comparison to that, but you’re right that it is an improvement.
In Chess, cheating is rampant not at the top professional level (probably) but at the level just below that — iirc there’s a lot of IMs banned for cheating on titled tuesday on chess.com? At least, many of the top players believe that cheating is rampant on online chess (though not amongst top players), and a lot of casual tournaments (eg between streamers) have had people get caught just aping stockfish. And there’s definitely a lot of accusations thrown around for online chess cheating that are generally considered unsubstantiated (the former world champion Kramnik being the most famous serial accuser).
Online chess tournaments not having rampant cheating seems to match the stuff Ashe is saying in their post:
The symbolic camera controls – which would be easy to circumvent for a dedicated cheater – seemed sufficient to curb almost all cheating in a way that threats or impotent references to “fair-play committees” were failing to.
when you add actual barriers to cheating, even if they‘re circumventable, cheating rates drop a lot, especially at the top level.
Of the factors you mention, I’m not sure how FIDE’s willingness to ban compares to Go organizations such as IGF or EGF. Plausible the unified nature might make a difference, but I suspect FIDE’s eagerness to strip titles is not any higher than the go equivalents. My guess is the other factors probably do little if anything: Magnus insinuating Hans Niemann was cheating (or Hikaru’s more direct accusations) probably had little effect in comparison, and Kramnik‘s accusations probably made the cheating problem worse if anything.
If you’re talking about OTB chess, then those tournaments have crazy amounts of security (some would say security theater) to prevent cheating: everyone has to leave their phone outside, the players are scanned with various tools, streams are on a long delay, and so forth.
(And like in Ashe’s post, when people are caught cheating in chess, their justification is normally “I just referenced stock fish occasionally” or “I just used it to suggest moves, I was playing”, and so forth)
I’m personally confused about whether to upvote or downvote this quick take himself.
My guess is Thomas joining OpenAI is probably a mistake, on the priors of “someone says they have a ‘particularly good reason to’ join a lab”. But I also want to encourage Thomas posting this here, because I don’t imagine the decision would’ve received much attention on LessWrong if not for this. (What of all the other people who joined the labs recently?)
In general, the incentives to join the labs are very strong and the incentives against are quite weak. (The model I have in my head is that of a strong current toward the labs—you can swim against it, but it’s oh so tempting to give up and let it take you.) One strategy to create incentives against joining the labs is to downvote and show general hostility about such posts. I’m not sure that’s a strategy that LW as an institution can afford to take.
As an aside, I find it amusing how much the incentives here mirror some of the incentives around voluntary disclosure of safety incidents/voluntary collaboration with METR. Insofar as you’d encourage METR to voluntary collaborate with labs on a limited scope investigation for the HuggingFace incident, or insofar as you’re okay with people praising OpenAI/Anthropic for disclosing safety incidents, my guess is you should encourage this sort of comment via upvoting it on LW.
(I’ve weakly upvoted but strongly disagreed with the top-level comment.)