Go ahead! Post a link here when you’re done.
alkjash
I wonder if the existence of evaluators at all will have a meaningful positive impact regardless of the competence of said evaluators.
https://www.anthropic.com/institute/measuring-pace-of-ai-development
One could read the above post as Anthropic trying to get ahead of possible problems revealed by evaluators, they sure spend a lot of paragraphs explaining why 6% of research compute to safety is more than it sounds.
“Safety research tends to use less compute than frontier training runs by its nature, so compute is an imperfect proxy for how much a company focuses on safety. This is because safety research consists of individual researchers designing experiments, which is time-consuming even though running the experiments is not particularly compute-intensive. The value of this metric, therefore, is less the absolute numbers and more that it provides a straightforward mechanism to compare like with like, across developers and over time.”
Not an expert but I could easily see the opposite spin being true: “safety researchers forced into time-consuming hand-crafted experiments because they’re bottlenecked by compute allocations.”
Sorry, just to reiterate, the LLM generated few of the words but did help with some minor editing.
My AI safety essay was flagged for LLM content (I think because I had Claude generate the references and links), so I started putting everything I write in LLM content blocks to not deal with the headache.
I think this was an overreaction.
Ok, thanks for the insight, I’ve decided the input of LLMs is insignificant enough that I should just drop the tags.
My last two posts seem to be the most unpopular things I’ve ever written. Are they actually unusually bad somehow or is it just the LLM blocks turning people off?
I use LLMs for some fact-checking and line-editing at the end of each writing session; both posts detect as 100% human-written in pangram. Would it be dishonest to just drop the LLM content blocks?
I’m much less sure of that, but maybe you have better models of Sam (and the class of tech oligarchs we’re really talking about)?
A substantial fraction of people I talk to fall back to “maybe it would be nice to be kept as a pet by a superintelligence.”
Thank you for the detailed overview!
I have added ILIAD onto the list and strongly updated towards ILIAD being one of the top places to point mathematicians to. I am glad this exists!
Fwiw, I did not come away with the same positive impression from browsing the website for one minute. Maybe I’m particularly bad at navigating websites but the only visible product I found was 6 youtube videos posted within the last month, one anonymous testimonial from an Intensive, and no concrete evidence of events occurring before Sep 2026.
Sorry, I agree with [it is possible to believe in higher marginal tax rate and pay extra voluntary tax] but disagree with [if you believe in higher marginal tax rate you should pay extra voluntary tax]. Giving examples of the former does little to change my mind on the latter.
Oh wow! I am actually noticeably confused by this and the other comment!
Some thoughts, not necessarily disagreements:
I consider LW as a noncentral element of “social media” (Fable agrees), so, well akshually you do use social media.
After thinking about it, I agree that the ceasefire one is just wrong.
I have not even considered that one would pay extra voluntary tax. I am confused about your conflation of “not exploiting tax loopholes” with “paying extra voluntary tax.” Separately, I would guess that the vast majority of people who genuinely believe in a higher marginal tax rate do not unilaterally pay voluntary tax, and I do not consider them to be hypocrites.
Hmm… I don’t know how much of a norm that is. If I was somehow capable of solving NS in two weeks and heard rumors that a colleague was getting close I would probably also sprint for 2 weeks to get a preprint out, and then send it over and ask to coauthor if they had enough of a writeup/ a similar proof. I don’t think anyone would fault me for this either.
Personally I was way more concerned about the shady language from Bubeck around “why would you ruin your career?” But it could be conceivably justified by context so I’m still not absolutely sure what to think about that.
I’m seeing a lot of great snappy responses to “How exactly would AI kill us all?” but not as many on “If they actually believed in slowdown they’d just slow down.” Here’s some attempts:
If you actually believed in higher marginal tax rate, you’d just pay extra voluntary tax.
If the US actually believes in nuclear nonproliferation, it would just dismantle its own nukes.
If you actually believed social media is bad for everyone, you’d just stop using social media.
If you actually believed in a ceasefire, you’d just stop shooting first.
I’m actually not sure which specific norm you think they broke, and you expand on this?
My impression is that the math research team at OpenAI are mostly mathematicians who genuinely care about academic norms. Several of my conversations with my friend on that team boil down to me asking him to stop worrying about doing right by normies and start worrying about alignment.
I think you are underestimating how much restraint they are already showing, as researchers who are used to sprinting to the arXiv the moment they have a proof, to sit on the gigantic hoard of results that they surely have stashed away, these results including problems many of their friends and collaborators have studied and pined for for decades.
(Pure speculation here), but I would guess that this is part of why they sprinted so hard on the NS project—finally they found an Anthropic project that they felt it was fair game for them to go all in on and take credit publicly for.
My two cents on the Navier-Stokes drama (as someone who knows mathematicians at both labs):
It seems that Levent and Tristan treat this project as a private personal research project that they’ve been working on for the past year, with some LLM assistance. They are incensed that a trillion-dollar company heard a rumor about their personal work and decided to devote unprecedented massive amounts of compute at breakneck speed to scoop them in particular, and see this behavior as a gross violation of academic research norms.It seems that Bubeck et al. at OpenAI are focused on the fact that Levent used Anthropic internal models in this private personal research project. Their story is that they have been extremely careful not to tread on math academia (my understanding is that they are sitting on a mountain of unreleased results and a substantial fraction of mathematical progress in 2026 is known only to OpenAI/Anthropic insiders). However, in their frame any project that uses Anthropic internal models is an Anthropic project (possibly legally so) and it is fair play for one trillion-dollar AI company to race literally as hard as possible against another. They are nonplussed that this race between two AI labs is being reframed as OpenAI bullying the little guy.
My issue with both Iliad and Jacob’s page and org is that afaict these have only existed for weeks to months, so I am hesitating to recommend them next to programs that have ever produced alumni … for example it looks like Jacob’s page is already a to-be-deprecated version of his new https://maisi.org/ which has existed for days.
Ok, I will at least look into it soon, thanks for the tip!
I do agree that a lot of low-hanging fruit comes from “lean on social capital to recruit people more than you’re normally comfortable with.” That being said, it does feel like almost always recruitment strategies actually scale off raw power level in some fair way (not sure I have the right vocab for this), e.g. I found that I have pretty high success x-risk-pilling my friends and colleagues and they have relatively low success pilling theirs afterwards, because they’ve thought about alignment for like a month. So the higher-order network effects are much smaller than I expected given the initial success.
Levels of Responsibility
I’ve been psychologically blocked for many years about contributing to alignment work, and recently unblocked. Not sure how generalizable this is, but wanted to share what was blocking me:
I had two modes of thinking that I label as [normal civic responsibility] and [heroic responsibility].
Roughly speaking, civic responsibility is act in such a way that if everyone acted this way, the problem would be solved. Going out to vote, emailing your senator, personally committing to not pushing the button.
Heroic responsibility is just solve the problem, all by yourself if necessary. Move p(doom) appreciably by yourself, maximize marginal impact, start a new moonshot org.
What I’m coming to realize is that there is space for a sliding scale, especially regards to alignment work, between civic and heroic responsibility. Starting with civic responsibility and going up, Responsibility Level n is act in such a way that if 1/n fraction of humanity acted this way, the problem would be solved. n = 1 is civic responsibility, n = 8 billion is heroic responsibility. Hill-climbing on n seems to be the right way for my brain to approach personal responsibility towards the world.
This situation seems more muddy than it first appeared to me.
In particular https://x.com/SebastienBubeck/status/2097379411691516310 claims that “it was admitted that internal Anthropic models had been used in [Levent’s and Tristan’s] proof of Euler blowup,” which essentially calls into question Buckmaster’s claim
“My work with Levent has been a purely personal collaboration, free of any institutional agreements or official involvement by either of our employers.”
Once internal models are in play, it seems to me that the character of the situation changes from “OpenAI bullied and scooped the little guy” to “OpenAI’s internal models scooped Anthropic’s internal models.”
I am a Chinese American math professor at Georgia Tech (https://alkjash.github.io/), and I have been passively interested in understanding how to message to China although it’s quite intimidating.
Please contact me if you’d like to chat.