AI 2040 seems substantially too pessimistic about interpretability. I’d be surprised if it was right about it.
For reference, the scenario describes MechInt as becoming useful in 2035 in the following way.
Other AIs work on mechanistic interpretability, the science of “mind reading” AIs from their weights and activations. Early versions of this were tried in the 2020s, but now it starts to significantly outperform common-sense observation of AI behavior. Researchers can often determine whether an AI is honest or lying; and sometimes trace the psychology that led to a decision.
First problem: I don’t think this is internally consistent.
The scenario attributes this progress to “AI work.” This seems fair.
What doesn’t seem reasonable is for it to take till 2035. According to the scenario, in 2033, two years earlier, only 50% of citizens people have employment, and the median US citizen is being paid 200k dollars a year from AI labor. It seems to me pretty implausible that we can have substitution of half of US human labor notably before we get gigantic levels of AI uplift from mechanical interpretability. I’d expect the opposite: enormous levels of uplift of MI notably before mass unemployment.
Or, in the scenario, I believe the intelligence explosion would have taken place in the area of 2029-2030 without intervention? But I expect AIs that could cause an intelligence explosion clearly could help a ton with mechanical interpretability / model internals stuff.
Second problem: I think like, the scenario isn’t adjusting for how insanely new the field is? Like Olah invented the term in 2020. So if it takes till 2035 for us to get extremely useful progress, then it will have taken 9 years—longer than the amount of time the field has really existed—to have gotten useful progress.
And several of those years the field existed it was like… a tiny handful of people. I think most progress was in the last three years (SAEs, j-space, NLA, etc), because three years ago it had a fraction of the resources. Even if we just account for field growth simply because of growth of human interest, I think I’d expect 1.5x-6x as much progress in the next three years as in the entire history of the field beforehand. Given AI assistance, I expect more like… 4x-80x? Something like that? I think this is a pretty tame assessment looking at lines on curves.
So yeah, I think AI 2040 is substantially underestimating MI (using MI as a broad term “models internals,” etc). The scenario has adversarially misaligned AIs in the ~2031 zone, which we only find out were adversarially misaligned afterwards—I think that’s quite unlikely to happen.
Interp via model internals is a bit useful now (in mid 2026)
It will be modestly useful in 2029 (assuming the AI 2040 scenario)
It will be very useful in 2035 (assuming the AI 2040 scenario)
(I don’t think this interp will that much be “mech interp” which seems to align with what you are saying.)
Thus, I think this text from AI 2040 that you quote probably somewhat understates how useful interp is at that point in the scenario. Your critique seems reasonable overall, though I’m probably moderately less bullish on interp than you seem to be.
AI 2040 seems substantially too pessimistic about interpretability. I’d be surprised if it was right about it.
For reference, the scenario describes MechInt as becoming useful in 2035 in the following way.
First problem: I don’t think this is internally consistent.
The scenario attributes this progress to “AI work.” This seems fair.
What doesn’t seem reasonable is for it to take till 2035. According to the scenario, in 2033, two years earlier, only 50% of citizens people have employment, and the median US citizen is being paid 200k dollars a year from AI labor. It seems to me pretty implausible that we can have substitution of half of US human labor notably before we get gigantic levels of AI uplift from mechanical interpretability. I’d expect the opposite: enormous levels of uplift of MI notably before mass unemployment.
Or, in the scenario, I believe the intelligence explosion would have taken place in the area of 2029-2030 without intervention? But I expect AIs that could cause an intelligence explosion clearly could help a ton with mechanical interpretability / model internals stuff.
Second problem: I think like, the scenario isn’t adjusting for how insanely new the field is? Like Olah invented the term in 2020. So if it takes till 2035 for us to get extremely useful progress, then it will have taken 9 years—longer than the amount of time the field has really existed—to have gotten useful progress.
And several of those years the field existed it was like… a tiny handful of people. I think most progress was in the last three years (SAEs, j-space, NLA, etc), because three years ago it had a fraction of the resources. Even if we just account for field growth simply because of growth of human interest, I think I’d expect 1.5x-6x as much progress in the next three years as in the entire history of the field beforehand. Given AI assistance, I expect more like… 4x-80x? Something like that? I think this is a pretty tame assessment looking at lines on curves.
So yeah, I think AI 2040 is substantially underestimating MI (using MI as a broad term “models internals,” etc). The scenario has adversarially misaligned AIs in the ~2031 zone, which we only find out were adversarially misaligned afterwards—I think that’s quite unlikely to happen.
I think I expect:
Interp via model internals is a bit useful now (in mid 2026)
It will be modestly useful in 2029 (assuming the AI 2040 scenario)
It will be very useful in 2035 (assuming the AI 2040 scenario)
(I don’t think this interp will that much be “mech interp” which seems to align with what you are saying.)
Thus, I think this text from AI 2040 that you quote probably somewhat understates how useful interp is at that point in the scenario. Your critique seems reasonable overall, though I’m probably moderately less bullish on interp than you seem to be.
I wouldn’t be surprised if AI is mildly superintelligent at most if not all of mech interp in 1-2 years when the ~100T models come out.