My personal website is https://jimfund.com/.
james oofou
why subject us to gary marcus tweets
Profile pictures solve this.
So, if Gemini 3.5 Flash is perhaps an 3.1-pro-sized model (is that what we mean by ‘big model’?), then might Gemini 3.5 Pro (scheduled for June and already being used internally at GDM) be a Mythos-sized model?
Here’s a very relevant article I wrote back then about LLM solving open problems (it references Have LLMs Generated Novel Insights) https://jimfund.com/math.html
i predict that on jan 1 2027 this will have already come to pass
I don’t get how any of these things are obviously nonabundant in an ASI world.
judging from Altman’s recent claims that something very capable was still being trained
Can you cite the timestamp or recall the quote you are referring to? I asked Gemini and I watched the video but didn’t find anything which pointed to something very capable still being trained.
Yes, that is implied.
I think it’s pretty clear we’re in the early stages of an intelligence explosion. Researchers are already sped up ~20% as of early February, so doubling times are shortening. Each doubling of time horizon increases researcher productivity by a greater amount than the previous doubling. There’s enough substitutivity that compute won’t be enough of a constraint to block a singularity.
We hit a bit of an inflection point starting in late November ’25 where AI systems provide decent uplift to engineers. Maybe a 5–10% uplift in total factor productivity for frontier AI research lab engineers. So that’s directly applicable to the rate of algorithmic progress, training run supervision, etc. but not overall progress (because increasing programming productivity doesn’t directly increase compute).
I expect uplift will increase superlinearly (and increasingly so) with model time horizon, and model time horizon to increase hyperexponentially (current doubling time of 2–3 months). Uplift with Opus 4.6 (released Feb 5) is probably about 20%. So, uplift will increase hyperexponentially, more than doubling every few months then weeks...
Uplift should reach a few hundred percent by Q4 this year (although uplift will obsolesce pretty quickly as a concept as models increasingly work independently of engineers). Then, conditional on limited compute not being too much of a constraint (and I don’t think it will be) we’ll get a singularity-style intelligence explosion before end of year. Timeline is pretty sensitive to the exact values of all these parameters.
I guess they’re losing money in the short-term but gaining training data and revenue (which helps them raise funds). It’s not clear to me that this is harming the lab in expectation.
Them: “I think X”
You: “That’s wrong because Z”
Them: “I think you’re just disagreeing because you’d not open-minded enough”
You: “What makes you think that?”
Them: “I think it because Y”
What do they say for ‘Y’? That seems the part that actually constitutes their argument and which you will be able to call out if they’re making a mistake.
Near-Future Fiction III
September 2026
Knowledge worker productivity has become relatively uncoupled from pre-ChatGPT levels, as the hardest technical tasks which these workers did at that point in time in a given working day can now in most cases be carried out autonomously by AI.
Programmers therefore begin to work at a higher level of abstraction, guiding AI workers, managing projects at a higher level.
Meanwhile, much progress is being made in robotics. Full self-driving has been achieved.
And AI has begun making novel breakthroughs. This enables continual learning: the AI’s new discoveries open up many new avenues for further discoveries, which open up many more such avenues, ad infinitum.
Image from a recent OpenAI talk
December 2026
Successful reinforcement learning on the September worker AIs has enabled AI to operate at that higher level of abstraction which software engineers had retreated to. Human knowledge workers are therefore relegated to maintenance work and helping out when the few remaining weak points in these AI systems cause trouble.
The difficulty of progress in AI intelligence relative to human intelligence begins reducing rapidly as time horizons extend beyond a few hours. At horizons of this length, human begin relying on caching tricks, iteration, brute force, etc. rather than, beyond a certain point, making fundamentally more difficult leaps of insight.
Early 2027
Humans are cut out of the loop entirely in knowledge work. The robotics explosion happens. Robots gradually replace humans in physical labour. AI progresses far beyond human-level.
Mid 2027
Humans fully obsolesce. Mind upload is achieved.
Notes
I assume a 3-month METR doubling time. We should expect lower doubling times over time given increased investment in AI, increased contribution by AI to progress, and decreased difficulty per double. Also, OpenAI has communicated that we should expect several major breakthroughs from them in 2026.
We should expect doubling times to decrease even further with time, although in a discontinuous way so it’s impossible to predict with much accuracy when it will happen.
It seems that your argument is based on high confidence in a METR time-horizon doubling time of roughly 7 months. But the available evidence suggests the doubling time is significantly lower.
In recent years we have observed shorter doubling times:
And what we know about labs’ internal models suggests this faster trend is holding up:
An important piece of evidence is OpenAI’s Gold performance at the International Mathematics Olympiad (IMO):
IMO participants get an average of 90 minutes per problem.
The gold medal cutoff at IMO 2025 was 35 out of 42 points (~83%)
They needed to get 5⁄6 problems fully correct (each question awards a maximum of 7 points), or a number of points equivalent to that.
This is a bit rough, but if their model had a METR-80 greater than 90 minutes then we would expect OpenAI to achieve Gold at least 50% of the time.
OpenAI staff members stated that a publicly released model of this capability could be expected at roughly the end of the year (and our METR trends are of course projections of publicly available models).
So, this implies a METR-80 greater than 90 minutes at December 2025.
The projected METR-80 according to a 3.45 month doubling time is 98 minutes.
So the Gold performance which was a massive surprise to many is actually right on-trend for 3.45 month doubling times. Of course, one might object that OpenAI may have just gotten lucky. But Google also got Gold! so we have two points of data.
And here’s a recent comment from Sam Altman where he states that he expects time-horizons days in length in 2026:and as these go from multi-hour tasks to multi-day tasks, which I expect to happen next year
Which a 7 month doubling time would not achieve, but which is in line with a doubling-time of 3.45 (that would get us to a time-horizon of roughly 3 days in December, 2026).
And here’s a recent comment from Jakub Pachocki:
Your argument that OpenAI stole money here is poorly thought-out.
OpenAI’s ~$500b valuation priced in a very high likelihood of it becoming a for-profit.
If it wasn’t going to be a for-profit its valuation would be much lower.
And if it wasn’t going to be a for-profit the odds of it having any control whatsoever over the creation of ASI would be very much reduced.
It seems likely public gained billions from this.
The text “the website of the venue literally says” appears twice in your post. The first time it appears seems to be a mistake and isn’t followed by a quotation.
Is this distinct from the problem of induction?
I’ve been looking for science fiction set in the late 2020s, and which addresses continued AI progress, for a few years now. Everything else just feels so totally disconnected from any plausible future. Very happy to have found your writing.
You are misunderstanding what METR time-horizons represent. The time-horizon is not simply the length of time for which the model can remain coherent while working on a task (or anything which corresponds directly to such a time-horizon).
We can imagine a model with the ability to carry out tasks indefinitely without losing coherence but which had a METR 50% time-horizon of only ten minutes. This is because the METR task-lengths are a measure of something closer to the complexity of the problem than the length of time the model must remain coherent in order to solve it.
Now, a model’s coherence time-horizon is surely a factor in its performance on METR’s benchmarks. But intelligence matters too. Because the coherence time-horizon is not the only factor in the METR time-horizon, your leap from “Anthropic claims Claude Sonnet can remain coherent for 30+ hours” to “If its METR time-horizon is not in that ballpark that means Anthropic is untrustworthy” is not reasonable.
You see. the tasks in the HCAST task set (or whatever task set METR is now using) tend to be tasks some aspect of which cannot be found in much shorter tasksanyany. That is, a task of length one hour won’t be “write a programme which quite clearly just requires solving ten simpler tasks, each of which would take about six minutes to solve”. There tends to be an overarching complexity to the task.
Maybe AI science as venture capital. People with money bet on which new area of science to spend compute on.