Blast from the past! Now that it’s been a few years, some of the stages in my mini-scenario have indeed come to pass, albeit a few months later than in the scenario. I’d say we are roughly at stage 4 now:
(1) Q1 2024: A bigger, better model than GPT-4 is released by some lab. It’s multimodal; it can take a screenshot as input and output not just tokens but keystrokes and mouseclicks and images. Just like with GPT-4 vs. GPT-3.5 vs. GPT-3, it turns out to have new emergent capabilities. Everything GPT-4 can do, it can do better, but there are also some qualitatively new things that it can do (though not super reliably) that GPT-4 couldn’t do.
(2) Q3 2024: Said model is fine-tuned to be an agent. It was already better at being strapped into an AutoGPT harness than GPT-4 was, so it was already useful for some things, but now it’s being trained on tons of data to be a general-purpose assistant agent. Lots of people are raving about it. It’s like another ChatGPT moment; people are using it for all the things they used ChatGPT for but then also a bunch more stuff. Unlike ChatGPT you can just leave it running in the background, working away at some problem or task for you. It can write docs and edit them and fact-check them; it can write code and then debug it.
(3) Q1 2025: Same as (1) all over again: An even bigger model, even better. Also it’s not just AutoGPT harness now, it’s some more sophisticated harness that someone invented. Also it’s good enough to play board games and some video games decently on the first try.
(4) Q3 2025: OK now things are getting serious. The kinks have generally been worked out. This newer model is being continually trained on oodles of data from a huge base of customers; they have it do all sorts of tasks and it tries and sometimes fails and sometimes succeeds and is trained to succeed more often. Gradually the set of tasks it can do reliably expands, over the course of a few months. It doesn’t seem to top out; progress is sorta continuous now—even as the new year comes, there’s no plateauing, the system just keeps learning new skills as the training data accumulates. Now many millions of people are basically treating it like a coworker and virtual assistant. People are giving it their passwords and such and letting it handle life admin tasks for them, help with shopping, etc. and of course quite a lot of code is being written by it. Researchers at big AGI labs swear by it, and rumor is that the next version of the system, which is already beginning training, won’t be released to the public because the lab won’t want their competitors to have access to it. Already there are claims that typical researchers and engineers at AGI labs are approximately doubled in productivity, because they mostly have to just oversee and manage and debug the lightning-fast labor of their AI assistant. And it’s continually getting better at doing said debugging itself.
(5) Q1 2026: The next version comes online. It is released, but it refuses to help with ML research. Leaks indicate that it doesn’t refuse to help with ML research internally, and in fact is heavily automating the process at its parent corporation. It’s basically doing all the work by itself; the humans are basically just watching the metrics go up and making suggestions and trying to understand the new experiments it’s running and architectures it’s proposing.
(6) Q3 2026 Superintelligent AGI happens, by whatever definition is your favorite. And you see it with your own eyes.
The caveat is that we don’t seem to have true continual learning yet. On the other hand, they ARE updating the models about once a month, by training them on additional data, and Anthropic researchers claim that continual learning will be solved this year. And the rest of the stage 4 description seems basically right.
If it was truly at “what people imagined stage 4 to be”, you might think that you/Ege/Ajeya are supposed to assign 90%/30%/75% to AGI within the next ~2.5 years. (Though ofc you could have had other updates that cancel out something here.) I think in fact all of you are lower than your own numbers there.
This newer model is being continually trained on oodles of data from a huge base of customers; they have it do all sorts of tasks and it tries and sometimes fails and sometimes succeeds and is trained to succeed more often.
My sense is that this isn’t a big part of the story for how new models’ capabilities are being increased. Though I don’t think we know for sure either way.
Now many millions of people are basically treating it like a coworker and virtual assistant. People are giving it their passwords and such and letting it handle life admin tasks for them, help with shopping, etc. and of course quite a lot of code is being written by it.
This seems accurate for coders. Is it true for people who aren’t coders? It’s not really true for my job or life admin tasks (like I use the models a fair bit, but it’s more in chat-bot mode than in agent-mode / trusting the models to do a lot of stuff for me) but maybe it’s more true for others.
unfortunately i think the scenarios are vague enough that as a practical matter it will be tricky to adjudicate or decide whether they’ve happened or not
has come true (unlike many of my actual predictions about the rate of progress, which were consistently too bearish about the median world due to how I was hedging against the current paradigm of scaling running into a bottleneck).
Looking back at the dialogue, I can’t actually remember how I interpreted Daniel’s stage 4 when I offered my probability estimate of 30%. I don’t currently think there’s a 30% chance of “superintelligent AGI” over the next few years, which makes me think what I had in mind for stage 4 was something more impressive than what actually ended up happening.
This also matches my rough sense that the current world is more like my 75th − 80th percentile world from late 2023 and not the 90th+ percentile that would be needed to justify an update from 6% to 30%, which is the update I said I would make if I observed (4) happen.
tbf I myself think that it’s not clear we’re in stage 4, I think stage 3 is arguably where we’re at. But we seem closer to 4 than to 3 imo. Maybe we should interpolate between them?
I think we’re in stage 4 in some ways and stage 3 in others, yeah.
Looking back at the scenario, I think if we just push your predictions forward by 70% (multiply the time gap from when they were made to when the predicted events would happen by 1.7) they look pretty good. Arguably (2) happened in Q2 2025 (first time I remember AI agents being a big deal was with o3) and (3) happened in Q4 2025 (I think there was a transition in how autonomously the agents could function around that time, it’s definitely the time when I felt comfortable just letting the AI write code without manually reviewing everything it was writing).
If we take that view, (4) will probably “really happen” in Q4 2026.
Note that the scenario I gave wasn’t actually a prediction, or at least, it wasn’t my median world. I said elsewhere in thread that my median was 2027 for AGI, and implied that my median for ASI was more like 27/28:
To be clear, my view is that we’ll achieve AGI around 2027, ASI within a year of that, and then some sort of crazy robot-powered self-replicating economy within, say, three years of that. So 1000x energy consumption around then or shortly thereafter (depends on the doubling time of the crazy superintelligence-designed-and-managed robot economy).
So my actual prediction would have been like the scenario but stretched out another 1-2 years.
Blast from the past! Now that it’s been a few years, some of the stages in my mini-scenario have indeed come to pass, albeit a few months later than in the scenario. I’d say we are roughly at stage 4 now:
The caveat is that we don’t seem to have true continual learning yet. On the other hand, they ARE updating the models about once a month, by training them on additional data, and Anthropic researchers claim that continual learning will be solved this year. And the rest of the stage 4 description seems basically right.
If it was truly at “what people imagined stage 4 to be”, you might think that you/Ege/Ajeya are supposed to assign 90%/30%/75% to AGI within the next ~2.5 years. (Though ofc you could have had other updates that cancel out something here.) I think in fact all of you are lower than your own numbers there.
My sense is that this isn’t a big part of the story for how new models’ capabilities are being increased. Though I don’t think we know for sure either way.
This seems accurate for coders. Is it true for people who aren’t coders? It’s not really true for my job or life admin tasks (like I use the models a fair bit, but it’s more in chat-bot mode than in agent-mode / trusting the models to do a lot of stuff for me) but maybe it’s more true for others.
I think my prediction about
has come true (unlike many of my actual predictions about the rate of progress, which were consistently too bearish about the median world due to how I was hedging against the current paradigm of scaling running into a bottleneck).
Looking back at the dialogue, I can’t actually remember how I interpreted Daniel’s stage 4 when I offered my probability estimate of 30%. I don’t currently think there’s a 30% chance of “superintelligent AGI” over the next few years, which makes me think what I had in mind for stage 4 was something more impressive than what actually ended up happening.
This also matches my rough sense that the current world is more like my 75th − 80th percentile world from late 2023 and not the 90th+ percentile that would be needed to justify an update from 6% to 30%, which is the update I said I would make if I observed (4) happen.
tbf I myself think that it’s not clear we’re in stage 4, I think stage 3 is arguably where we’re at. But we seem closer to 4 than to 3 imo. Maybe we should interpolate between them?
I think we’re in stage 4 in some ways and stage 3 in others, yeah.
Looking back at the scenario, I think if we just push your predictions forward by 70% (multiply the time gap from when they were made to when the predicted events would happen by 1.7) they look pretty good. Arguably (2) happened in Q2 2025 (first time I remember AI agents being a big deal was with o3) and (3) happened in Q4 2025 (I think there was a transition in how autonomously the agents could function around that time, it’s definitely the time when I felt comfortable just letting the AI write code without manually reviewing everything it was writing).
If we take that view, (4) will probably “really happen” in Q4 2026.
Yep! Thanks.
Note that the scenario I gave wasn’t actually a prediction, or at least, it wasn’t my median world. I said elsewhere in thread that my median was 2027 for AGI, and implied that my median for ASI was more like 27/28:
So my actual prediction would have been like the scenario but stretched out another 1-2 years.
Makes sense.