Having just played around with the model for the first time, I think this sort of scribbling is probably susceptible to cognitive biases like “I want the lines to evenly cover the space of strictly possible futures” or something. There are probably dozens of possible short-timelines futures, but once you scribble a single line to represent short timelines, it feels a little stupid to draw ten more lines covering the exact same timeline.
Also, I think your forecast was fine. Here’s some math.
Below is a basic exponential:
In which is a time horizon,[1] is the number of time horizon doublings per year (i.e. one doubling every 6 months = two doublings per year ⇒ = 2), and is the number of years since .
If we take the METR graph at face value, time horizons double ~twice per year[2]:
And in April 2026, the state-of-the-art model was Mythos Preview, which could do ~3 hour tasks at 80% reliability:
So by October 2028, 2.5 years later, we would expect SOTA to have a time horizon of 96 hours:
And we don’t get >1 month time horizons until April 2030:
Alternatively, if we plug in METR’s exact values (, ), we find that SOTA time horizons won’t hit > 1 month until ~4.6 years, aka 4 years and 7 months, aka November 2030:
Reaching a >1 month time horizon by October 2028 would require timelines to double 3.2 times per year, aka every 3.75 months, aka almost twice as fast as reality. We could also be wrong about (e.g. = 24 hours, still assuming 2 doublings per year) or we could be wrong about both parameters (e.g. = 6 hours, = 2.8[3]).
I think this sort of scribbling is probably susceptible to cognitive biases like “I want the lines to evenly cover the space of strictly possible futures” or something.
Well, the idea of the scribbles was to try to elicit my subjective beliefs, so in some sense, I don’t want to avoid cognitive biases. But I think what you’re saying is that the collection of lines might not actually represent my subjective beliefs? I think that’s true, and in fact the failure to account for enough chance of acceleration is a valid example of that.
I think you’re also suggesting specifically that looking at the collection of lines might cause that problem. I actually agree that that would be a danger, but in the actual interface I used, you could only see one line at a time, which was supposed to fix that issue. (The idea was to “sample many I.I.D. futures”, without thinking about collective behavior, if that makes sense.)
(Also, could you have factored acceleration into your scribbles? Probably. But then you wouldn’t be outside-view forecasting anymore, right?)
That would be less outside-view, certainly, but I wasn’t necessarily trying to avoid using the inside view. So I don’t think I can use that as an excuse!
But I think what you’re saying is that the collection of lines might not actually represent my subjective beliefs?
Something like that, yes.
in the actual interface I used, you could only see one line at a time, which was supposed to fix that issue. (The idea was to “sample many I.I.D. futures”, without thinking about collective behavior, if that makes sense.)
Ah, I see. I used the interface too, I just didn’t get the memo and kept looking at / thinking about the other lines I’d drawn.
Having just played around with the model for the first time, I think this sort of scribbling is probably susceptible to cognitive biases like “I want the lines to evenly cover the space of strictly possible futures” or something. There are probably dozens of possible short-timelines futures, but once you scribble a single line to represent short timelines, it feels a little stupid to draw ten more lines covering the exact same timeline.
Also, I think your forecast was fine. Here’s some math.
Below is a basic exponential:
In which is a time horizon,[1] is the number of time horizon doublings per year (i.e. one doubling every 6 months = two doublings per year ⇒ = 2), and is the number of years since .
If we take the METR graph at face value, time horizons double ~twice per year[2]:
And in April 2026, the state-of-the-art model was Mythos Preview, which could do ~3 hour tasks at 80% reliability:
So by October 2028, 2.5 years later, we would expect SOTA to have a time horizon of 96 hours:
And we don’t get >1 month time horizons until April 2030:
Alternatively, if we plug in METR’s exact values ( , ), we find that SOTA time horizons won’t hit > 1 month until ~4.6 years, aka 4 years and 7 months, aka November 2030:
Reaching a >1 month time horizon by October 2028 would require timelines to double 3.2 times per year, aka every 3.75 months, aka almost twice as fast as reality. We could also be wrong about (e.g. = 24 hours, still assuming 2 doublings per year) or we could be wrong about both parameters (e.g. = 6 hours, = 2.8[3]).
If you want to play around with the math yourself, here’s a link on Desmos.
(Also, could you have factored acceleration into your scribbles? Probably. But then you wouldn’t be outside-view forecasting anymore, right?)
I’ll be using hours as my unit, but it doesn’t really matter what this unit is: you could use anything and the equation would still work.
I’m being generous here: their original doubling time claim was actually 7 months, but the math is cleaner this way.
Note that is by far the most important variable: doesn’t change the graph that much because exponentials are weird.
Well, the idea of the scribbles was to try to elicit my subjective beliefs, so in some sense, I don’t want to avoid cognitive biases. But I think what you’re saying is that the collection of lines might not actually represent my subjective beliefs? I think that’s true, and in fact the failure to account for enough chance of acceleration is a valid example of that.
I think you’re also suggesting specifically that looking at the collection of lines might cause that problem. I actually agree that that would be a danger, but in the actual interface I used, you could only see one line at a time, which was supposed to fix that issue. (The idea was to “sample many I.I.D. futures”, without thinking about collective behavior, if that makes sense.)
That would be less outside-view, certainly, but I wasn’t necessarily trying to avoid using the inside view. So I don’t think I can use that as an excuse!
I wonder if you would get better results if you scribbled with both log and constant axis
Something like that, yes.
Ah, I see. I used the interface too, I just didn’t get the memo and kept looking at / thinking about the other lines I’d drawn.