I’d probably want to interpret 8 doublings in no-CoT serial reasoning ability as enabling hidden reasoning equivalent to using 2^8=256 tokens in current models.
This interpretation seems incorrect to me. 8 doublings roughly gives you GPT-5.5′s with-CoT (i.e., unlimited CoT) time horizon, but it certainly can’t do everything it can do with unlimited CoT in 256 tokens. I think the problematic assumption is that one doubling of CoT tokens amounts to one doubling in the time horizon: I wouldn’t expect models to perform twice as well on our single-pass tasks when allowed to output two tokens instead of one.
I think the problematic assumption is that one doubling of CoT tokens amounts to one doubling in the time horizon
I’m not assuming that; I point out the contradiction this leads to in my comment above. The biggest problem with this is that there’s a finite time horizon with unlimited CoT.
My claim is that intuitively, I’d expect that doubling the looping or the number of layers is more similar to doubling the number of CoT tokens than to doubling the time horizon.
Perhaps confusingly, when I said “8 doublings in no-CoT serial reasoning ability,” I meant “whatever vague notion of serial reasoning ability we’re increasing with looping,” not time horizon.
This interpretation seems incorrect to me. 8 doublings roughly gives you GPT-5.5′s with-CoT (i.e., unlimited CoT) time horizon, but it certainly can’t do everything it can do with unlimited CoT in 256 tokens. I think the problematic assumption is that one doubling of CoT tokens amounts to one doubling in the time horizon: I wouldn’t expect models to perform twice as well on our single-pass tasks when allowed to output two tokens instead of one.
I’m not assuming that; I point out the contradiction this leads to in my comment above. The biggest problem with this is that there’s a finite time horizon with unlimited CoT.
My claim is that intuitively, I’d expect that doubling the looping or the number of layers is more similar to doubling the number of CoT tokens than to doubling the time horizon.
Perhaps confusingly, when I said “8 doublings in no-CoT serial reasoning ability,” I meant “whatever vague notion of serial reasoning ability we’re increasing with looping,” not time horizon.