I feel like looped recurrence isn’t analogous to the paper’s empirical trendline, which is based on a different scaling process. I’m not sure how to model a process that can be roughly estimated as giving some multiplier in layers while maintaining the number of parameters.
It seems likely to me that it is analogous, in the sense that increasing parameters, or increasing recurrence, or increasing CoT length can all increase the task-horizon length. I think that the main question here is how does recurrence fit into the model? From the test-time compute paper, it seems plausible that
(1) looped recurrence up to a few times can be seen as maybe increasing effective parameter count by a few times, and also
(2) there is a pretty hard ceiling on the amount of extra compute this buys you. There might be very fast diminishing returns on recurrence.
Note that if (2) is true then my above point (using recurrence to 8x effective model depth destroys CoT monitoring) is wrong, or at least the premise is impossible—it might be that an effective 8x depth would be very very bad, but recurrence alone can (as far as we know) only buy labs an effective 2x depth/param multiplier
I feel like looped recurrence isn’t analogous to the paper’s empirical trendline, which is based on a different scaling process. I’m not sure how to model a process that can be roughly estimated as giving some multiplier in layers while maintaining the number of parameters.
It seems likely to me that it is analogous, in the sense that increasing parameters, or increasing recurrence, or increasing CoT length can all increase the task-horizon length. I think that the main question here is how does recurrence fit into the model? From the test-time compute paper, it seems plausible that
(1) looped recurrence up to a few times can be seen as maybe increasing effective parameter count by a few times, and also
(2) there is a pretty hard ceiling on the amount of extra compute this buys you. There might be very fast diminishing returns on recurrence.
Note that if (2) is true then my above point (using recurrence to 8x effective model depth destroys CoT monitoring) is wrong, or at least the premise is impossible—it might be that an effective 8x depth would be very very bad, but recurrence alone can (as far as we know) only buy labs an effective 2x depth/param multiplier