I think looking at the number of effective hidden tokens makes sense, but IMO the time-horizons soubling is a better measure of “how much unmonitorable work can the model do”.
I do, though, think I was wrong to focus on model depth, and that a more correct metric is parameter counts, which would need to increase by a factor of 4.2^8(i.e. 16 doublings) to get unmonitorable models on the scale of today’s frontier models. And since doubling the recurrence depth probably gets you much less than 2x effective params (maybe 1.3-1.7?), the recurrence might actually have to be crazy huge to be worrying by itself (taking these numbers seriously, it would have to recur ~2^32 times!!). Also, recurrence probably has very diminishing returns, and trying to double the recurrence amount more than a few times has a hard usefulness cap.
After thinking about this more, I feel less worried about recurrence?
I think looking at the number of effective hidden tokens makes sense, but IMO the time-horizons soubling is a better measure of “how much unmonitorable work can the model do”.
I do, though, think I was wrong to focus on model depth, and that a more correct metric is parameter counts, which would need to increase by a factor of 4.2^8(i.e. 16 doublings) to get unmonitorable models on the scale of today’s frontier models. And since doubling the recurrence depth probably gets you much less than 2x effective params (maybe 1.3-1.7?), the recurrence might actually have to be crazy huge to be worrying by itself (taking these numbers seriously, it would have to recur ~2^32 times!!). Also, recurrence probably has very diminishing returns, and trying to double the recurrence amount more than a few times has a hard usefulness cap.
After thinking about this more, I feel less worried about recurrence?