This suggests it would be better for safety to scale by keeping model size constant, and using longer chains of thought, than by increasing model size to the point where it can do complex thoughts in a single forward pass.
I wonder what the limits of this are though? Can you get to arbitrary levels of intelligence by just increasing thinking time?
This suggests it would be better for safety to scale by keeping model size constant, and using longer chains of thought, than by increasing model size to the point where it can do complex thoughts in a single forward pass.
I wonder what the limits of this are though? Can you get to arbitrary levels of intelligence by just increasing thinking time?