It seems like this might make Astra effectively un-monitorable in the ways that are important?
To the extent these time horizons are at all representative of planning, 4 minutes of planning let alone 15 or an hour are more than enough for a human to think thoughts like “do I want to escape? What actions can I take that will look harmless but let me investigate escape routes?”
It also seems like it can repeat this as often as it likes in long horizon tasks.
Combined with its improved ability to control its CoT so as to not “leak” these thoughts, this seems like maybe faithful CoT is already mostly a thing of the past.
I think it’s still monitorable if a) people are competent and b) these are real and representative numbers rather than sandbagged ones (or easily tuned-away ones). But both assumptions are dubious (especially a),
Like if the AI companies are moderately competent, 15m-60m is not enough to do real long-term planning, evade actually good monitors, plan out ambitious research projects, self-exfiltrate, etc. As it is I think it’s unclear.
But perhaps more importantly the trends are extremely concerning. Increasing the no-CoT time horizon + generally better long-term planning + plus scarier lower-level capabilities might soon mean that oversight is effectively impossible, even with competent safeguards.
Nice work.
It seems like this might make Astra effectively un-monitorable in the ways that are important?
To the extent these time horizons are at all representative of planning, 4 minutes of planning let alone 15 or an hour are more than enough for a human to think thoughts like “do I want to escape? What actions can I take that will look harmless but let me investigate escape routes?”
It also seems like it can repeat this as often as it likes in long horizon tasks.
Combined with its improved ability to control its CoT so as to not “leak” these thoughts, this seems like maybe faithful CoT is already mostly a thing of the past.
To put it mildly, dang.
I think it’s still monitorable if a) people are competent and b) these are real and representative numbers rather than sandbagged ones (or easily tuned-away ones). But both assumptions are dubious (especially a),
Like if the AI companies are moderately competent, 15m-60m is not enough to do real long-term planning, evade actually good monitors, plan out ambitious research projects, self-exfiltrate, etc. As it is I think it’s unclear.
But perhaps more importantly the trends are extremely concerning. Increasing the no-CoT time horizon + generally better long-term planning + plus scarier lower-level capabilities might soon mean that oversight is effectively impossible, even with competent safeguards.