Autonomy/agency/long-term planning—seems like a likely but-for cause of most takeover stories. Most takeover stories I can think about entail superhuman long-term planning, or at minimum peak human planning + coordination across instances. Takeovers without that always seem either flimsy or entail seemingly a “magical” technological edge in one or multiple other domains.
According to Reuters investigating the OpenAI HuggingFace incident: “agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI’s infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints”
So it does appear to be something todays models are trying!
Yeah I think of it as a pretty continuous thing right now, though future leaps are of course possible. Today’s models are more agent-y than the models of a year ago, and the models of a year ago more agent-y than the models in 2024.
This is conceptually different quite from but maybe meaningfully related to the METR time horizons.
According to Reuters investigating the OpenAI HuggingFace incident: “agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI’s infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints”
So it does appear to be something todays models are trying!
Yeah I think of it as a pretty continuous thing right now, though future leaps are of course possible. Today’s models are more agent-y than the models of a year ago, and the models of a year ago more agent-y than the models in 2024.
This is conceptually different quite from but maybe meaningfully related to the METR time horizons.