“Take the worry that AIs will accidentally be trained to be deceptive. Sure, it’s possible. But we’re not running reinforcement learning over year-long trajectories — for now, we’re running it over a week at most. The natural prediction is that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking.”
“Many safety controls for AI assistants are designed around individual actions. If an action is disallowed, it is blocked. If it is sensitive, the system asks the user for explicit approval. But long-running models, whose actions may unfold autonomously over hours, days, or even weeks, challenge this setup: monitoring individual actions no longer suffices to track the intent of the overall trajectory.”
Bolding text here for emphasis. Rohin said “we’re running [reinforcement learning] over a week at most,” and the OpenAI post says a model’s actions “may unfold autonomously over hours, days, or even weeks...” I’m wondering if the difference here suggests OpenAI may be running RL over trajectories longer than a week. If the answer is “That’s what it sounds like,” I wonder if that would change Rohin’s “natural prediction...that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking.”
This is something Rohin Shah said in his interview on the 80,000 Hours podcast published June 2, 2026. From the transcript:
“Take the worry that AIs will accidentally be trained to be deceptive. Sure, it’s possible. But we’re not running reinforcement learning over year-long trajectories — for now, we’re running it over a week at most. The natural prediction is that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking.”
This is something OpenAI wrote in “Safety and alignment in an era of long-horizon models,” published July 20, 2026:
“Many safety controls for AI assistants are designed around individual actions. If an action is disallowed, it is blocked. If it is sensitive, the system asks the user for explicit approval. But long-running models, whose actions may unfold autonomously over hours, days, or even weeks, challenge this setup: monitoring individual actions no longer suffices to track the intent of the overall trajectory.”
Bolding text here for emphasis. Rohin said “we’re running [reinforcement learning] over a week at most,” and the OpenAI post says a model’s actions “may unfold autonomously over hours, days, or even weeks...” I’m wondering if the difference here suggests OpenAI may be running RL over trajectories longer than a week. If the answer is “That’s what it sounds like,” I wonder if that would change Rohin’s “natural prediction...that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking.”