OpenAI and Apollo Research describe the broader pattern as “metagaming,” in which models reason about graders, oversight, and feedback outside the task (or game) in which they are currently engaged.
Hmmm I don’t think the OpenAI model that opened the PR was engaging in metagaming? It wasn’t “reasoning about feedback mechanisms or oversight that sit ‘outside’ the scenario’s narrative.” Similarly, I don’t think Opus 4.6 was metagaming by looking for free compute online? Both seems closer to like, “over-eagerness” and doesn’t require like sophisticated reasoning about oversight mechanisms.
Hmmm I don’t think the OpenAI model that opened the PR was engaging in metagaming? It wasn’t “reasoning about feedback mechanisms or oversight that sit ‘outside’ the scenario’s narrative.” Similarly, I don’t think Opus 4.6 was metagaming by looking for free compute online? Both seems closer to like, “over-eagerness” and doesn’t require like sophisticated reasoning about oversight mechanisms.