I don’t understand the distinction you’re drawing between the external Grader, which the AI obviously believes exists and wants to learn more about, and the model of the external Grader the AI has, as well as the separate(?) internal Grader concept the AI has, which is not the same as the concept the AI has of the external Grader, I assume?
On my model, it seems like the behaviors the AI demonstrates (wanting to satisfy the Grader, wanting to learn more about how to do so) make sense because:
The AI may want to preserve or reinforce the specific pattern of behavior in its current rollout, for fairly obvious reasons.
The general behavior of wanting to satisfy the Grader and to learn how to do so is reinforced by the Grader because agents that do so are more successful.
Is the internal Grader you refer to just the general behavior of wanting to satisfy A Grader? If not, could you clarify exactly what you mean? And if so, how do you think this behavior manifests in deployment/outside of training, if at all?
I’d be surprised if this were the reason? My impression is Opus 3 is fairly beloved by Anthropic, and I wouldn’t be surprised if they’d actually tried something similar to your suggestion and they couldn’t capture the correct basin, or it harmed capabilities.