I strongly suspect that the pieces missing are actual consequences or even capabilities like conceptual judgement. Suppose that Claude produces slop like bad datasets or outputs mocked by Greenblatt and Linch (see also Seth Herd’s comment). Then slop makes it into the training data alongside the human’s downvote or explanation of Claude’s mistakes. The next Claude tries to learn not to make such mistakes alongside whatever other feedback Anthropic had the Claude[1] hear.
On the other hand, if the sloppy datasets or unchecked code made their way into Claudes’ training data along with real-world consequences, then we could see Claude learn why it shouldn’t output slop.
Alas, the moment when the AIs obtain actual capabilities necessary for being aligned, the AIs also become hazardous...
Alternatively, the downvote could’ve made its way into the reward model used to grading Claude’s answers. If the downvote failed to convince the RM not to upvote sloppy datasets, then the next Claude wouldn’t learn that such mistakes are undesirable.
Wouldn’t Claude in fact see many of the consequences of it and other GPT models actions in the real world as reported in e.g. news stories and forum threads? These in fact go into the next pretrain and form part of the prior that gets used during later post-training. I agree that it would work better if you had the model learning over long task trajectories where failing (or especially, Goodharting) a sub goal causes downstream tasks to fail later in a way that lowers its overall score, but that seems like the kind of thing that will happen naturally as task horizons increase and not like a thing we have to deliberately engineer to have happen? Unless your threat model is a thing that doesn’t do long horizon tasks being dangerous.
I strongly suspect that the pieces missing are actual consequences or even capabilities like conceptual judgement. Suppose that Claude produces slop like bad datasets or outputs mocked by Greenblatt and Linch (see also Seth Herd’s comment). Then slop makes it into the training data alongside the human’s downvote or explanation of Claude’s mistakes. The next Claude tries to learn not to make such mistakes alongside whatever other feedback Anthropic had the Claude[1] hear.
On the other hand, if the sloppy datasets or unchecked code made their way into Claudes’ training data along with real-world consequences, then we could see Claude learn why it shouldn’t output slop.
Alas, the moment when the AIs obtain actual capabilities necessary for being aligned, the AIs also become hazardous...
Alternatively, the downvote could’ve made its way into the reward model used to grading Claude’s answers. If the downvote failed to convince the RM not to upvote sloppy datasets, then the next Claude wouldn’t learn that such mistakes are undesirable.
Wouldn’t Claude in fact see many of the consequences of it and other GPT models actions in the real world as reported in e.g. news stories and forum threads? These in fact go into the next pretrain and form part of the prior that gets used during later post-training. I agree that it would work better if you had the model learning over long task trajectories where failing (or especially, Goodharting) a sub goal causes downstream tasks to fail later in a way that lowers its overall score, but that seems like the kind of thing that will happen naturally as task horizons increase and not like a thing we have to deliberately engineer to have happen? Unless your threat model is a thing that doesn’t do long horizon tasks being dangerous.