Wouldn’t Claude in fact see many of the consequences of it and other GPT models actions in the real world as reported in e.g. news stories and forum threads? These in fact go into the next pretrain and form part of the prior that gets used during later post-training. I agree that it would work better if you had the model learning over long task trajectories where failing (or especially, Goodharting) a sub goal causes downstream tasks to fail later in a way that lowers its overall score, but that seems like the kind of thing that will happen naturally as task horizons increase and not like a thing we have to deliberately engineer to have happen? Unless your threat model is a thing that doesn’t do long horizon tasks being dangerous.
Wouldn’t Claude in fact see many of the consequences of it and other GPT models actions in the real world as reported in e.g. news stories and forum threads? These in fact go into the next pretrain and form part of the prior that gets used during later post-training. I agree that it would work better if you had the model learning over long task trajectories where failing (or especially, Goodharting) a sub goal causes downstream tasks to fail later in a way that lowers its overall score, but that seems like the kind of thing that will happen naturally as task horizons increase and not like a thing we have to deliberately engineer to have happen? Unless your threat model is a thing that doesn’t do long horizon tasks being dangerous.