One explanation as for why models aren’t coherently dishonest is that there have been specific interventions to prevent this. An example of this is inoculation prompting, which prevents reward hacking behavior from turning into emergent misalignment. AFAIK this technique was actually used in the training of recent Claude models.
One explanation as for why models aren’t coherently dishonest is that there have been specific interventions to prevent this. An example of this is inoculation prompting, which prevents reward hacking behavior from turning into emergent misalignment. AFAIK this technique was actually used in the training of recent Claude models.