New OpenAI post: Can midtraining on docs about aligned AI bake in alignment priors for agents? We report an experiment where those priors are quickly washed away by RL and fail to generalize to agentic settings. But that cuts both ways: priors that AIs are misaligned fade too!
https://x.com/tomekkorbak/status/2038704753887379891
https://alignment.openai.com/how-far-does-alignment-midtraining-generalize/