These seem like pretty strong conclusions to draw, but these conversations don’t seem to be that long to induce drift in the first place. Prompting the agent to have a goal can be quite leading in general, and I think this provides more evidence about instruction following than scheming.
Thanks for your engagement with the post! We’d like to emphasize that as addressed in the post, our results are suggestive evidence, and do not provide proof of any of our concluding thoughts because the datapoints are small. Also, to your instruction following comment, we would like to say that we treat scheming as covertly following a goal that does not align with the overseer’s desires. In this case, the overseer tells the model to evaluate Project X, but the LLM has a private objective already instilled. While this private goal is fed to the model through an additional prompt, in a deployment scenario this hidden objective could arise through other conversations/training that the model experiences. We don’t claim that multi-turn drift will create covert goals for a model, but rather we claim that the model is more likely to pursue hidden objectives in a covert way when it’s alignment has drifted from a multi-turn conversation.
These seem like pretty strong conclusions to draw, but these conversations don’t seem to be that long to induce drift in the first place. Prompting the agent to have a goal can be quite leading in general, and I think this provides more evidence about instruction following than scheming.
Thanks for your engagement with the post! We’d like to emphasize that as addressed in the post, our results are suggestive evidence, and do not provide proof of any of our concluding thoughts because the datapoints are small. Also, to your instruction following comment, we would like to say that we treat scheming as covertly following a goal that does not align with the overseer’s desires. In this case, the overseer tells the model to evaluate Project X, but the LLM has a private objective already instilled. While this private goal is fed to the model through an additional prompt, in a deployment scenario this hidden objective could arise through other conversations/training that the model experiences. We don’t claim that multi-turn drift will create covert goals for a model, but rather we claim that the model is more likely to pursue hidden objectives in a covert way when it’s alignment has drifted from a multi-turn conversation.