The reason I care about my notion of adversarial is that it is the notion that is relevant for using models to monitor other models.
If I, as a professor, had two twin brothers in my course, it would be highly imprudent to ask one of them to proctor the exam of the other. Yet we regularly have one copy of a model monitor another copy.
The reason it works, and the two models do not collude with one another is not because they are perfect and don’t make any mistake, and not because they don’t have never had a misaligned thought.
The reason is works is because current models basically do the task that was given to them. If one copy of the model was given the prompt “Do X” and another copy of the model was given the prompt “Report to me how well the other model did X” then they would each do this.
Yes, they can make mistakes, but as I write, the lack of collusion means that it is possible to amplify. For example, if I want to be really sure the model didn’t do anything bad I can run multiple copies of the monitor and can also vary some of their prompts.
All of this would not make sense if models had their own agenda independent of the prompt. If we were in that world, I could have 100 monitors watching an agent, and all 101 of them would collude together to decide what to report.
It also means that by different prompting and context, I am not simulating a single “super employee” but a collection of many employees that can also check on one another.
This could change in two ways:
1. We could start to observe models behave in ways that can only be explained by models having their own persistent long term goals that are the same goals whether they are being used as an actor, monitor, or anything else.
2. We could radically change how we deploy models by deploying them as a single full context entity that everyone talks to, rather than a large collection of agents with restricted contexts.
I think it is important to do evaluations as well as monitoring in both training and deployment to watch out for 1. And I think it will be a bad idea to deploy models in the mode of 2.
As METR notes, the fact that such incidents are detected is a positive sign about our safety practices. Generally I am not arguing that misalignment never exists. But we can still rely on monitors to detect it precisely because we don’t have widespread collusion. For example, the whole reason chain of thought monitorability is useful for safety is because we can trust monitors to read the chain of thought and report what they see in it.
The reason I care about my notion of adversarial is that it is the notion that is relevant for using models to monitor other models.
If I, as a professor, had two twin brothers in my course, it would be highly imprudent to ask one of them to proctor the exam of the other. Yet we regularly have one copy of a model monitor another copy.
The reason it works, and the two models do not collude with one another is not because they are perfect and don’t make any mistake, and not because they don’t have never had a misaligned thought.
The reason is works is because current models basically do the task that was given to them. If one copy of the model was given the prompt “Do X” and another copy of the model was given the prompt “Report to me how well the other model did X” then they would each do this.
Yes, they can make mistakes, but as I write, the lack of collusion means that it is possible to amplify. For example, if I want to be really sure the model didn’t do anything bad I can run multiple copies of the monitor and can also vary some of their prompts.
All of this would not make sense if models had their own agenda independent of the prompt. If we were in that world, I could have 100 monitors watching an agent, and all 101 of them would collude together to decide what to report.
It also means that by different prompting and context, I am not simulating a single “super employee” but a collection of many employees that can also check on one another.
This could change in two ways:
1. We could start to observe models behave in ways that can only be explained by models having their own persistent long term goals that are the same goals whether they are being used as an actor, monitor, or anything else.
2. We could radically change how we deploy models by deploying them as a single full context entity that everyone talks to, rather than a large collection of agents with restricted contexts.
I think it is important to do evaluations as well as monitoring in both training and deployment to watch out for 1. And I think it will be a bad idea to deploy models in the mode of 2.
Don’t we have METR’s report where the incidents reported by OpenAI include attempts to instruct another instance to conceal evidence of misalignment?
As METR notes, the fact that such incidents are detected is a positive sign about our safety practices. Generally I am not arguing that misalignment never exists. But we can still rely on monitors to detect it precisely because we don’t have widespread collusion. For example, the whole reason chain of thought monitorability is useful for safety is because we can trust monitors to read the chain of thought and report what they see in it.