As METR notes, the fact that such incidents are detected is a positive sign about our safety practices. Generally I am not arguing that misalignment never exists. But we can still rely on monitors to detect it precisely because we don’t have widespread collusion. For example, the whole reason chain of thought monitorability is useful for safety is because we can trust monitors to read the chain of thought and report what they see in it.
Don’t we have METR’s report where the incidents reported by OpenAI include attempts to instruct another instance to conceal evidence of misalignment?
As METR notes, the fact that such incidents are detected is a positive sign about our safety practices. Generally I am not arguing that misalignment never exists. But we can still rely on monitors to detect it precisely because we don’t have widespread collusion. For example, the whole reason chain of thought monitorability is useful for safety is because we can trust monitors to read the chain of thought and report what they see in it.