Control (AI)

TagLast edit: 9 Mar 2024 0:47 UTC by Lauro Langosco

AI Control in the context of AI Alignment is a category of plans that aim to ensure safety and benefit from AI systems, even if they are goal-directed and are actively trying to subvert your control measures. From The case for ensuring that powerful AIs are controlled:

In this post, we argue that AI labs should ensure that powerful AIs are controlled. That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures. We think no fundamental research breakthroughs are required for labs to implement safety measures that meet our standard for AI control for early transformatively useful AIs; we think that meeting our standard would substantially reduce the risks posed by intentional subversion.

There are two main lines of defense you could employ to prevent schemers from causing catastrophes.
Alignment: Ensure that your models aren’t scheming.^[2]
Control: Ensure that even if your models are scheming, you’ll be safe, because they are not capable of subverting your safety measures.^[3]