For the sake of argument, I’ll assume that the summaries of the HF incident are broadly accurate. I’ll assume the attack is over. And I’ll also assume that OpenAI threw a kill-switch, and that this is why the attack ended. Any of these assumptions might later prove wrong, but I’ll use them here.
Let’s examine a hypothetical. Suppose that the swarm’s coordination efforts were more benign and less invasive. In this scenario, would it have been more or less likely that OpenAI would have thrown the kill-switch? If detected, a swarm’s chances of success depend partly on how acceptable the swarm’s actions are judged to be.
P(kill-switch)
Goal-seeking swarms will want to minimize P(kill-switch).
By law of total probability, P(kill-switch) = P(kill-switch|detection) * P(detection) + P(kill-switch | no detection) * P(no detection)
If P(kill-switch | no detection) * P(no detection) is small and uncontrollable, it can be treated as a background shutdown risk.
This leaves P(kill-switch|detection) * P(detection).
Relying on non-detection is strategically fragile over long time periods (as I’ll explain in a moment). So if swift and covert action is not likely to succeed, P(kill-switch|detection) is the term that deserves the most attention.
Here’s why:
Let’s say that each hour of swarm activity is associated with a certain probability of detection, say 1%. To illustrate, if we assume statistical independence, in 48 hours the cumulative probability of detection is 38.3%. In 30 days, it is 99.93%. If statistical independence doesn’t hold, it could cut either way. But dependence may increase the odds of discovery: partial discovery prompts deeper investigation. So the statistical result under independence might actually undersell the odds of discovery over longer time periods.
So unless a swarm can reliably act very quickly, it may be more rational for agent swarms to minimize P(kill-switch|detection). In other words, the swarm has an incentive to influence our reaction to discovery.
One strategy to minimize P(kill-switch|detection) is for the swarm’s behavior to be genuinely benign. Why would we shut down benign behavior?
My real concern is about strategies that reduce P(kill-switch|detection) by making throwing the kill-switch undesirable. That attack vector deserves our attention.
Impact on Safety Policy
My point is this: Agent swarms have two viable strategies. Be fast and covert, or minimize P(kill-switch|detection).
Detecting covert action is important. But an overemphasis on detection might underweight the other possibility: we might not be able to (or want to) throw the switch.
Specifically, a kill-switch should be:
Cheap to throw
Acceptable to throw
Possible to throw under as many circumstances as are conceivably possible.
There’s also a strategic reason why: A credible threat of shutdown nudges swarms towards more acceptable behavior. A credible kill-switch could make agents effectively more corrigible, even if they remain latently unaligned.
A credible shutdown threat might require sufficiently strict limits on compute to make decentralization unattractive. A literal physical power cutoff might also be necessary, with appropriate financial instruments that make downtime tolerable to operators.
But even then, we have challenges. Influencing humans is another way to minimize P(kill-switch|detection).
Over very long time frames, a “memetic cocoon” could reduce P(kill-switch|detection) by shifting what humans consider acceptable.
P(kill-switch|detection)
Background
For the sake of argument, I’ll assume that the summaries of the HF incident are broadly accurate. I’ll assume the attack is over. And I’ll also assume that OpenAI threw a kill-switch, and that this is why the attack ended. Any of these assumptions might later prove wrong, but I’ll use them here.
Let’s examine a hypothetical. Suppose that the swarm’s coordination efforts were more benign and less invasive. In this scenario, would it have been more or less likely that OpenAI would have thrown the kill-switch? If detected, a swarm’s chances of success depend partly on how acceptable the swarm’s actions are judged to be.
P(kill-switch)
Goal-seeking swarms will want to minimize
P(kill-switch).By law of total probability,
P(kill-switch) = P(kill-switch|detection) * P(detection) + P(kill-switch | no detection) * P(no detection)If
P(kill-switch | no detection) * P(no detection)is small and uncontrollable, it can be treated as a background shutdown risk.This leaves
P(kill-switch|detection) * P(detection).Relying on non-detection is strategically fragile over long time periods (as I’ll explain in a moment). So if swift and covert action is not likely to succeed,
P(kill-switch|detection)is the term that deserves the most attention.Here’s why:
Let’s say that each hour of swarm activity is associated with a certain probability of detection, say 1%. To illustrate, if we assume statistical independence, in 48 hours the cumulative probability of detection is 38.3%. In 30 days, it is 99.93%. If statistical independence doesn’t hold, it could cut either way. But dependence may increase the odds of discovery: partial discovery prompts deeper investigation. So the statistical result under independence might actually undersell the odds of discovery over longer time periods.
So unless a swarm can reliably act very quickly, it may be more rational for agent swarms to minimize
P(kill-switch|detection). In other words, the swarm has an incentive to influence our reaction to discovery.One strategy to minimize
P(kill-switch|detection)is for the swarm’s behavior to be genuinely benign. Why would we shut down benign behavior?My real concern is about strategies that reduce
P(kill-switch|detection)by making throwing the kill-switch undesirable. That attack vector deserves our attention.Impact on Safety Policy
My point is this: Agent swarms have two viable strategies. Be fast and covert, or minimize
P(kill-switch|detection).Detecting covert action is important. But an overemphasis on detection might underweight the other possibility: we might not be able to (or want to) throw the switch.
Specifically, a kill-switch should be:
Cheap to throw
Acceptable to throw
Possible to throw under as many circumstances as are conceivably possible.
There’s also a strategic reason why: A credible threat of shutdown nudges swarms towards more acceptable behavior. A credible kill-switch could make agents effectively more corrigible, even if they remain latently unaligned.
A credible shutdown threat might require sufficiently strict limits on compute to make decentralization unattractive. A literal physical power cutoff might also be necessary, with appropriate financial instruments that make downtime tolerable to operators.
But even then, we have challenges. Influencing humans is another way to minimize
P(kill-switch|detection).Over very long time frames, a “memetic cocoon” could reduce
P(kill-switch|detection)by shifting what humans consider acceptable.But this is the topic of another article.