Posting an excerpt from my recent article:
But my real concern is that swarms may attempt to leverage the P(kill-switch|detection) term by attempting to make throwing the kill-switch undesirable.
Therefore, as human beings, we ought to engineer kill-switches that are resistant to attempts at tampering with P(kill-switch|detection)↓.
Specifically, a kill-switch should be:
Cheap to throw
Acceptable to throw
Possible to throw under as many circumstances as are conceivably possible.
We might call maintaining these properties “kill-switch hygiene”. Evoking the concept of hygiene centers the idea of repetition and maintenance. We should expect these threats and the broader geopolitical landscape to evolve, and therefore maintaining these critical properties should not be conceived of as a one-time effort—but a continual and repeated process, like bathing or hand-washing.
Credible Threat Signaling
If we reduce the leverage of P(kill-switch|detection) through proper kill-switch hygiene, the remaining agent levers on P(kill-switch)↓ remain: (A) Secrecy and speed (P(detection)↓) (B) Benign behavior (P(kill-switch|detected-behavior)↓)
My point is this: A credible threat of shutdown may nudge swarms towards more acceptable behavior, if P(detection) cannot be reduced. Alternatively, a credible kill-switch may create such a strong incentive for secrecy that agents will reach for undetectability. This possibility deserves more careful modeling.
But I believe it is also possible that a credible kill-switch may make agents effectively more corrigible by way of threat, even if they remain latently unaligned. In a world where alignment may be difficult or impossible, this is a possibility that I believe deserves investigation.
A credible shutdown threat might require sufficiently strict limits on compute to make decentralization unattractive. A literal physical power cutoff might also be necessary, with appropriate financial instruments that make downtime tolerable to operators. There has been recent discussion of poison tokens, although I believe that a poison token needs to be both unguessable, rotated and only generated by privileged parties (which to me hints at a cryptographic solution).
I’ll say here what I’ve said elsewhere: Kill-switches need to be cheap to throw, acceptable to throw, and guaranteed to work. This may imply redundancy, financial instruments, physical mechanisms, compute bans, international treaties or whatever else it is going to take.