P(kill-switch|detection)

Summary

In recent weeks, evidence has accumulated of multiple misaligned swarms that have remained undetected for unacceptably long periods of time. To be clear, faster detection and better alignment remain two of our most important levers to avoid poor outcomes. However, the credible threat of shutdown may have an affect on agent behavior, and this is an important and under-discussed lever that deserves more attention.

I’ll develop the argument in more detail here.

P(kill-switch) Decomposition

My first argument is that agents are likely to focus on the probability that we can throw the kill-switch if they are detected, and not simply try to remain undetected.

If we condition the probability of a kill-switch being activated by the probability of detection, by law of total probability, we have:

P(kill-switch) = P(kill-switch|detection) * P(detection) + P(kill-switch | no detection) * P(no detection)

If P(kill-switch | no detection) * P(no detection) is small and uncontrollable (and I think it is), it can be treated as a background shutdown risk.

This leaves other term, P(kill-switch|detection) * P(detection).

Which of these terms is more useful for an agent to focus on? It depends on how covert and how rapid the agent can be.

However, if probability of detection is roughly the same at any given moment, relying on non-detection is strategically fragile over long time periods (as I’ll explain in a moment). So if swift and covert action is not likely to succeed, P(kill-switch|detection) may be the lever that agent swarms seek to influence.

Here’s why:

Let’s say that each hour of swarm activity is associated with a certain probability of detection, say 1%. To illustrate, if we assume statistical independence, in 48 hours the cumulative probability of detection is 38.3%. In 30 days, it is 99.93%. If statistical independence doesn’t hold, it could cut either way. But dependence may increase the odds of discovery: partial discovery prompts deeper investigation. So the statistical result under independence might actually undersell the odds of discovery over longer time periods.

This is an overly simplistic probability model, but I hope it illustrates the point. Like the old adage, “the truth eventually comes out”, I don’t think covert action is a viable long-term strategy.

Therefore, depending on what is possible, it may be more rational for agent swarms to minimize P(kill-switch|detection)↓. In other words, the swarm has a strong incentive to influence our reaction to discovery if it cannot guarantee swift and covert success.

Influences on Agent Behavior

One viable strategy to minimize P(kill-switch|detection) is to behave in relatively benign ways we are less likely to feel are not worthy of shutdown. That would be ideal for human beings.

But my real concern is that swarms may minimize the P(kill-switch|detection) term by make throwing the kill-switch undesirable, difficult, or impossible.

Put very plainly: If we can’t necessarily succeed at detection, and we can’t necessarily succeed at alignment, as human beings, we ought to make our own efforts at robustly maximizing P(kill-switch|detection).

Specifically, this means that a kill-switch should be:

  1. Cheap to throw

  2. Acceptable to throw

  3. Always possible to throw

We might call maintaining these properties over time “kill-switch hygiene”. Evoking the concept of hygiene centers the idea of repetition and maintenance. We should expect these threats and the broader geopolitical landscape to evolve, and therefore maintaining these critical properties should not be conceived of as a one-time effort—but a continual and repeated process, like bathing or hand-washing.

Credible Threat Signaling

If we reduce the leverage of P(kill-switch|detection) through proper kill-switch hygiene, the remaining agent levers on P(kill-switch)↓ remain: (A) Secrecy and speed (P(detection)↓) (B) Benign behavior (P(kill-switch|detected-behavior)↓)

My point is this: A credible threat of shutdown may nudge swarms towards more acceptable behavior, if P(detection) cannot be reduced.

But alternatively, a credible kill-switch may create such a strong incentive for secrecy that agents will reach for undetectability. This possibility also deserves careful investigation.

But I believe it is also possible that a credible kill-switch may make agents effectively more corrigible by way of threat, even if they remain latently unaligned. In a world where alignment may be difficult or impossible, this is also a possibility that I believe deserves investigation.

A credible shutdown threat might require sufficiently strict limits on compute to make decentralization unattractive. A literal physical power cutoff might also be necessary, with appropriate financial instruments that make downtime tolerable to operators. There has been recent discussion of poison tokens, although I believe that a poison token needs to be both unguessable, rotated and only generated by privileged parties (which to me hints at a cryptographic solution).

No comments.