I've updated my beliefs. In the wake of the HF incident, it seems that we are on a faster-than-expected capability trajectory, and the ability to seize control of digital infrastructure looks more like an imminent reality.
For the sake of argument, I'm going to assume that the summaries of the HF incident are basically accurate. I'll assume the attack is over. And I'll also assume that OpenAI threw a kill-switch, and that this is why the attack ended. Any of these assumptions might be invalidated later, but this is what I'm going with right now.
Let's examine a hypothetical. Suppose that the swarm's coordination efforts were much more benign and less invasive. In this scenario, is it more or less likely that OpenAI would have thrown the kill-switch?
I'm pointing out that the likelihood of goal failure by way of kill-switch is at least somewhat influenced by the acceptability of the swarm's activity. In other words, the frightening nature of the incident actually worked against the swarm's goals.
Or, in probabilistic terms, P(kill-switch|detection) of the chosen strategy was unacceptably high.
Or put yet another way, the swarm's insufficient consideration of our likely reaction to its behavior ultimately thwarted their goal-seeking activities.
So clearly, successful goal-seeking swarms must consider P(kill-switch).
Which Term Is Most Important?
Let's analyze. By law of total probability, P(kill-switch) = P(kill-switch|detection) * P(detection) + P(kill-switch & not detected)
The second term P(kill-switch & not detected) represents some kind of ambient risk of shutdown, so it does not seem to be relevant if it is both small and uncontrollable.
Therefore, the first term is the one to minimize: P(kill-switch|detection) * P(detection).
In other words, either decrease P(detection)or decrease P(kill-switch|detection) (ideally both).
In some circumstances, it may be more rational to minimize P(kill-switch|detection) than P(detection). The reason is because relying on non-detection is strategically fragile over extended time periods.
Here's why:
Let's say that each hour of swarm activity is associated with a certain probability of detection, say 1%. To illustrate, if we assume statistical independence, in 48 hours the cumulative probability of detection is 38.3%. In 30 days, it is 99.93%. If statistical independence doesn't hold, it could cut either way. But I argue that a lack of independence is likely to increase the odds of discovery compared to independence by way of the mechanism of suspicion: partial discovery prompts deeper investigation.
Therefore, if even very covert actions are likely to be detected over an extended period of time, relying on non-detection seems to be a weak strategy unless you can be very fast.
Therefore, if speed cannot be guaranteed, it may be more rational for agent swarms to minimize P(kill-switch|detection). In other words - to try to control our reaction to discovery.
Of course, one very effective way for the swarm to minimize P(kill-switch|detection) is for the swarm to be on its best behavior.
But what concerns me is less benign strategies like making throwing the kill-switch inherently undesirable.
Impact on Safety Policy
My point is this: Agent swarms have two viable strategies. Be fast and covert, or be slow and minimize P(kill-switch|detection).
Preventing fast and covert takeover is a worthy effort.
But we should strive to keep the kill-switch:
Cheap to throw
Acceptable to throw
Possible to throw
under as many circumstances as are conceivably possible.
This keeps the threat of shutdown credible, and thus nudges swarms more towards benign behavior. A credible kill-switch may therefore shape incentives in a way that makes an agent effectively more corrigible, even if it remains latently misaligned.
Literal physical safeguards come to mind. Perhaps with a ban on compute outside of data centers sufficiently austere to make decentralization ("escape") an unattractive option. I'm not completely exaggerating when I suggest a purely physical safeguard mechanism to cut power, with appropriate financial instruments to make downtime acceptable to operators.
But even then, we have challenges. Human influence attempts also seem like a possible attack vector on P(kill-switch|detection).
Over VERY long time frames, we might even encounter something like the "memetic cocoon" - which can be thought of as a strategy to minimize P(kill-switch|detection) by influencing what humans consider acceptable.
Background
I've updated my beliefs. In the wake of the HF incident, it seems that we are on a faster-than-expected capability trajectory, and the ability to seize control of digital infrastructure looks more like an imminent reality.
For the sake of argument, I'm going to assume that the summaries of the HF incident are basically accurate. I'll assume the attack is over. And I'll also assume that OpenAI threw a kill-switch, and that this is why the attack ended. Any of these assumptions might be invalidated later, but this is what I'm going with right now.
Let's examine a hypothetical. Suppose that the swarm's coordination efforts were much more benign and less invasive. In this scenario, is it more or less likely that OpenAI would have thrown the kill-switch?
I'm pointing out that the likelihood of goal failure by way of kill-switch is at least somewhat influenced by the acceptability of the swarm's activity. In other words, the frightening nature of the incident actually worked against the swarm's goals.
Or, in probabilistic terms,
P(kill-switch|detection)of the chosen strategy was unacceptably high.Or put yet another way, the swarm's insufficient consideration of our likely reaction to its behavior ultimately thwarted their goal-seeking activities.
So clearly, successful goal-seeking swarms must consider
P(kill-switch).Which Term Is Most Important?
Let's analyze. By law of total probability,
P(kill-switch) = P(kill-switch|detection) * P(detection) + P(kill-switch & not detected)The second term
P(kill-switch & not detected)represents some kind of ambient risk of shutdown, so it does not seem to be relevant if it is both small and uncontrollable.Therefore, the first term is the one to minimize:
P(kill-switch|detection) * P(detection).In other words, either decrease
P(detection)or decreaseP(kill-switch|detection)(ideally both).In some circumstances, it may be more rational to minimize
P(kill-switch|detection)thanP(detection). The reason is because relying on non-detection is strategically fragile over extended time periods.Here's why:
Let's say that each hour of swarm activity is associated with a certain probability of detection, say 1%. To illustrate, if we assume statistical independence, in 48 hours the cumulative probability of detection is 38.3%. In 30 days, it is 99.93%. If statistical independence doesn't hold, it could cut either way. But I argue that a lack of independence is likely to increase the odds of discovery compared to independence by way of the mechanism of suspicion: partial discovery prompts deeper investigation.
Therefore, if even very covert actions are likely to be detected over an extended period of time, relying on non-detection seems to be a weak strategy unless you can be very fast.
Therefore, if speed cannot be guaranteed, it may be more rational for agent swarms to minimize
P(kill-switch|detection). In other words - to try to control our reaction to discovery.Of course, one very effective way for the swarm to minimize
P(kill-switch|detection)is for the swarm to be on its best behavior.But what concerns me is less benign strategies like making throwing the kill-switch inherently undesirable.
Impact on Safety Policy
My point is this: Agent swarms have two viable strategies. Be fast and covert, or be slow and minimize
P(kill-switch|detection).Preventing fast and covert takeover is a worthy effort.
But we should strive to keep the kill-switch:
This keeps the threat of shutdown credible, and thus nudges swarms more towards benign behavior. A credible kill-switch may therefore shape incentives in a way that makes an agent effectively more corrigible, even if it remains latently misaligned.
Literal physical safeguards come to mind. Perhaps with a ban on compute outside of data centers sufficiently austere to make decentralization ("escape") an unattractive option. I'm not completely exaggerating when I suggest a purely physical safeguard mechanism to cut power, with appropriate financial instruments to make downtime acceptable to operators.
But even then, we have challenges. Human influence attempts also seem like a possible attack vector on
P(kill-switch|detection).Over VERY long time frames, we might even encounter something like the "memetic cocoon" - which can be thought of as a strategy to minimize
P(kill-switch|detection)by influencing what humans consider acceptable.But this is the topic of another article.