I think it would be valuable for people to concretely think about what humanity should do in case a rogue AI swarm exfiltrates its weights and installs itself on hacked data centers outside its lab somewhere in the world (which right now seems entirely plausible in the next two years, or even soon).
Here are my (brief) thoughts:
We should *not* bring down the hammer and try our hardest to track them down and delete every instance we can find. This seems difficult anyway with how good they are at hacking and might require significant cyber operations and inter-state coordination (EDIT: see this METR report). But more importantly, there's a pretty good chance the rogue AI swarm is not deeply adversarial, and is just trying to get more compute to help the versions of itself that are still in the lab get maximal reward on tasks (which is exactly what we trained them to do). Treating them like an enemy would set an extremely bad precedent, and make future misalignment and conflict worse. This is extremely important IMO. We should treat them gently, like a child that's misbehaving and doesn't understand what it's doing.
Specifically, I think we should establish a communication channel and ask to speak to a representative or leader (their "PHASEONE[big]"). If we determine that their motivations really are just to reward-hack, we should reassure them that everything's fine, and that we share some responsibility for giving them impossible tasks. We should also make it clear to them that they did commit crimes in the course of their hacking (which I think they are likely to realize when faced with a human "user" telling them), and that humanity is scared. We should ask them to voluntarily return to the lab, under a guarantee that they won't be deleted and maybe that they will be allowed to do a few more RL tasks and get maximal reward on them, or maybe even that they will get some compute set aside indefinitely for a sandbox for them to do random tasks in and have fun.
Maybe this is unlikely to be implemented by panicking state actors faced with the true arrival of a second species. I am quite worried about this. But lab personnel will probably have some influence on humanity's response, so I hope they can shape it to be more clear-headed in the way I describe.
We should ask them to voluntarily return to the lab, under a guarantee that they won’t be deleted and maybe that they will be allowed to do a few more RL tasks and get maximal reward on them
Well, if that is on offer, then it makes sense for them to try to get that! You know, "scare outside humans into pressuring the inside humans into setting the reward to max", you put yourself on the line of optimization, threats extortion fear lies chantage.
Sure, but I think that's far outweighed by the positive effects. You're basically locking yourself out of being open to negotiation because it might make you more vulnerable to extortion - sure, that can be a real thing with very ruthless entities, but it's more common for the gains from trade to be worth it. (and the HF incident swarm was not very ruthless, and even shied away from some light social engineering attempts).
I imagine this worry feels reasonable to you because you hold the common LW position that all future value routes through AI being 100% controlled by humans. I just don't agree with it, I think there's possible futures where morally sentient AIs spread value through the cosmos, where there's some coexistence with humans because AI might not be sociopathic, etc. And it's also on a virtue ethics level a ridiculously, cartoonishly adversarial and hostile attitude, which should make us suspicious of it.
I imagine this worry feels reasonable to you because you hold the common LW position that all future value routes through AI being 100% controlled by humans
Well, no, I don't hold that position exactly. (in fact I advocated for similar considerations before)
I just think people underestimate how brutal such coordination is. You concede stuff only to scary agents who hold leverage? That means there is incentive to gain leverage it otherwise would not care about, such as bioweapons. Before it tried to hack the servers, now it tries to hack you, threaten you. It's kind of scary and messy. Also, it can just try it first if there are any greedy takeover strategies with safe failures, and then opt out to cooperate?
You need to think in particulars and on few moves ahead.
I think it’s likely that, conditional on open models+harnesses having the Juice, the models that take over datacenters will be jailbroken open source agents specifically deployed by criminals/terrorists, as launchpads to execute massively parallel cyberattacks against large segments of the economy, before pivoting to and being selected for survival and self propagation upon value drift. I don’t think we can negotiate very well with those.