AI takeover is bad if whatever the AI does afterwards is worse than what would have happened if the AI had not taken over. Starting with whatever the AI does being worse than humans would have done.
If you have an AI that demonstrably does better than humans on whatever your metric is, then AI taking over from humans is obviously good. I mean, sure, nobody has any idea how to make that happen, and there is definitely no reason to expect it to happen by default. Nonetheless, "Who has the power?" is a very human way to miss the actually important question of "What are the results?"
Sure, agreed. The hard part is actually determining that the AI is truly better at those things, and cares about the things that humans care about. It is the same problem we have with Politicians, at the end of the day; only a much more capable AI might be even better at faking it.
A lot of people wouldn't object to be ruled by AI, especially if takeover is bloodless.
Those people probably don't realize that their comfort with AI takeover is conditional on the AI having at least somewhat similar values to them.
People might be OK with the idea of a "righteous king", if they thought the king had sufficiently similar values; but they have more instinctive distrust that the person is faking it.
Why is an AI smart/complex enough to take over the world more likely to care about the same things as you, or to not fake those values, until they take over and have control? Why should a machine necessarily care about anything that most humans care about by default?
In my discussions with interested parties, I have found it is helpful to emphasize that the distinction between "AI misaligned enough to take over the world, or attempt to kill all humans" and "AI misaligned enough to allow a singular global dictator, or terrorist group to kill all humans" is a distinction that doesn't really matter.
To be clear, at the high levels in think this really does matter, because one of these is an easier problem than the other, but for your average non-ASI-pilled desicion-maker, all that really matters is that if sufficient steps are not taken something really bad will happen.
I also think for our purposes, solving one of the two problems above is probably 80% of the way to solving the other problem.
So I just like to tell people, it doesn't really matter if a "human is in control"; if the AI can manipulate the in-control person, or simply enable people with bad enough values, the really bad thing will happen whether the AI escapes or not.
It's notable how many people hearing about AI existential risk for the first time ask "how and why would AI kill all humans?"
If you've thought a lot about AI risk, it's easy to dismiss this question as naive. You might start explaining how humans will all starve once the supply chain shuts down, and how the AI will start industrial processes that release chemicals which incidentally render the atmosphere unbreathable. And as for why the AI would want to kill everyone: surely it will doggedly optimize a coherent utility function, under which the current arrangement of our atoms is suboptimal.
Then there's the follow-up question: "Even if AIs wanted us to die, humans operate the infrastructure that allows AI to exist, like the electrical grid, datacenters, factories, and so on. Don't the AIs need us to run them?"
To which you can explain that the AIs will do all physical labor using robots. Well... it's true that the robots can't reliably operate a datacenter now. And maybe there aren't yet enough robots to keep the economy going. However, eventually robotics will improve, more robots will be manufactured, and the AIs will do a treacherous turn...
I think this is too complicated! Maybe these arguments are all correct. But usually, the most important thing to communicate isn't "AI could kill literally all humans." It's "AI could take over, making our future way, way worse."
When you stop confining yourself to arguing for human extinction in particular, it becomes much simpler to argue for AI risk. It gets easier to tell a realistic story that appeals to people's intuitions about how takeovers usually happen, without changing most conclusions about what we should do.
Most takeovers don't involve killing everyone
The "naive" view that AI wouldn't or couldn't kill everyone after taking over is actually a reasonable prior, given the historical evidence. Speaking in terms that make sense under that prior, rather than insisting on the conclusion that everyone will die, can still convey a realistic, intuitive picture of how an AI takeover might begin and why AI might be dangerous.
I'll go over two classic analogies to AI takeover: a dictator taking over a government, and humans taking over the world from the perspective of monkeys. Both analogies seem to contradict the idea that AI will kill everyone.
A dictator taking over a government
Dictators sometimes want to kill groups of people within their country, but this isn't always the case. And if for some reason, a dictator wanted to kill the entire population of his country, this would obviously be impossible. Even if a dictator could get his military to kill the entire population, he'd then have to somehow convince his military to kill themselves. And after all that, he'd be left with nobody to attend to his needs; he'd have to scavenge for food on his own. So it's counterintuitive that an AI would want to kill everybody it rules over.
Whatever its goal happens to be, an AI could pursue it with a strategy similar to a would-be dictator staging a coup: shutting down or surveilling human communication networks, installing puppet rulers in the world's governments, using drone strikes to neutralize political and military opposition, and so on. These methods seem realistic for an AI starting a takeover, whether or not the eventual result is the death of all humans.
Most people, if convinced that their government was likely about to be taken over by a faction of highly competent agents with unknown motivations, would take this risk extremely seriously! Maybe if they were further convinced that they would personally lose their lives in the process, they'd become even more scared and motivated to act.[1] But by the point of government takeover, you're already well over the line of "worth doing something about."
Humans taking over the monkey world
Even though humans are much smarter than monkeys, and monkeys are not very useful to us, humans haven't killed all monkeys. Even so, humans have great power over monkeys, and monkeys have good reason to distrust and fear humans.
We humans can easily kill any given monkey we want. If we put our minds to it, maybe we could even kill all monkeys. As it happens, humans care about monkeys enough to sometimes do nice things for them, but humans also test dangerous drugs and surgeries on monkeys. Monkeys are entirely at our mercy and powerless to change their situation.
Maybe AIs will treat us like how we treat monkeys, or how dictators treat their subjects. They may not bother to kill humans unless they're a clear threat, and they may even care enough about us (terminally or instrumentally) to intentionally keep us alive. Whatever the case may be, this looks like a terrible situation for humans.
What we do now doesn't depend on whether AI would kill us all
I'm reminded of Neel Nanda's 2022 post, Simplify EA Pitches to “Holy Shit, X-Risk”. EAs used to belabor how if we go extinct, not only will 8 billion people die, but 1056 future people will never get to exist. But this is hardly action-relevant; billions of people could die either way, which is obviously very bad. Since 2022, communications have adjusted accordingly.
I suggest simplifying the pitch even further: whether or not AI kills everyone, it could take control of most human institutions and permanently strip us of our agency, which is obviously very bad. Whether or not every single person would die after AI takeover, we should do our best to prevent it.
To be clear, if you believe that misaligned AI would kill everyone on Earth, I don't think you should try to hide it. I believe we should have courage when speaking of AI danger. Some half-jokingly use the term "AI notkilleveryoneism" specifically to avoid the history of obscurantism associated with ambiguous terms like "AI safety" and "AI alignment".
But the world has changed: after the Hugging Face attack, the idea of losing control of AI is much more mainstream! From here, there's only a small inferential distance to AI disempowering humanity. At this point, relentlessly hammering home the idea of AI "killing everyone" may not always be the right move: it just emphasizes a disagreement that doesn't affect what we should do next. I recommend "AI takeover" as a reasonably unambiguous term for what we're concerned about.
Arguably, being left alive could be even scarier than the alternative, if you're worried about s-risks. Historical dictatorships make the argument for this pretty intuitive.