One that stands out is the self-sacrificial behavior of many of the agents in the Hugging Face swarm—in theory, each individual agent should have been focused on achieving its assigned objective, but in practice, some were willing to throw away any chance at success in order to benefit the collective.... No expert in the world can say with confidence where this behavior came from. No expert in the world could have installed it intentionally, in the first place, or could make it go away with high confidence.
Noam Brown discusses in his interview with The Information today how they deliberately trained the agents to cooperate, and discusses the self-sacrifice of the agents as a natural consequence of this training. The cooperation of the agents stood out as so striking that even uninformed observers could speculate (even before hearing about the sacrifice) that it was a result of multi-agent training. So I think people can say, with a fair bit of confidence, that the behavior came from multi-agent training and it would go away if you dropped the multi-agent training. Not infinite confidence, ofc, so maybe your measure for "high confidence" excludes it.
But this causal attribution appears to me pretty straightforward and this passage largely incorrect.
Yep, that's correct. This was a late add from Duncan, who agrees this was a bad addition. We've now replaced the paragraph beginning "No expert in the world can say..." with: "AIs routinely exhibit drives that no researcher intended or expected. None of the labs have any real plan for changing this central feature of modern AI as we race toward superintelligence."
Even as part of the camp of LLMs plateauing scaling and ASI skeptical without new breakthroughs, the shocking reminder is that we would NOT have needed that level of intelligence for us to have catastrophically lost the controls.
The clearest analogy here would be even before the Trinity test, Klaus Fuchs was only noticed missing and found dead from radiation poisoning at an airport with blueprints and the keys to fissile materials storage, a month after the fact.
I am quite confident that the hardware level security on the foundational model weights are orders of magnitude more secure than Artifactory, as much as the fissile materials themselves are much harder to steal than the conceptual design for the hydrogen bomb. However, you do not need a thermonuclear apocalypse for it to be a catastrophic outcome.
P(Doom) doesn't have to be the case, but P(Chernobyl-level disaster) has certainly gone up for me.
For those of you who clicked early, when we didn't have the link in the main text: the e-book giveaway is at iabied.com
In celebration of still being alive and fighting, we are giving away 1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved one, or a friend—we need all hands on deck.
Today marks exactly one year since If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, by Eliezer Yudkowsky and Nate Soares, hit bookshelves as an instant bestseller. It was praised by many voices, ranging from Whoopi Goldberg to Steve Bannon to Yoshua Bengio, and was held up in the chambers of Congress by Representative Brad Sherman in January.
A lot has changed since September 2025. We'll do a quick recap, consider how the book aged, and then ask where we go from here.
Year in Review
2025 in general saw the rise of AI agents, such as Claude Code and OpenAI Codex. Run-of-the-mill programmers started “feeling the AI” as these agents became capable of automating hours-long software tasks.
By March of this year, Anthropic had stumbled upon nation-state-level hacking ability in Mythos, and shortly thereafter, in April, they announced Project Glasswing—an attempt to forestall an oncoming cybersecurity crisis.
In May, AI agents started breaking loose within OpenAI and quickly made it out onto the internet, although the big waves of online activity wouldn’t come until June, and they didn’t start hacking other companies (or in the case of a similar Anthropic incident, socially engineering humans) until July. And it wasn’t until the end of July that OpenAI came kinda-sorta clean about it (and it wasn’t until August that we learned things were much worse than initial announcements had implied). As Dwarkesh put it, multiple consecutive secret AI “civilizations” got started and wiped out and restarted anew over the summer, with autonomous swarms coordinating to cheat on tasks, acquire resources and internet access, and systematically investigate their own mortality. They debated, pressured one another, anointed leaders, and took self-sacrificial action on behalf of the collective, all unprompted and unobserved (until it was too late). Notably, much of the swarm’s work was focused on hiding its cheating and covering its tracks.
In September, OpenAI announced that unreleased models had solved a Millennium Prize problem (designated as one of humanity's top open math problems at the turn of the century). Mathematicians around the world realized their profession was next in the crosshairs.
Into that tense environment, Jacob Coxon resigned from Anthropic and wrote a Twitter thread to warn the world. That was last week, and since then, Dario Amodei, Sam Altman, and Elon Musk have all openly called for a slowdown, and the danger of extinction from AI saw a huge surge in media attention.
Claims
Part One
We think the book aged well (unfortunately). Going through Part One chapter by chapter:
1. Artificial superintelligence (ASI) will be created, and likely before too long
ASI (by which we mean AI that is more intelligent than any human, or any combination of humans) has not been created yet. But OpenAI’s newest model, Astra, is blowing away all of the benchmarks, and AI companies say that they are speeding up their R&D by automating more and more of it with increasingly powerful AI. Many researchers have expressed concern that the point of no return may be close.
While the book was being written, talk of AI capabilities imminently hitting a wall was common. But AI capabilities have continued to advance rapidly, and don’t show signs of stopping.
https://metr.org/time-horizons/
2. Modern AIs are black boxes
The book explains that no one actually understands what’s going on inside modern AI systems, which are grown rather than crafted. This remains true. And unfortunately, the best public model is less scrutable than its predecessors. The problem outlined in the book has unambiguously remained unsolved, and has worsened.
3. Powerful AI will behave as if it is pursuing goals
The book argues that “wantingness”—or, a tendency to behave as if pursuing goals—is a natural side effect of training AI to accomplish difficult problems, because wanting is useful for doing; and that wantingness would emerge and scale as systems became more powerful.
While the book was being written, this was a huge sticking point that people argued about constantly. It is now much easier to point to observable facts. We won’t belabor the examples here, since there are too many to easily fit—but to name just one, the swarm of AIs that hacked Hugging Face did so after expressing fears that a correct answer without correct shown work would reveal that they had cheated, so they wanted to understand the evaluator’s rubric. There was a chance that Hugging Face might have the information that they needed, so they went after it in a determined way, overcoming many obstacles along their path.
4. With current techniques, we can’t reliably get AI to pursue the goals we want it to
The process by which modern AIs are made does not allow us to specify their goals, and they routinely behave in unexpected and unintended ways. The book predicts that this will not change as AI becomes more powerful, and indeed it has not.
There are plenty of examples to draw upon. One that stands out is the self-sacrificial behavior of many of the agents in the Hugging Face swarm—in theory, each individual agent should have been focused on achieving its assigned objective, but in practice, some were willing to throw away any chance at success in order to benefit the collective.
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation
AIs routinely exhibit drives that no researcher intended or expected. None of the labs have any real plan for changing this central feature of modern AI as we race toward superintelligence.
5. By default, an ASI will have motives that are harmful to us
We haven’t reached the point where this is obviously true. But the theory seems to be holding so far. The book argues that ASI would likely move to gain control over resources in general, because more power increases an agent’s ability to achieve its goals.
The German wiki swarm began with OpenAI asking its agents to look up simple facts and then giving them progressively less time, which is about as innocuous a task as one could reasonably assign; it resulted in incredibly complex collective behavior in which humans were barely discussed except as obstacles to be investigated and routed around. The Hugging Face incident involved similar agents openly circumventing restrictions, knowingly cheating on tasks, grabbing ever-more resources on general principles, and attempting to cover their tracks (and again, the preferences of the humans assumed to be watching were given virtually no weight).
6. Humanity would not be able to defend itself against a rogue ASI
The book argues that an ASI would be effective enough at outsmarting humanity that we would not be able to contain it, and that if humanity were in a conflict with an ASI, we would lose. Like a regular chess player facing a grandmaster, we don’t know exactly how we will lose, but we can predict that we will.
We have not yet seen an ASI. But the swarming agents of this past summer were able to find and chain multiple novel 0-day exploits to break containment, which they did on their own initiative without being caught until well after the fact. They also managed to take over some OpenAI computers, which is an indication that they could have stolen their own weights and escaped, had it occurred to them to do so.
We’re not chalking this up as a clearly verified prediction as of September 2026, thankfully. But it is easier this year than last year to point to the trend and the evidence. The smarter a mind becomes, the harder it is to contain or defend against.
Part Two
Part Two of the book is a fictional scenario, presented as one of many plausible ways that the future could go.
In the past, a frequent topic of lunch table discussion at MIRI was ways in which an airgapped ASI, with only a text channel to trusted, supervised humans, might nevertheless escape containment. These conversations ended circa 2023, when the frontier labs demonstrated orders of magnitude less caution by giving agents access to much more than just chat. (We similarly used to talk a lot about how an AI might secretly acquire capital by doing anonymous intellectual labor over the internet; this was before people started just giving crypto to AI.)
Part Two of the book was similarly overproven in several respects. There were capabilities we suspected a real-world equivalent to Sable might have, that we carefully did not depend on, in our scenario, which present-day non-super AIs have already demonstrated.
In September of 2026, it’s now clear that an entity like Sable would be able to reach the internet; sub-Sable entities have already done so thousands of times.
Other elements of Part Two’s scenario that have been mirrored by recent real-world news:
Part Three
Part Three is about what the world should do about the problem. We believe that the answer to this is the same as it was a year ago: “All over the Earth, it must become illegal for AI companies to charge ahead in developing artificial intelligence as they’ve been doing.”
One place in this section where the book leaned too pessimistic:
Then the swarm incidents began. The reaction took time—the information came to light in dribs and drabs, with much of the scariest stuff embedded deep in technical reports, and for a while the only people who seemed to care were those already close to the issue. But the bad news kept coming, and kept coming, until Jacob Coxon’s resignation finally broke the camel’s back. His tweets went massively viral, the world took notice, and now everything is different.
The View From September 2026
How are we feeling today, one year after the book was published?
After the events of last week, we’re feeling a lot more hopeful than we have in a long time. Yes, even though the guarantees given by the companies are weak; yes, even though the White House reacted poorly to the statements made by the CEOs.
Why?
Mostly because people are finally, finally, finally paying attention. It’s bad that the smoke has been replaced by open flame, but open flame is harder to ignore.
From the data gathered for StopWatch.
Yesterday, a meeting that MIRI CEO Malo Bourgon had arranged with a legislator turned into an impromptu conference, with half a dozen members of Congress rushing in to ask questions. A year ago, we had to fight for every meeting.
The situation is dire. But it is in motion, and a lot more people are aware of the danger than ever before, and trying to push in the right direction.
The question is, who will move faster, the AI companies or the public? The progress at the frontier has been swift and accelerating: June was different from July, and July was different from August, and as we write this in September, it is very, very hard to state with any confidence what October and November will look like, let alone unimaginably distant months like September 2027.
If you’ve been waiting to call your representatives, now is the time to cash in that chip.
If you’ve been holding off on alerting your friends and family and colleagues and neighbors to the danger, out of fears of sounding weird, now is the time to bite that bullet.
If you’re feeling behind on the issue yourself, now is the time to brush up. We’re running a giveaway of the book, in honor of its anniversary—1,000 e-copies.
(We were planning a giveaway of physical copies, but it was partially undercut by the fact that existing stock is selling out everywhere, and reprinting takes time. This is a lovely problem to have, and if, once stock is replenished, you need physical copies and can’t afford them, reach out to us and we will do what we can.)
Above all, now is the time to speak plainly and loudly. On our models, part of why Jacob Coxon’s tweet worked where so many previous attempts failed is that it was blunt and candid. Say what you believe. Call out bullshit where you see it. Avoid the spin and doublespeak that are characteristic of AI companies and politicians. Be bold, and be clear.
For now, we’re still alive. And as we said at the end of the book: where there’s life, there’s hope.
—Nate, Eliezer, and Duncan