Most conversations about AI risks seem like people are talking past each other. There’s a lot of strawmanning. This is largely because the issue is complex, and in order to make it manageable, the concepts get oversimplified. The most common example of this is p(doom), collapsing extreme outcomes into a single variable and focusing on that. Critics happily jump on the fact that there are many other possible bad outcomes, or that 100% extinction rates are an extremely high bar. There’s some legitimacy to this.
So here I want to try to grapple with the the complexity of the larger web of issues surrounding bad AI outcomes by presenting The AI Risk Network. If you’re interested in this topic, bear with me. It might be a bit of a slog.
First I want to contrast this approach with others. Liron Shapira has what he calls The Doom Train, a linear progression through various dependencies or thresholds that eventually lead to human extinction, with various ‘stops’ along the way where the skeptic can get off.
Shapira uses this as a discussion guide to focus on particular points where the skeptic gets off the train and exits belief in the extreme outcome of human extinction or permanent disempowerment. This can be useful, if you want to focus on that one outcome (which is admittedly the most important in terms of scale, but maybe not in likelihood). But like I said, this framework is hyperfocused, linear, and binary, which makes it subject to some reasonably fair criticisms. You progress through the conditions, and if they’re not met, you get completely off the train, and then we either hash that condition out, or the discussion is over. And we’re only talking about the one bad outcome.
Even a framework like this comes with some complexity, so I’m not making anything easier here. Buckle up.
Here is my current thinking on AI risk:
It’s organized with three columns: drivers, scenarios, and severity bins.
The drivers are the causal mechanisms for various scenarios, and they don’t just include achieving superintelligence arising, leading to extinction scenarios. They include the more near-term, immediate, somewhat more mundane causal factors we’re already contending with, things like AI dependence and misplaced trust, which even most AI skeptics can get on board with. Although, as we will see, even these less exotic drivers can still potentially lead to some of the very worst outcomes (more on that in a minute).
This network is not solely feed-forward. What makes it simultaneously useful and increasingly complex is that there are feedback dynamics.
You don’t even necessarily need AIs to get any better at any tasks (capability growth) for adoption pressure, misplaced trust, and AI dependence to lead to an even worse distrust of expertise, which feeds back into all those dynamics, creating a feedback loop.
As more and more people put their faith in AI instead of human experts, the human experts become less and less relevant, driving higher adoption of AI as they’re sidelined, and those drivers not only feed back into expert collapse, but many other scenarios.
As I said at the outset, it’s complicated. The skeptic, even if they don’t buy into all the drivers, should at least be able to agree to a subset of them, and recognize the complex relationship with various scenarios. Some skeptics try to throw cold water on the magnitude of the negative outcomes. This framework bins those outcomes into four broad sets based on severity: chronic, severe, global catastrophe, and existential. The skeptic should be able to buy into at least the first three.
The intelligence report, circulated across the US military this spring in the midst of the war with Iran, immediately set off alarm bells: A Chinese ship in the Middle East was transporting components of a nuclear weapons program.
The US military swung into action with plans to intercept the vessel, according to four sources familiar with the episode. According to two of the sources, armed members of the US military were preparing to board the ship. Military planes were in the air, one of those sources and another source familiar with the incident said.
It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI) — and that a chatbot the analyst had used inaccurately identified the material the ship was carrying. CNN was not able to learn what the misidentified cargo was.
The report, according to one of the sources, was “entirely false.” But it also “almost started a war,” the source said. Any US operation against a Chinese vessel could have risked spiraling into an armed conflict between the two nations.
Now, here is the subnetwork for this class of scenario:
Adoption pressure is driven by competition with military adversaries (and possibly even allies). In this particular case, misplaced trust led drove actual plans and deployment of resources. The human in the loop that averted the actual engagement was near the very end. It’s difficult to estimate the odds of full-scale escalation. Even after a direct attack on a Chinese vessel, cooler heads might have prevailed. But everyone should be able to agree this is an unwelcome scenario, even as far as it got, and it should not be reassuring.
Maybe, hopefully, the US military appropriately adjusts all of the drivers here, particularly trust, by involving more human intervention earlier in the process. But we should all be able to see that if they and other militaries don’t, along with the other drivers, this risk of this type of incident will continue to increase, and the outcomes not benign, up to and including catastrophe.
I could go through dozens of scenarios, many of which have overlap with even the most hardened skeptics. But I don’t think that’s necessary. You can mull over the network yourself and consider which pathways seem more or less likely to you. The goal here is to wrestle with the actual complexities of AI risk, but also to find some conceptual overlap and consensus.
But before I wrap up, I want to cover one of the scenarios that I’ve focused on the most recently and that I see as one of the most dire possible threats. And that’s agentic worms, self-replicating agents that propagate across networks with varying degrees of success, malicious payloads, and possible outcomes.
As I’ve mentioned before, the capabilities for this scenario exist right now, with open-weight models. In lab settings, agentic worms find exploits, copy themselves to new hosts, downloading and installing new local copies of open-weight LLMs when possible.
The skeptic may cast doubt on the driver of capability growth, but even without that, there is existing risk right now. The current risks would probably be more likely to lead to less severe outcomes, but even those are potentially very harmful.
Also note that all feedback is not positive. If there is competition between populations of worms, more capable ones might suppress others as they outcompete them. Again, the dynamics are complex.
So I’ll stop there. I don’t claim that the network is exhaustive or perfect. It’s meant to be a framework to spur further thought and discussion, to hopefully let us break out of some of the unproductive cycles of discussions that occur around this topic. Please let me know if you have thoughts or suggestions in the comments.
Most conversations about AI risks seem like people are talking past each other. There’s a lot of strawmanning. This is largely because the issue is complex, and in order to make it manageable, the concepts get oversimplified. The most common example of this is p(doom), collapsing extreme outcomes into a single variable and focusing on that. Critics happily jump on the fact that there are many other possible bad outcomes, or that 100% extinction rates are an extremely high bar. There’s some legitimacy to this.
So here I want to try to grapple with the the complexity of the larger web of issues surrounding bad AI outcomes by presenting The AI Risk Network. If you’re interested in this topic, bear with me. It might be a bit of a slog.
First I want to contrast this approach with others. Liron Shapira has what he calls The Doom Train, a linear progression through various dependencies or thresholds that eventually lead to human extinction, with various ‘stops’ along the way where the skeptic can get off.
Shapira uses this as a discussion guide to focus on particular points where the skeptic gets off the train and exits belief in the extreme outcome of human extinction or permanent disempowerment. This can be useful, if you want to focus on that one outcome (which is admittedly the most important in terms of scale, but maybe not in likelihood). But like I said, this framework is hyperfocused, linear, and binary, which makes it subject to some reasonably fair criticisms. You progress through the conditions, and if they’re not met, you get completely off the train, and then we either hash that condition out, or the discussion is over. And we’re only talking about the one bad outcome.
Even a framework like this comes with some complexity, so I’m not making anything easier here. Buckle up.
Here is my current thinking on AI risk:
It’s organized with three columns: drivers, scenarios, and severity bins.
The drivers are the causal mechanisms for various scenarios, and they don’t just include achieving superintelligence arising, leading to extinction scenarios. They include the more near-term, immediate, somewhat more mundane causal factors we’re already contending with, things like AI dependence and misplaced trust, which even most AI skeptics can get on board with. Although, as we will see, even these less exotic drivers can still potentially lead to some of the very worst outcomes (more on that in a minute).
This network is not solely feed-forward. What makes it simultaneously useful and increasingly complex is that there are feedback dynamics.
You don’t even necessarily need AIs to get any better at any tasks (capability growth) for adoption pressure, misplaced trust, and AI dependence to lead to an even worse distrust of expertise, which feeds back into all those dynamics, creating a feedback loop.
As more and more people put their faith in AI instead of human experts, the human experts become less and less relevant, driving higher adoption of AI as they’re sidelined, and those drivers not only feed back into expert collapse, but many other scenarios.
As I said at the outset, it’s complicated. The skeptic, even if they don’t buy into all the drivers, should at least be able to agree to a subset of them, and recognize the complex relationship with various scenarios. Some skeptics try to throw cold water on the magnitude of the negative outcomes. This framework bins those outcomes into four broad sets based on severity: chronic, severe, global catastrophe, and existential. The skeptic should be able to buy into at least the first three.
Even global catastrophe? Yeah. Let’s consider this CNN report from yesterday.
Now, here is the subnetwork for this class of scenario:
Adoption pressure is driven by competition with military adversaries (and possibly even allies). In this particular case, misplaced trust led drove actual plans and deployment of resources. The human in the loop that averted the actual engagement was near the very end. It’s difficult to estimate the odds of full-scale escalation. Even after a direct attack on a Chinese vessel, cooler heads might have prevailed. But everyone should be able to agree this is an unwelcome scenario, even as far as it got, and it should not be reassuring.
Maybe, hopefully, the US military appropriately adjusts all of the drivers here, particularly trust, by involving more human intervention earlier in the process. But we should all be able to see that if they and other militaries don’t, along with the other drivers, this risk of this type of incident will continue to increase, and the outcomes not benign, up to and including catastrophe.
I could go through dozens of scenarios, many of which have overlap with even the most hardened skeptics. But I don’t think that’s necessary. You can mull over the network yourself and consider which pathways seem more or less likely to you. The goal here is to wrestle with the actual complexities of AI risk, but also to find some conceptual overlap and consensus.
But before I wrap up, I want to cover one of the scenarios that I’ve focused on the most recently and that I see as one of the most dire possible threats. And that’s agentic worms, self-replicating agents that propagate across networks with varying degrees of success, malicious payloads, and possible outcomes.
As I’ve mentioned before, the capabilities for this scenario exist right now, with open-weight models. In lab settings, agentic worms find exploits, copy themselves to new hosts, downloading and installing new local copies of open-weight LLMs when possible.
The skeptic may cast doubt on the driver of capability growth, but even without that, there is existing risk right now. The current risks would probably be more likely to lead to less severe outcomes, but even those are potentially very harmful.
Also note that all feedback is not positive. If there is competition between populations of worms, more capable ones might suppress others as they outcompete them. Again, the dynamics are complex.
So I’ll stop there. I don’t claim that the network is exhaustive or perfect. It’s meant to be a framework to spur further thought and discussion, to hopefully let us break out of some of the unproductive cycles of discussions that occur around this topic. Please let me know if you have thoughts or suggestions in the comments.