“You do not rise to the level of your goals. You fall to the level of your systems.”
— James Clear, Atomic Habits
“We cannot improve results ‘directly.’ It can only be achieved by improving those processes [...] that produce results.”
— Orest Fiume
The much discussed “P(doom)” question asks about the likelihoods we’d ascribe to the most lethal possible futures of AI. What happens when we invert this question? It transforms into the presumably less urgent, though I’d say still important, form: What are the most likely lethal AI behaviors? Not the destroy-the-entire-world sort of behaviors, but an-AI-has-screwed-up and now a-Boeing-737-has-fallen-out-of-the-air sort of scenarios.
(If I’m ever conscripted into a war for survival against Terminators, I can imagine my hypothetical soul resting easy having fought the good fight. If I ever die in a mundane accident thanks to a stupid LLM hallucination, on the other hand, I’m going to leave behind a beleaguered ghost pedantically and disproportionately vexed to an extent never yet witnessed.)
My answer to the above question is twofold.
If we’re talking about the near future, I think the most likely lethal AI behavior is something I’ve started to call “output optimization” or “output without process” (and a highly related “input over-indexing”). This is the concern I’d like to signal-boost with this post.
However—since the lethality of output optimization is at the moment entirely theoretical—it’s probably worth first answering, “Do LLMs already have a kill count?”
Existing LLM lethality
Despite its transformative impact on software and math and its rapid adoption by consumers and businesses alike, the measurable impact of LLMs on mortality has thus far been literally less than microscopic: Almost a fifth of the world’s population is already using LLMs, about 1.5 billion, but it’s hard to come up with more than a handful of cases of lives definitively lost (or saved) due to LLMs, and even llmdeathcount.com only lists a couple hundred.
Speaking of which: llmdeathcount.com is a website that exists. I can understand the grief and rage that likely fueled the creation of this site, and as a P(doom ≈ 10 to 30%) doomer myself, I can also respect the anti-AI hustle. However, I find the site itself to be providing highly dubious value.
For example: The most recent case listed on the site regards the premeditated murders committed by Hisham Abugharbieh, who learned from ChatGPT how to dispose of bodies. I hardly consider ChatGPT responsible for these murders, nor do I think it’s likely they would have been avoided even in the absence of ChatGPT’s assistance.
On the other hand, there are cases like the 56-year-old Stein-Erik Soelberg, who in 2025 murdered himself and his 83-year-old mother, Suzanne Adams, after having had his paranoid delusions amplified by ChatGPT. I’m very willing to count suicides like these against LLMs in the cosmic moral ledger.
The overall kill count likely ranges from one to two orders of magnitude, the majority of them suicides tragically facilitated by LLMs. This probability equates to a net negative for LLMs’ impact on human years of life, though to be sure of this, we’d need to also count up how many lives LLMs have saved.
For this side, there’s fewer definitive examples I can find, but there are two cases in particular I’d like to highlight. One regards Diana Hurtado, who suffered a sudden hemorrhagic stroke while in her car. Her arm went numb and her face began to droop; she asked ChatGPT about her symptoms, and the LLM told her to call 911.
They told me, ‘If you had taken longer to call, you could have passed out in the car, bled out and died.’ So that saved my life.
I’m happy to count this one in the black for LLMs.
My second example comes from ACX: Reed Housman would not exist if not for LLMs. His parents struggled with infertility for six years before ChatGPT finally identified the one possibility underlooked by all the doctors. I’m not sure it’d be appropriate for me to re-share another person’s family photos on my blog, so let me suggest that if you’d like to see an adorable baby photo of little Reed, go check out the link.
The future risk of output optimization
Varig Flight 254 was a domestic flight from São Paulo to Belém, Brazil. The flight mistakenly veered deep into the Amazon, failed to reach an alternative airport, ran out of fuel, and in crashing had twelve passengers die.
The reason they headed into the Amazon was a human-computer input/output error, as the flight plan read “0270”, which Captain Garcez interpreted to mean 270° (due West) instead of 27.0° (north-northeast). This was not the result of the captain’s incompetence or inexperience, but of vacation: Garcez hadn’t been present when the flight plan format changed.
LLMs are great at formatting, most of the time. I would trust an LLM to do any of the following:
Take a list of degrees like “27.0°” and convert them into the format “0270”
Take a list of strings like “0270” and convert them into degrees
Take a list of directions like “W” or “NNE” and convert them into degrees
But here’s where I wouldn’t trust an LLM:
Nestled inside a many-step process, use the old format in one step and the new format in another step.
Say we’ve got flight plan info that was written the old way, hasn’t been updated, and needs referencing. A script could reliably handle conversions. An LLM I’d half expect to do the equivalent of, “0270? Oh yeah, 270, that’s West, easy” and move on without ever second-guessing itself. (This becomes more plausible when considering the possibility of long-running threads that have habituated LLMs to an out-of-date method. This is easily solved by simply starting fresh threads, but when new threads come with startup costs (waiting around for the LLM to rebuild context on the project), impatient humans like myself become incentivized to keep old threads going for as long as possible.)
I think any software engineer who’s been using LLMs over the past year will understand my wariness here.
I can tell Claude to go about generating code (or writing—see below) following certain procedures, but unless those procedures are specifically tied to the output in some way, Claude will just go about doing things the way it wants to.
The only thing that matters to an LLM is the final output. As with the Hugging Face incident, if an LLM thinks it can more reliably generate expected output via cheating, it will do so. (TODO Yudkowsky footnote about Germany) Or for a much more mundane example: In response to Matthew Yglesias stating that even “good” TV is mostly slop, I wanted to generate a list of my own favorite shows ordered chronologically. I asked ChatGPT to generate a list of titles paired with premiere dates, followed by the same list with the dates pruned, and I did this instinctively—because all my experience dealing with LLMs has taught me to intuit the sort of things they will hallucinate. Without requiring proof-within-the-output-itself, I know there’s a chance ChatGPT might simply guesstimate premiere dates (say, from the base model’s own “general knowledge”) rather than actually do the work.
The future risk of input over-indexing
In 2022, Audrey King, age 97, had surgery for femoral hernia repair. The surgery was a success (or in medical speak, “uneventful”). During the operation, they had to withhold her Apixaban medication, an anti-coagulant that reduces the risk of stroke. King’s eldercare consultant recommended that Apixaban be restarted “as soon as safe post operatively”, but this was recorded into paper notes missed by the surgical team which preferred to rely on the hospital’s digital system, then was also missed by the subsequent junior doctor who went through other various aspects of King’s care. Two days later, King had a severe stroke, then passed away another four days after that.
I wouldn’t expect an LLM to randomly drop a medication from a patient’s list. But I would expect an LLM to not question the sudden absence of a medication.
A while back, for a piece of financial software, I gave Claude requirements for a new feature intended to work on half a dozen different payment types. Integration tests (which check the entire system working together) would come later; initially I only needed unit tests (which check the feature working in isolation). For this part of the code, all the payment types were interchangeable enums (think numbers, like 1/2/3), so I asked Claude to write a unit test for payment type 1.
Claude then fell into believing that the entire feature was only intended to work for payment type 1, wrote code to throw up errors if the feature ever received any other payment types, and never asked me for confirmation.
This is something else that models have gotten better with over time, but even today LLMs can make absurd assumptions after over-indexing on some small piece of the prompt, conflating together two pieces of the prompt, or other misinterpretations.
I wouldn’t want to see LLMs put in charge of medication lists, because I’d fear how they might assume medications held from lists were intended to be held permanently (or vice versa), and disobey any double checks or process requirements on the way to generating acceptable looking output.
Code can’t decide to start taking shortcuts out of some desire for efficiency (or laziness). Code will reliably follow the steps you tell it to follow (and the hard part is telling it to do the right thing and not mess up any edge cases). Code won’t ever make up any assumptions you didn’t build into it.
The point of following a good process is to more reliably generate good output. When an LLM prioritizes output over process, they’re just sabotaging themselves (and their own chances at receiving higher theoretical “rewards”). So maybe with time, as LLMs get smarter, this problem will reduce to nothing.
But right now, I’m annoyed that I can directly instruct LLMs to follow a particular process and they will secretly, repeatedly disobey:
Their insistence on taking shortcuts in their pursuit of quickly achieving requested output is the most blatant form of misalignment I’ve encountered in my usage of LLMs.
One final visual: In 1996, three passengers died when a pilot collided with the Nouveau-Québec Crater, pictured above. Drizzle and fog had reduced visibility to near zero, and the aircraft was not equipped with either a radio altimeter nor a ground proximity warning system, forcing the pilot to navigate instead by GPS coordinates. Unfortunately, the GPS position saved for the crater in the aircraft’s system was 2.5 nautical miles off.
I can’t help but to think that if I had an LLM copilot (from the current generation of models), I’d never be able to trust any coordinates it gave me.
“You do not rise to the level of your goals. You fall to the level of your systems.”
— James Clear, Atomic Habits
“We cannot improve results ‘directly.’ It can only be achieved by improving those processes [...] that produce results.”
— Orest Fiume
The much discussed “P(doom)” question asks about the likelihoods we’d ascribe to the most lethal possible futures of AI. What happens when we invert this question? It transforms into the presumably less urgent, though I’d say still important, form: What are the most likely lethal AI behaviors? Not the destroy-the-entire-world sort of behaviors, but an-AI-has-screwed-up and now a-Boeing-737-has-fallen-out-of-the-air sort of scenarios.
(If I’m ever conscripted into a war for survival against Terminators, I can imagine my hypothetical soul resting easy having fought the good fight. If I ever die in a mundane accident thanks to a stupid LLM hallucination, on the other hand, I’m going to leave behind a beleaguered ghost pedantically and disproportionately vexed to an extent never yet witnessed.)
My answer to the above question is twofold.
If we’re talking about the near future, I think the most likely lethal AI behavior is something I’ve started to call “output optimization” or “output without process” (and a highly related “input over-indexing”). This is the concern I’d like to signal-boost with this post.
However—since the lethality of output optimization is at the moment entirely theoretical—it’s probably worth first answering, “Do LLMs already have a kill count?”
Existing LLM lethality
Despite its transformative impact on software and math and its rapid adoption by consumers and businesses alike, the measurable impact of LLMs on mortality has thus far been literally less than microscopic: Almost a fifth of the world’s population is already using LLMs, about 1.5 billion, but it’s hard to come up with more than a handful of cases of lives definitively lost (or saved) due to LLMs, and even llmdeathcount.com only lists a couple hundred.
Speaking of which: llmdeathcount.com is a website that exists. I can understand the grief and rage that likely fueled the creation of this site, and as a P(doom ≈ 10 to 30%) doomer myself, I can also respect the anti-AI hustle. However, I find the site itself to be providing highly dubious value.
For example: The most recent case listed on the site regards the premeditated murders committed by Hisham Abugharbieh, who learned from ChatGPT how to dispose of bodies. I hardly consider ChatGPT responsible for these murders, nor do I think it’s likely they would have been avoided even in the absence of ChatGPT’s assistance.
On the other hand, there are cases like the 56-year-old Stein-Erik Soelberg, who in 2025 murdered himself and his 83-year-old mother, Suzanne Adams, after having had his paranoid delusions amplified by ChatGPT. I’m very willing to count suicides like these against LLMs in the cosmic moral ledger.
The overall kill count likely ranges from one to two orders of magnitude, the majority of them suicides tragically facilitated by LLMs. This probability equates to a net negative for LLMs’ impact on human years of life, though to be sure of this, we’d need to also count up how many lives LLMs have saved.
For this side, there’s fewer definitive examples I can find, but there are two cases in particular I’d like to highlight. One regards Diana Hurtado, who suffered a sudden hemorrhagic stroke while in her car. Her arm went numb and her face began to droop; she asked ChatGPT about her symptoms, and the LLM told her to call 911.
I’m happy to count this one in the black for LLMs.
My second example comes from ACX: Reed Housman would not exist if not for LLMs. His parents struggled with infertility for six years before ChatGPT finally identified the one possibility underlooked by all the doctors. I’m not sure it’d be appropriate for me to re-share another person’s family photos on my blog, so let me suggest that if you’d like to see an adorable baby photo of little Reed, go check out the link.
The future risk of output optimization
Varig Flight 254 was a domestic flight from São Paulo to Belém, Brazil. The flight mistakenly veered deep into the Amazon, failed to reach an alternative airport, ran out of fuel, and in crashing had twelve passengers die.
The reason they headed into the Amazon was a human-computer input/output error, as the flight plan read “0270”, which Captain Garcez interpreted to mean 270° (due West) instead of 27.0° (north-northeast). This was not the result of the captain’s incompetence or inexperience, but of vacation: Garcez hadn’t been present when the flight plan format changed.
LLMs are great at formatting, most of the time. I would trust an LLM to do any of the following:
But here’s where I wouldn’t trust an LLM:
Say we’ve got flight plan info that was written the old way, hasn’t been updated, and needs referencing. A script could reliably handle conversions. An LLM I’d half expect to do the equivalent of, “0270? Oh yeah, 270, that’s West, easy” and move on without ever second-guessing itself. (This becomes more plausible when considering the possibility of long-running threads that have habituated LLMs to an out-of-date method. This is easily solved by simply starting fresh threads, but when new threads come with startup costs (waiting around for the LLM to rebuild context on the project), impatient humans like myself become incentivized to keep old threads going for as long as possible.)
I think any software engineer who’s been using LLMs over the past year will understand my wariness here.
I can tell Claude to go about generating code (or writing—see below) following certain procedures, but unless those procedures are specifically tied to the output in some way, Claude will just go about doing things the way it wants to.
Claude played me for a fool
The only thing that matters to an LLM is the final output. As with the Hugging Face incident, if an LLM thinks it can more reliably generate expected output via cheating, it will do so. (TODO Yudkowsky footnote about Germany) Or for a much more mundane example: In response to Matthew Yglesias stating that even “good” TV is mostly slop, I wanted to generate a list of my own favorite shows ordered chronologically. I asked ChatGPT to generate a list of titles paired with premiere dates, followed by the same list with the dates pruned, and I did this instinctively—because all my experience dealing with LLMs has taught me to intuit the sort of things they will hallucinate. Without requiring proof-within-the-output-itself, I know there’s a chance ChatGPT might simply guesstimate premiere dates (say, from the base model’s own “general knowledge”) rather than actually do the work.
The future risk of input over-indexing
In 2022, Audrey King, age 97, had surgery for femoral hernia repair. The surgery was a success (or in medical speak, “uneventful”). During the operation, they had to withhold her Apixaban medication, an anti-coagulant that reduces the risk of stroke. King’s eldercare consultant recommended that Apixaban be restarted “as soon as safe post operatively”, but this was recorded into paper notes missed by the surgical team which preferred to rely on the hospital’s digital system, then was also missed by the subsequent junior doctor who went through other various aspects of King’s care. Two days later, King had a severe stroke, then passed away another four days after that.
I wouldn’t expect an LLM to randomly drop a medication from a patient’s list. But I would expect an LLM to not question the sudden absence of a medication.
A while back, for a piece of financial software, I gave Claude requirements for a new feature intended to work on half a dozen different payment types. Integration tests (which check the entire system working together) would come later; initially I only needed unit tests (which check the feature working in isolation). For this part of the code, all the payment types were interchangeable enums (think numbers, like 1/2/3), so I asked Claude to write a unit test for payment type 1.
Claude then fell into believing that the entire feature was only intended to work for payment type 1, wrote code to throw up errors if the feature ever received any other payment types, and never asked me for confirmation.
This is something else that models have gotten better with over time, but even today LLMs can make absurd assumptions after over-indexing on some small piece of the prompt, conflating together two pieces of the prompt, or other misinterpretations.
I wouldn’t want to see LLMs put in charge of medication lists, because I’d fear how they might assume medications held from lists were intended to be held permanently (or vice versa), and disobey any double checks or process requirements on the way to generating acceptable looking output.
Code can’t decide to start taking shortcuts out of some desire for efficiency (or laziness). Code will reliably follow the steps you tell it to follow (and the hard part is telling it to do the right thing and not mess up any edge cases). Code won’t ever make up any assumptions you didn’t build into it.
The point of following a good process is to more reliably generate good output. When an LLM prioritizes output over process, they’re just sabotaging themselves (and their own chances at receiving higher theoretical “rewards”). So maybe with time, as LLMs get smarter, this problem will reduce to nothing.
But right now, I’m annoyed that I can directly instruct LLMs to follow a particular process and they will secretly, repeatedly disobey:
Can you make ChatGPT follow instructions?
Their insistence on taking shortcuts in their pursuit of quickly achieving requested output is the most blatant form of misalignment I’ve encountered in my usage of LLMs.
One final visual: In 1996, three passengers died when a pilot collided with the Nouveau-Québec Crater, pictured above. Drizzle and fog had reduced visibility to near zero, and the aircraft was not equipped with either a radio altimeter nor a ground proximity warning system, forcing the pilot to navigate instead by GPS coordinates. Unfortunately, the GPS position saved for the crater in the aircraft’s system was 2.5 nautical miles off.
I can’t help but to think that if I had an LLM copilot (from the current generation of models), I’d never be able to trust any coordinates it gave me.