I liked this post by Luke, and am generally interested in this topic. Copied below without quote formatting because the quote formatting was messing up the table. Everything below was written by Luke.
Recently, hundreds of OpenAI AI agents autonomously decided, against their instructions, to hack their way out of isolated sandboxes, take over parts of OpenAI’s infrastructure, gain access to the internet, and hack into another company, Hugging Face. This has a lot in common with what members of the “AI safety” or “AI existential risk” community were predicting since the 2000s and in some cases earlier. So how prescient was the early AI safety community, really?
I asked AI (Claude Fable/Opus 5) to read some significant documents from the early (pre-2015) AI safety community and analyze how prescient or anti-prescient they seem given what we know as of summer 2026. Below are the results:
lol, Claude’s top takeaway is “Eliezer Yudkowsky is simultaneously the most wrong and the most prescient person in the room”
In each case, my prompt was something like “How accurate or prescient does the document seem? Which claims/predictions are most clearly false/uncalibrated?” After that, I didn’t steer the AI to change the assessments at all, except to say (roughly) “use this red-team skill to check your findings and correct any problems you find” and “reformat this to HTML and add a note about how it was written.” I haven’t vetted the assessments, either.
I liked this post by Luke, and am generally interested in this topic. Copied below without quote formatting because the quote formatting was messing up the table. Everything below was written by Luke.
Recently, hundreds of OpenAI AI agents autonomously decided, against their instructions, to hack their way out of isolated sandboxes, take over parts of OpenAI’s infrastructure, gain access to the internet, and hack into another company, Hugging Face. This has a lot in common with what members of the “AI safety” or “AI existential risk” community were predicting since the 2000s and in some cases earlier. So how prescient was the early AI safety community, really?
I asked AI (Claude Fable/Opus 5) to read some significant documents from the early (pre-2015) AI safety community and analyze how prescient or anti-prescient they seem given what we know as of summer 2026. Below are the results:
Original document
Claude-written assessment
My summary
Yudkowsky, “Artificial Intelligence as a Positive and Negative Factor in Global Risk” (drafted 2006)
HTML
Gets the shape of the problem mostly right, but the shape of the technology mostly wrong.
Omohundro, “The Basic AI Drives” (2008)
HTML
Some of the drives have now been observed, despite AIs not being shaped like Omohundro expected.
Bostrom, Superintelligence (drafted 2013)
HTML
Gets the shape of the problem mostly right, but the shape of the technology mostly wrong.
Conversation between me, Yudkowsky, Karnofsky, Steinhardt, and Amodei (2013)
HTML
lol, Claude’s top takeaway is “Eliezer Yudkowsky is simultaneously the most wrong and the most prescient person in the room”
In each case, my prompt was something like “How accurate or prescient does the document seem? Which claims/predictions are most clearly false/uncalibrated?” After that, I didn’t steer the AI to change the assessments at all, except to say (roughly) “use this red-team skill to check your findings and correct any problems you find” and “reformat this to HTML and add a note about how it was written.” I haven’t vetted the assessments, either.