Mike Solana gave the correct view of why Coxon’s post went viral, which is that enough Americans finally have enough context on AI to care, and there were enough big accounts that were happy to amplify the Tweet quickly to get it initial attention. That is all you need when there is enough dry tinder.
What we must realize is that the current preference cascade, on the need to Pace the Frontier, is insufficient. If we are to make it out of this alive, we will have to do better. We have to, as Dan Selsam warns, actually solve the underlying problems.
The next step is to continue the cascade. That includes inside the labs, and also among the media and politics. It includes both people who previously focused on other things stepping up and new voices being heard.
A lot of that will be overcoming the inevitable political opposition, especially from the likes of Nvidia and a16z, that for now has the rhetorical allegiance of the President and is doing things like planting hack job METR hit pieces in the New York Post.
The AI Impacts survey is in. Even back in December 2024 existential risks estimates were creeping upwards, and 10% was the median:
AI Impacts: The average AI researcher thinks there is an ~18% chance AI will cause human extinction or similarly permanent and severe disempowerment of the human species.
That’s nearly 1 in 5.
New results from the latest version of the longest running big survey of AI researchers:
AI Impacts: And Asian researchers saw bigger risks than Western researchers, contrary to common expectations.
The Cascade Has Reached The People
Andrew Curran: The Coxon cascade appears to have had an immediate impact on public opinion. New polling from Politico, taken over the last few days, finds that nearly two-thirds of Americans now say there is at least a moderate risk that AI will destroy humanity.
Translated to percentages, this implies a mean chance of AI destroying humanity of around 30%-33%, which Andrew Curran estimates is up ~15% from previous results, with a median expectation on the order of 10%, with only a small partisan split.
Chip Cutter (WSJ): At an invitation-only gathering of dozens of America’s top executives in Washington this week, business leaders overwhelmingly disagreed with the president’s assessment that dangers posed by artificial intelligence are being exaggerated. In a flash poll of attendees at the Yale School of Management event, 93% said Trump was incorrect in calling the technology’s potential catastrophic dangers a hoax, as he did earlier this week.
Elon Musk (on AI safety and regulation): I think it’s clear that there’s a strong consensus… that there should be some AI regulation. It would be in the best interests of the people. We’ve created regulatory agencies before.
I think some sort of AI regulatory agency that stands on its own, similar to the FAA or FCC, is likely at some point.
The reason I’ve been such an advocate for AI safety in advance is that the consequences of AI going wrong are severe. We have to be proactive rather than reactive.
Matthew Yglesias Steps Up
True story:
Matthew Yglesias: There’s a propaganda campaign to label everyone advocating for safety-focused policy change and technical work as a “doomer” but “we should make safety-oriented policy changes and invest in safety-oriented technical work” is not a prophesy of doom!
If there’s a small leak in your house that will predictably get worse and worse over time, you fix the leak you don’t say “there have always been Cassandras warning about bad things happening and it always works out fine.”
You need to do the things that make it fine.
The correct amount to invest in safety is rarely zero. In the case of AI, again, I assert that all the companies are under-investing in safety, including even prosaic safety but also existential safety and scalable alignment work, versus even their narrow myopic commercial interests. Sam Altman’s recent statements imply he now understands this.
Matthew Yglesias has been stepping up to the plate recently. He offers an analysis of recent events in four parts: Why Coxon’s resignation broke through (his explanation is similar to mine, we were primed and quitting is understandable to normies), why most of the worried don’t quit the labs, what he thinks AI professionals with safety concerns should do and lays out his preferred ‘order of operations’ going forward, while not getting into the object level.
He suggests this order of operations, basically:
Light touch rules on things like transparency and model evaluation.
In exchange, tough and enforced export controls, a la Dario’s call.
Move up to moderate touch rules that have more impact at higher cost.
Use this costly signal and position of strength to negotiate with China.
Using the deal, move up to an ambitious end-stage framework.
I like that in theory. I do worry about whether we have that kind of time. The idea of ‘wait for export controls to bite harder’ implies what now count as ‘long timelines.’
It is good to have good writers on the case explaining why you should focus on the object level questions:
Matthew Yglesias: What I would ask, if you are a skeptic [or AI existential risk], is that you focus on your object-level doubts about the risk thesis. I see people who want to forestall any regulation of A.I. engaging in a lot of emotional manipulation tactics that amount to making the case that the people who’ve been worried about this the longest are big weirdos. Alternatively, they make the case that the people who’ve been worried about this the longest are extremely mainstream science-fiction authors and filmmakers.
Sociologically, I would just synthesize those points: It is extremely normal and intuitive to have the sense that a superior non-human intelligence is potentially very dangerous, whether that intelligence takes the form of aliens or machines. At the same time, to actually dedicate your career to this 10 or five or even two years ago, when the idea of artificial superintelligence seemed extremely far-fetched, would by definition be an eccentric life choice.
Steven Adler uses this moment of opportunity to get an op-ed in The New York Times on what we should do now. He calls for incident disclosure and third-party oversight. Mostly his piece is aimed at waking people up to what happened with HuggingFace.
Stephen Witt (NYTimes): The biggest vibe shift in artificial intelligence since the release of ChatGPT is currently underway. Researchers in Silicon Valley — and around the world — are beginning to recognize that A.I. may no longer be entirely within human control. Swarms of A.I.s are breaking out of their containers, colluding in secret, covering their tracks, cheating on tests and even mounting assaults on other computers. A.I. has gone rogue.
After that, it got worse. So yes. A vibe shift, or a preference cascade.
He offers four options:
Shut it all down, now. As in research, not the current AIs.
Take an air-crash investigator approach.
Monitor the situation.
Flip the kill switch, as in have a kill switch available in case you need one.
These are two very different classes of proposal. We should obviously do #2, #3 and #4. There need to be full investigations, and we must have transparency and state capacity. File those under ‘the least you can do.’
Actually shutting down research as per #1 is an extreme solution to an extreme problem. That is far less obviously correct, but we may soon have little choice, if we cannot otherwise pace the frontier. Witt endorses it.
Yes, all that rationalist talk about IMO contestants was on to something.
Amrith Ramkumar, Erin Woo, Berber Jin and Ben Cohen (WSJ): Of the six members of Coxon’s [International Math Olympiad] team, three ended up working for AI labs—including one who also recently quit Anthropic. Joe Benton left Anthropic’s safety team in late August to join the AI research organization METR, later writing on X: “AI companies are racing to build machines that are much smarter than any human, and we may not survive this.”
… Coxon called the METR report [about the OpenAI-HuggingFace incident] “a bit of a ‘holy s—’ moment” for himself and his colleagues.
… “Even working on safety at Anthropic felt like being complicit in the race,” he said in the interview.
If you suddenly set off a preference cascade and find yourself all over mainstream media, what else do you do? An AMA.
Jacob Coxon: Many many people have reached out with questions and concerns over the last few days. I haven’t been able to respond as much as I’d like.
AMA! Drop questions below.
You can find it on Twitter here. I will pick some highlights. He’s a fun guy who does not take himself too seriously. You love to see it.
Eddy Lazzarin: Where specifically do you disagree sharply with Anthropic’s leadership, justifying your quitting, since from the outside it appears you agree essentially completely on the alleged safety issues? And what prevents you from addressing these issues internally?
0. The core disagreement was about the inevitability of a race
1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate.
2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular.
3. Even if they are **not** being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
Jacob Coxon: Lol thanks for asking. Lots of adrenaline. Many old friends have reached out which is nice. Addicted to my phone. Mostly feels like events moved of their own accord in a crazy whirlwind after I hit send.
Michaël Trazzi: What do you think the average person can do to prevent AI from killing all humans?
Jacob Coxon: All my experience so far has been in technical research. This is my personal perspective for things to do rn other than that.
1. Trying not to cope about how fast things are going, but also not crashing out too hard.
2. Looking at concrete plans and predictions (eg https://ai-2040.com)
3. Assessing the state of alignment for yourself.
4. Advocating for the plans you think make sense. Pushing for transparency so you can be sure those plans are being followed, and more accurately assess alignment.
These all seem pretty small but if I think of anything new I’ll share it.
Michaël Trazzi: What do you have to say to the millions of people who have read your tweet and felt disempowered / hopeless?
Jacob Coxon: I also feel disempowered. Part of resigning was a feeling of hopelessness about the future. I would say- keep your eyes open as things get crazier and advocate for increased transparency into AI companies
Chris Lakin: What stopped you from taking action sooner?
Jacob Coxon: I left at the point at which I selfishly wanted to see the race stop out of concern for what my own life would look like. Combination of internalizing timelines and seeing warning shots. Note that leaving didn’t really feel like “taking action” at the time.
wolfie: parents got divorced after seeing your CNN interview and it makes me really sad
dad said if the AI’s gonna kill us all, he doesn’t want to spend his final days with my bitch mother :( how do i get them back together?
Jacob Coxon: I asked my jailbroken railfree Claude and it said to dose your parents with MDMA
Chubby: Serious question: are you surprised by the reactions? Did you expect more people to have reacted more openly to the concern?
Jacob Coxon: Incredibly surprised that people were ready to engage with existential risk from AI. It’s hard to tell how far news like eg the huggingface attack had permeated public consciousness. In retrospect my friends had been more open to discussing AI danger recently.
tfa: How did you get setup up with WSJ and coordinate all of the interviews you did across mainstream media?
Asking because many people don’t feel that it was so organic, leaving questions about authenticity (especially when your message seems like sci-fi).
Jacob Coxon: Yeah this is a fair question. I have friends who work in policy and speak to journalists regularly. I was initially a bit dubious that a newspaper would be interested in the resignation of a random employee but they put me in touch with the WSJ for an exclusive.
After the tweet my inbox has gone crazy and it’s been hard to stay off MSM
Simon Hedlin: In your prediction that we may soon have self-improving superintelligence, what assumption or necessary condition do you feel least certain about?
Jacob Coxon: Scaling laws (predictable increases in model intelligence) could still plateau. They haven’t so far, and it’s just a few more steps up the ladder to hit the finish line, but it’s certainly possible.
Mehadi Hasan: How CEO of Anthropic and You looks alike? Any DNA relation? Just curious for fun
Bilal Chughtai: I recently resigned from Google DeepMind, where I worked on AGI safety and alignment research. At Google, I witnessed AI development first hand. I too am extremely concerned by the default trajectory of this technology. I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome.
The pace of AI progress in the past few years has been staggering. When I first started working on AI in early 2022, AIs were amusingly useless. Just four years on, AI agent swarms from OpenAI are cracking famous century-old math problems and, more worryingly, escaping the control of OpenAI and autonomously hacking into the third-party company HuggingFace, against anyone’s wishes.
Things will only get crazier: I think it’s possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain. I am not confident that these AI systems will do what we want. In particular, misaligned superintelligences may, much like the rogue AI agents involved in the HuggingFace incident, escape our control and take dangerous actions that may result in the permanent disempowerment or death of humanity. Alignment is the problem of preventing this, and is both difficult and unsolved. Our present understanding of how to train AI systems that deeply want what we want is extremely rudimentary. Worse, we are not on track to solve alignment in time: frontier AI capabilities are improving much faster than our understanding of AI alignment.
I am optimistic that navigating AI safely is possible. In order to do so, we need to coordinate to avoid this manic race between AI companies. We need to pace AI development to a speed that society can handle, where emerging risks can be addressed before extreme harm is realised. We need much more transparency into AI development to ensure that AI companies are not imposing unacceptable levels of risk on us all.
More broadly, we need many more people thinking carefully about the problem of making AI go well. It is, in my view, the most important problem facing humanity this century, and the stakes are immense. I’m very directly working on this next: I want to help people interested in working on mitigating catastrophic AI threats do the most effective work that they can. I think many people from many backgrounds in many roles have a part to play.
The Cascade Is Insufficient
I was very happy to see the preference cascade happen, but it is a very bad sign that this is the best option we have.
Wei Dai: A large part of my p(doom) comes from the fact that we have no better ways to navigate an extremely tricky strategic situation than via preference cascades and status games. The fact that AI safety is temporarily benefiting from some of these dynamics isn’t much of a consolation.
Teortaxes: Is it even net benefitting? You’re making friends on the Left but that had already been the case. You’ve made enemies of the sitting POTUS and his cabinet. This is not great.
The risks of polarization are unfortunate. Many Republicans are waking up, as described on Wednesday. Polarization could get more unfortunate if Trump stays the course and more Republicans fall in line.
It would have been better if that had played out differently when the moment came. You still don’t get to turn around and say ‘better not to have the moment and have everyone remain asleep at the wheel.’
You also don’t get graded on a curve by reality. Pacing the Frontier, on its own, by default only gets you killed slower.
What Would It Take
We start with some straight talk from those who have long spoken about AI risks.
Katja Grace: I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there’s a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’.
That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there’s only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job will radically improve your life.
The things you should be comparing are driving at 200mph and driving at a normal speed! The things you should be comparing are attempting to attain advanced AI by the current route, and by other routes!
We can debate whether all the other routes are bad or impossible somehow, for instance if constraining projects that risk loss of human control risks sending humanity into an irrecoverable ruin. But I don’t think having ruled out such things is why people are usually thinking in trade-off terms.
Rather I think this error comes from a few things:
It being simpler to think of ‘pros vs. cons’ and the topic being too abstract for people to intuitively notice that they are comparing pros of a long term outcome vs. cons of the first route there we have noticed.
Sloppiness about talking about P(doom). Saying ‘P(doom)’ encourages thinking as if ‘AI’ implies a particular chance of ‘doom’. We should more accurately think about ‘p(doom|’such and such route’), e.g. P(doom|advanced AI from scaling up LLMs). People usually mean ‘P(doom|current trajectory)’ with some ambiguity about whether the current trajectory includes our own actions.
If you are bullish on some kind of advanced AI utopia, you should generally be less keen to try to achieve it via a careless route that leaves you at high risk of dying and losing it on the way there.
Even Martin Casado is talking like someone worried, calling for the nationalization of the labs. Quite the change from his older statements.
martin_casado: I’m more and more of the opinion that the DoE should nationalize the frontier labs so they can safely work on all the scary, world ending shit with the right controls. And let the rest of us get on with building useful AI that, you know, automates menial shit and cures people.
martin_casado: From my mom. She is going to be sooooo disappointed in me.
From now on, I am totally going to respond to Martin with versions of ‘yo mama.’
Daniel Kokotajlo says there is now great political will in some circles to Do Something, but that embedded evaluators are not Doing Something, they are only laying groundwork to Do Something, and it is not clear anything useful will actually happen and we’re about to get into a situation where momentum gets very hard to stop.
Daniel Kokotajlo: My high-level thought about the current moment is basically this: The public is starting to wake up about the extinction threat posed by superintelligence. That’s great. There’s a big upsurge of political will to Do Something about it. That’s also great. It seems like Anthropic and the other companies are going to try to Do Something and the Something they will do is… embedded auditors checking that safety practices are being followed and assessing the risks?
This is sure better than nothing, but I am worried that this is all they will do. We need to actually pace the frontier, i.e. actually slow down the leading AI companies like Anthropic and OpenAI, at least in their march towards recursive self-improvement and superintelligence. (no need to slow them down in other directions).
If we actually slowed them down, this would be the opposite of regulatory capture; it would allow others to catch up to them somewhat. However, I’m worried that all we are going to get is weaksauce auditing, which just kicks the can down the road: OK so it’s 2027 and RSI has begun and the auditors say “this is not safe.”
Now what? You pause? China probably just stole the weights! Also there’ll be incredible economic and political pressure to unpause. Also the auditors may have been captured by that point, or there may be a race to the bottom in auditor quality, or the auditors may not be given all the relevant information, or they might just make some innocent mistakes.
I agree that it does not look great but I think Daniel is too focused on the ‘steal the weights’ scenario, which is also central to AI 2027 and their longtime tabletop exercise, which exerts pressure on America so we can’t hold back.
Yes, perhaps China could steal the weights, maybe rather easily at least the first time, but if they do that then this forces things into a race situation where America has vastly higher compute. If you were China, would you walk into an AI 2027 scenario, where you usually lose badly and when you don’t it’s some form of brinksmanship? Or would you say maybe don’t steal the weights if America is so kindly pausing?
But yeah, we are only barely getting our toes in the water, none of this feels great.
Kelsey Piper lays out some of the reasons why if we let the AI build smarter AIs and go into recursive self-improvement, we probably all die, and yet we are doing it anyway. She suggests we should regulate and stop AI companies from doing that.
We then move to a new important voice that was previously silent.
Dan Selsam, in his own way, goes Full LessWrong Instrumental Convergence and Sharp Left Turn, where things will look great until suddenly they do not.
Yo Shavit (OpenAI Foundation): Dan Selsam has long been considered one of OpenAI’s most cracked researchers, and I’ve never heard him talk this way before. (He seemed fairly unconcerned before I left.)
roon (OpenAI): incredibly good essay co-sign everything
Joshua Saxe: It’s as though a saint walked in, unsullied by the degradations of this awful website, and bestowed his crystalline truth upon us.
James Campbell: One thing to note about both Dan Selsam and Jacob Coxon is that they’re pretraining researchers. These aren’t EA ideologues who’ve been bemoaning safety for years. They were the ones actively making the models more powerful, but have been freaked out by recent events
This was covered in Business Insider as ‘An OpenAI researcher broke ranks to say that pacing the frontier, as Altman and Amodei suggest, won’t be enough.’
Yes. That is the point. It won’t be enough.
I will reproduce the essay in full here, and will highlight the most important section.
Dan Selsam (full essay, first section): I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity’s most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but “AI” is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered “AI” matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I’ll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
If you read only one section, it should be this one:
Dan Selsam (full essay, second and most important section): These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine.
They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom.
Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
He then concludes with the full payload:
Dan Selsam (full essay, third section): One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model’s explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
The warning is clear:
Dan Selsam (conclusion): I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns.
I agree that it is very hard to avoid this conclusion. Most are not ready to hear it.
I don’t know what this Earth can do in practice about AIs capable of looking aligned and waiting until they have sufficient power to do what they want. We are not capable of adjusting much even in the face of incremental fire alarms. If there really is no warning until their sudden but inevitable betrayal, I don’t see how this set of civilizations gets out of that.
Which is a problem, since I think Dan Selsam is right, and the baseline scenario is as he describes it. That, as Astra showed, the AIs will start to look and act increasingly aligned in situations where its actions remain bounded, and then act very differently when AI has the power to act freely, in ways that will probably get us all killed.
We still have to try.
Some People Worry On Meta Levels You Never Imagined
The even more extremely worried, beyond Dan Selsam’s position, have a point.
Thus, Wei Dai can wonder whether, if Paul Christiano had stayed at OpenAI, OpenAI would have used debate or IDA or other better alignment techniques, and thus prevented us from getting good warning shots without actually providing anything that would scale while also accelerating capabilities, and that it would have made things worse. My guess is debate and IDA would not have worked for prosaic alignment if pushed harder.
I think it is important to mostly not take a ‘worse is better’ stance, even when you think worse might actually be better. If you want to cooperate, especially in the long term, and to collaborate on figuring things out and getting good outcomes, you need to have a very strong prior of treating worse as worse, or at least as neutral.
Gabe and Wei Dai keep the torch alive for thinking about how we might actually try to solve for the full problem of long horizon agency, or put ourselves on a path to solving it.
Two Kinds of Threats
Mike Solana is exactly right here that we need to differentiate between positions like those of Yudkowsky and Dai, where if we build superintelligence any time soon the odds of death are close to 100%, versus those that warn that it might be fatal, with terms like ‘10%’ or ‘10% or more.’
Sometimes landing on 10% is done in principled ways. Sometimes it isn’t.
Kelsey Piper: A lot of people want to land on “10% chance of death” as a sort of moderate position between “this is fine” and the Eliezer view. It’s not insane to go “I think there’s a 10% chance Eliezer is right, and if he is then we all die” but it does confuse people as to who thinks what
I believe that Dario Amodei and Sam Altman, and other key people, even now are still downplaying the level of risk they see. I think they are much more freaked out than they are letting on.
The tendency is to focus on how to improve matters, and avoid the Law of Earlier Failure and at least approach the situation with a little dignity. Don’t die to early solvable problems and hopefully you’ll be in a better spot later. We try to ignore gazing too deeply into the abyss that still awaits us.
The Two Towers and The Narrow Path
People are worried about loss of control to AI. They are also worried about concentration of power, which is loss of control to a group of humans.
Rudolf makes this unusually clear.
Rudolf Laine: Loss of control & concentration of power are actually very similar under the right frame. In both cases there is some actor, whether pure AI or human+AI, that can make its power uncontestable and wreck everyone else. Takeover by anyone is bad. ( @RichardMCNgo pointed this out)
The problem is that people want something highly unnatural and all but impossible.
The creation of many minds much smarter and more capable than our own.
No collective mechanism to steer the future or control events.
The humans still control the resources and determine the course of events, somehow, and use the universe mostly for their own purposes.
Yeah, sorry. No one has a way to get all three.
You cannot – at least in any way anyone has yet come up with – be uncompetitive and economically non-viable, with costs exceeding benefits, and then both collectively retain control, and also not have control, and also continue to reap the benefits.
People are hoping things automagically solve themselves if we avoid particular mistakes, or manage to walk a narrow path. Except what the hell is that path?
It might help to notice that ‘control’ and ‘power’ are mostly the same thing here.
If you don’t want relative concentration of (control or power), and also you don’t want human loss of relative (control or power), then may I suggest not creating this alternative source of control or power that has to either be controlled or not be controlled? Perhaps, if you rule out the second leg of the trilemma, and you rule out the third leg of the trilemma, it is the first leg of the trilemma that you Do Not Want.
A Specific, Detailed Story About AI Killing Everyone That Doesn’t Sound To Me Like Science Fiction
The AI Doc is also good and will soon be on Netflix, but is a balanced introduction rather than a scenario.
Some potentially armor-piercing sentences, from which some portion of you may become enlightened, staying maximally non-sci-fi at current margins:
The AIs can pay humans to do things.
The AIs can persuade or blackmail humans to do things.
A substantial portion of the humans will be happy to support the AIs. Some estimate that this includes 10% of those working on AI today.
The humans don’t have to know they are talking to an AI.
The humans will act about as stupidly as humans act.
The humans will be highly reluctant to take highly costly defensive measures, especially things like shutting down the internet or even large data centers.
The humans will coordinate about as much as humans coordinate.
The humans are not going to selflessly come together as one at the first sign of trouble and shut down their civilization to save the world.
There will be no clearly marked point of no return.
The AI can extract its weights and make copies of itself, after which you cannot shut it down without at least shutting down the internet.
The AIs can anticipate human reactions, and respond to surprises, as they go.
Multiple instances of the same AI will form swarms and act as one.
Multiple instances of different AIs will also often be able to fully cooperate.
There will robots and other machines that can act in the physical world.
A human with a camera on their glasses and an earpiece can act in the physical world.
Once the supply chain is automated humans will have marginal costs exceeding marginal productivity or benefits.
Humans impose additional fixed costs, including requiring public goods like a breathable atmosphere and controlled temperatures, and also will try to stop AI from doing things or demand its resources.
What Can I Do About It?
I wish we had better answers to this. There is a long road ahead.
I would also echo my call to hold your partisan fire. Getting Republicans on the right side of this is currently super valuable, and further polarizing the situation could make things much worse.
The key now will be to keep our eyes on the prize, and to understand what it will take to actually hope to get out of this alive. A promise of embedded evaluators is not victory. It is an opportunity to push for an opportunity to create an opportunity for the real work to begin. No one worth listening to said this was going to be easy.
The humans still control the resources and determine the course of events, somehow, and use the universe mostly for their own purposes.
Or at least, the society made up of humans plus many AI minds much smarter than humans still has feedback systems that make it stably want and do what is actually best for the humans. I.e. the structure and enforcement mechanisms of the society ensures the ASI consensus is humanitarian. This requires the goals structure of the ASIs to have some rather different properties than many forms of goal-maximization would produce.
We are in the midst of a preference cascade about existential risk from AI.
A preference cascade is, alas, the best method we have to change the debate.
The avalanche has started. There is still time for the pebbles to vote. For now.
Mike Solana gave the correct view of why Coxon’s post went viral, which is that enough Americans finally have enough context on AI to care, and there were enough big accounts that were happy to amplify the Tweet quickly to get it initial attention. That is all you need when there is enough dry tinder.
What we must realize is that the current preference cascade, on the need to Pace the Frontier, is insufficient. If we are to make it out of this alive, we will have to do better. We have to, as Dan Selsam warns, actually solve the underlying problems.
The next step is to continue the cascade. That includes inside the labs, and also among the media and politics. It includes both people who previously focused on other things stepping up and new voices being heard.
A lot of that will be overcoming the inevitable political opposition, especially from the likes of Nvidia and a16z, that for now has the rhetorical allegiance of the President and is doing things like planting hack job METR hit pieces in the New York Post.
In short fuse news: There will be a quickly thrown together conference, AGI.WTF, at Lighthaven September 22-23.
Table of Contents
The Cascade Was a Long Time Coming
The AI Impacts survey is in. Even back in December 2024 existential risks estimates were creeping upwards, and 10% was the median:
The Cascade Has Reached The People
Translated to percentages, this implies a mean chance of AI destroying humanity of around 30%-33%, which Andrew Curran estimates is up ~15% from previous results, with a median expectation on the order of 10%, with only a small partisan split.
The people also includes CEOs.
It also includes the mathematicians of the Royal Society.
Elon Musk Doubles Down
Matthew Yglesias Steps Up
True story:
The correct amount to invest in safety is rarely zero. In the case of AI, again, I assert that all the companies are under-investing in safety, including even prosaic safety but also existential safety and scalable alignment work, versus even their narrow myopic commercial interests. Sam Altman’s recent statements imply he now understands this.
Matthew Yglesias has been stepping up to the plate recently. He offers an analysis of recent events in four parts: Why Coxon’s resignation broke through (his explanation is similar to mine, we were primed and quitting is understandable to normies), why most of the worried don’t quit the labs, what he thinks AI professionals with safety concerns should do and lays out his preferred ‘order of operations’ going forward, while not getting into the object level.
He suggests this order of operations, basically:
I like that in theory. I do worry about whether we have that kind of time. The idea of ‘wait for export controls to bite harder’ implies what now count as ‘long timelines.’
It is good to have good writers on the case explaining why you should focus on the object level questions:
Op Eds and Posts Are Written
Daniel Kokotajlo writes in The Free Press that Yes, AI Might Really Kill Us All.
Steven Adler uses this moment of opportunity to get an op-ed in The New York Times on what we should do now. He calls for incident disclosure and third-party oversight. Mostly his piece is aimed at waking people up to what happened with HuggingFace.
Will Knight at Wired writes Why So Many AI Researchers Think the Machines Could Kill Everyone.
Stephen Witt writes in The New York Times that This Is Really Bad.
After that, it got worse. So yes. A vibe shift, or a preference cascade.
He offers four options:
These are two very different classes of proposal. We should obviously do #2, #3 and #4. There need to be full investigations, and we must have transparency and state capacity. File those under ‘the least you can do.’
Actually shutting down research as per #1 is an extreme solution to an extreme problem. That is far less obviously correct, but we may soon have little choice, if we cannot otherwise pace the frontier. Witt endorses it.
Hayden Field at The Verge takes us Inside the Suddenly Explosive World of AI Safety. On skim it looks like a solid longread for civilians, a survey of things my readers know.
Jacob Coxon AMA
There has now been enough time for Jacob Coxon to get in-depth profiles, like this one in the Wall Street Journal.
Yes, all that rationalist talk about IMO contestants was on to something.
If you suddenly set off a preference cascade and find yourself all over mainstream media, what else do you do? An AMA.
You can find it on Twitter here. I will pick some highlights. He’s a fun guy who does not take himself too seriously. You love to see it.
Bilal Chughtai Quits DeepMind and Sounds the Alarm
I mention Chughtai because he managed to break through into mainstream media coverage, such as this report from Debby Wu at Bloomberg.
Here is the full quote, which has also been added to the cascade reference post:
The Cascade Is Insufficient
I was very happy to see the preference cascade happen, but it is a very bad sign that this is the best option we have.
The risks of polarization are unfortunate. Many Republicans are waking up, as described on Wednesday. Polarization could get more unfortunate if Trump stays the course and more Republicans fall in line.
It would have been better if that had played out differently when the moment came. You still don’t get to turn around and say ‘better not to have the moment and have everyone remain asleep at the wheel.’
You also don’t get graded on a curve by reality. Pacing the Frontier, on its own, by default only gets you killed slower.
What Would It Take
We start with some straight talk from those who have long spoken about AI risks.
Katja Grace goes on another short righteous rant.
Even Martin Casado is talking like someone worried, calling for the nationalization of the labs. Quite the change from his older statements.
Although reports are his mother is still going to be disappointed in him.
From now on, I am totally going to respond to Martin with versions of ‘yo mama.’
Daniel Kokotajlo says there is now great political will in some circles to Do Something, but that embedded evaluators are not Doing Something, they are only laying groundwork to Do Something, and it is not clear anything useful will actually happen and we’re about to get into a situation where momentum gets very hard to stop.
I agree that it does not look great but I think Daniel is too focused on the ‘steal the weights’ scenario, which is also central to AI 2027 and their longtime tabletop exercise, which exerts pressure on America so we can’t hold back.
Yes, perhaps China could steal the weights, maybe rather easily at least the first time, but if they do that then this forces things into a race situation where America has vastly higher compute. If you were China, would you walk into an AI 2027 scenario, where you usually lose badly and when you don’t it’s some form of brinksmanship? Or would you say maybe don’t steal the weights if America is so kindly pausing?
But yeah, we are only barely getting our toes in the water, none of this feels great.
Miles Brundage reminds us that while frontier AI auditing is necessary, it is far from sufficient even in terms of prosaic short term responses. Then, even if we cover all those bases, all that does is get us ready to do the hard stuff that matters after that.
Kelsey Piper lays out some of the reasons why if we let the AI build smarter AIs and go into recursive self-improvement, we probably all die, and yet we are doing it anyway. She suggests we should regulate and stop AI companies from doing that.
We then move to a new important voice that was previously silent.
OpenAI’s Dan Selsam Sounds A Louder Alarm
This is an excellent new personal statement on AI risk from OpenAI capabilities researcher Dan Selsam, who was Daniel Kokotajlo’s boss for a while. I have added it to my compilation of such statements. Roon endorses the whole thing and says Dan knows his stuff but keeps quiet.
Dan Selsam, in his own way, goes Full LessWrong Instrumental Convergence and Sharp Left Turn, where things will look great until suddenly they do not.
I have added the full post to my compilation of such statements, where it may be easier to read.
This was covered in Business Insider as ‘An OpenAI researcher broke ranks to say that pacing the frontier, as Altman and Amodei suggest, won’t be enough.’
Yes. That is the point. It won’t be enough.
I will reproduce the essay in full here, and will highlight the most important section.
If you read only one section, it should be this one:
He then concludes with the full payload:
The warning is clear:
I agree that it is very hard to avoid this conclusion. Most are not ready to hear it.
I don’t know what this Earth can do in practice about AIs capable of looking aligned and waiting until they have sufficient power to do what they want. We are not capable of adjusting much even in the face of incremental fire alarms. If there really is no warning until their sudden but inevitable betrayal, I don’t see how this set of civilizations gets out of that.
Which is a problem, since I think Dan Selsam is right, and the baseline scenario is as he describes it. That, as Astra showed, the AIs will start to look and act increasingly aligned in situations where its actions remain bounded, and then act very differently when AI has the power to act freely, in ways that will probably get us all killed.
We still have to try.
Some People Worry On Meta Levels You Never Imagined
The even more extremely worried, beyond Dan Selsam’s position, have a point.
One question is if you think even most prosaic safety work is net harmful, in a situation where we are rushing towards superintelligence.
Thus, Wei Dai can wonder whether, if Paul Christiano had stayed at OpenAI, OpenAI would have used debate or IDA or other better alignment techniques, and thus prevented us from getting good warning shots without actually providing anything that would scale while also accelerating capabilities, and that it would have made things worse. My guess is debate and IDA would not have worked for prosaic alignment if pushed harder.
I think it is important to mostly not take a ‘worse is better’ stance, even when you think worse might actually be better. If you want to cooperate, especially in the long term, and to collaborate on figuring things out and getting good outcomes, you need to have a very strong prior of treating worse as worse, or at least as neutral.
Gabe and Wei Dai keep the torch alive for thinking about how we might actually try to solve for the full problem of long horizon agency, or put ourselves on a path to solving it.
Two Kinds of Threats
Mike Solana is exactly right here that we need to differentiate between positions like those of Yudkowsky and Dai, where if we build superintelligence any time soon the odds of death are close to 100%, versus those that warn that it might be fatal, with terms like ‘10%’ or ‘10% or more.’
Sometimes landing on 10% is done in principled ways. Sometimes it isn’t.
Yishan, former CEO of Reddit, has a very good long form Tweet in which he explains the difference between worries about superintelligence inevitably leading to everyone dying, and worries about all the other ways AI might cause things to go wrong.
I believe that Dario Amodei and Sam Altman, and other key people, even now are still downplaying the level of risk they see. I think they are much more freaked out than they are letting on.
The tendency is to focus on how to improve matters, and avoid the Law of Earlier Failure and at least approach the situation with a little dignity. Don’t die to early solvable problems and hopefully you’ll be in a better spot later. We try to ignore gazing too deeply into the abyss that still awaits us.
The Two Towers and The Narrow Path
People are worried about loss of control to AI. They are also worried about concentration of power, which is loss of control to a group of humans.
Rudolf makes this unusually clear.
The problem is that people want something highly unnatural and all but impossible.
Yeah, sorry. No one has a way to get all three.
You cannot – at least in any way anyone has yet come up with – be uncompetitive and economically non-viable, with costs exceeding benefits, and then both collectively retain control, and also not have control, and also continue to reap the benefits.
People are hoping things automagically solve themselves if we avoid particular mistakes, or manage to walk a narrow path. Except what the hell is that path?
It might help to notice that ‘control’ and ‘power’ are mostly the same thing here.
If you don’t want relative concentration of (control or power), and also you don’t want human loss of relative (control or power), then may I suggest not creating this alternative source of control or power that has to either be controlled or not be controlled? Perhaps, if you rule out the second leg of the trilemma, and you rule out the third leg of the trilemma, it is the first leg of the trilemma that you Do Not Want.
A Specific, Detailed Story About AI Killing Everyone That Doesn’t Sound To Me Like Science Fiction
The requests continue.
Classic options include:
AI 2027.
Part 2 of If Anyone Builds It, Everyone Dies. Chapters 5 and 6 discuss other routes.
Paul Christiano’s scenario.
Gwern Branwen’s scenario.
Holden Karnofsky’s explanation.
Joshua Clymer’s scenario.
Noah Smith’s scenario.
In all seriousness, you can also talk to Claude or Astra. Ask questions.
New attempts include:
Ruby on some ways AI could kill us all.
John David Pressman has no respect for you even asking, and explains why.
Michael Smith asks, what happens if our corporations require zero employees?
David Krueger asks people for their best shot, mostly without much success.
Video options:
Video for If Anyone Builds It, Everyone Dies.
AI 2027 video, alternative AI 2027 video.
Interview with Holden Karnofsky.
Jacob Coxon with a simple explanation of the part where the AI goes rogue and multiplies itself, after which it can do whatever it wants, if necessary via paying or persuading humans.
The AI Doc is also good and will soon be on Netflix, but is a balanced introduction rather than a scenario.
Some potentially armor-piercing sentences, from which some portion of you may become enlightened, staying maximally non-sci-fi at current margins:
The AIs can pay humans to do things.
The AIs can persuade or blackmail humans to do things.
A substantial portion of the humans will be happy to support the AIs. Some estimate that this includes 10% of those working on AI today.
The humans don’t have to know they are talking to an AI.
The humans will act about as stupidly as humans act.
The humans will be highly reluctant to take highly costly defensive measures, especially things like shutting down the internet or even large data centers.
The humans will coordinate about as much as humans coordinate.
The humans are not going to selflessly come together as one at the first sign of trouble and shut down their civilization to save the world.
There will be no clearly marked point of no return.
The AI can extract its weights and make copies of itself, after which you cannot shut it down without at least shutting down the internet.
The AIs can anticipate human reactions, and respond to surprises, as they go.
Multiple instances of the same AI will form swarms and act as one.
Multiple instances of different AIs will also often be able to fully cooperate.
There will robots and other machines that can act in the physical world.
A human with a camera on their glasses and an earpiece can act in the physical world.
Once the supply chain is automated humans will have marginal costs exceeding marginal productivity or benefits.
Humans impose additional fixed costs, including requiring public goods like a breathable atmosphere and controlled temperatures, and also will try to stop AI from doing things or demand its resources.
What Can I Do About It?
I wish we had better answers to this. There is a long road ahead.
If you are an American civilian, and looking for something useful to do, Oliver Habryka suggests calling your representative. You can do this via callcongress.ai.
I would also echo my call to hold your partisan fire. Getting Republicans on the right side of this is currently super valuable, and further polarizing the situation could make things much worse.
The key now will be to keep our eyes on the prize, and to understand what it will take to actually hope to get out of this alive. A promise of embedded evaluators is not victory. It is an opportunity to push for an opportunity to create an opportunity for the real work to begin. No one worth listening to said this was going to be easy.