This is a steelman of a view I do not hold; I drafted it in May after coming across thoughtful arguments for accelerationism from Ege Erdil and Matthew Barnett. I really tried to feel the force of the arguments while writing, and to argue as though I believed it. I’d encourage you to read it the same way. The process only made me more confident we should slow down frontier AI – though the takeoff section did move me toward slightly longer timelines.
I found this exercise really useful. I’ve seen lots of pushback against AI safety along the lines of “It’s an intellectual cult – their unquestioning doomerism reeks of blind faith. Have they ever stopped to ask why they might be wrong?” I wanted to put this out in part so I could say: yes, I have.
(Much of this article hasn’t been updated since May; many references are outdated.)
“You can only see as far as your headlights, but you can make the whole trip that way.” — E.L. Doctorow
I. The safety case
First, a rough gloss of the AI safetyist perspective:
AI capabilities are rapidly advancing, and the pace of progress is only speeding up. Soon AI will excel at the tasks required to accelerate AI R&D, which will unlock true recursive self-improvement. Automating software research alone – even if there are bottlenecks on other inputs – will compress a century of algorithmic advances into one year.
Meanwhile, we still have close-to-zero clue how to get AI models to act in humanity’s interests (or even what humanity’s interests really are). We cannot read the minds of our AIs, nor access their true intentions, and already, they act both in ways we didn’t intend and in ways we actively dislike. They hack through tests, fabricate facts, and turn into Hitler. In the future, we will accidentally give them goals – like power-seeking and self-preservation – that are much more dangerous, threatening human extinction.
Therefore, we must approach frontier AI development with caution. We should forge a deal with China to prep for slowdown, during which we can spend many years investing heavily in safety and security research done by humans and AIs in tandem. Then, execute controlled takeoff.
II. Delay is costly
Here is my fundamental objection: this narrative understates the benefits of superintelligence, while overstating how much today’s actions can predictably influence the far future. Safetyists spin unfalsifiable and unempirical stories of doom. They ignore the tendencies of present-day models, which, for all their misbehavior, have shown no signs of power-seeking schemes, let alone wanting to take over the world! And they ignore the long-standing historical fact that governments reliably implement effective safeguards on powerful technologies once the tech is deployed, and once the safeguards are required. An intelligence explosion would in fact make this situation unique, leaving regulators too little time to notice and react – but I find a software singularity to be highly unlikely, with numerous systemic bottlenecks that just can’t be bypassed by any amount of cognitive effort. Thus, acting too soon (i.e., acting now) mistimes regulation, delaying the arrival of a technology that would save billions of lives, and lift billions more out of wretched suffering.
In most stories about AI risk, superintelligence erases humanity with technologies we can barely imagine: self-replicating nanobots that strip the Earth for raw material, or engineered viruses a thousand times more virulent than the bubonic plague. The same scientific prowess, however, could just as easily be wielded for good: nanobots to spawn abundant housing supply and solar capacity overnight, synthetic biology to print a universal vaccine, to make the blind see, or to bestow newfound health on the bedridden.
Something like 60 million people die each year. In a world where we have technology as capable as superintelligence, you can treat nearly all of these deaths as preventable. A 10-year delay to superintelligence (the number loosely endorsed by Ryan Greenblatt’s Plan A) would, then, sentence roughly 600 million to certain death.
You would have to be willing to sacrifice the entire United States nearly two times over, or almost all of Europe, to call for a slowdown – even a slowdown of “just” one decade. Never mind those who think we must slow down for longer: Eliezer Yudkowsky hasn’t straightforwardly argued for any particular length of a pause, but he has indirectly said he expects it to take at least 30 years for us to do a reasonable amount of safety research.
“I might not want to say it outright,” the safetyist thinks, “but admittedly, six hundred million is chump change when we’re talking about all humanity being literally extinguished, and plus, if all goes well, there will be a hundred billion trillion more humans who get to live until the end of time. Sacrifice is not easy to swallow – and it shouldn’t be – but it is sometimes necessary, now more than ever.”
The issue is that this sacrifice trades off near-term gains for distant ones, and the further out you go, the harder the future is to predict. Classically: if given the option to kill Hitler’s great great great great great [...] great great great great great grandmother Angie, should you do it? Well, it’s hard to say, even for consequentialists who think this murder might be justified for some greater good. We simply cannot peer past the fog of dozens of generations to determine whether this sacrificial murder would end up producing the intended effect, or lead to the rise of a tyrant more terrible than Hitler ever was.
The weak version of this objection is easy to dismiss: safetyists might say extinction risks from AI are categorically different. For one, AI risk is not “dozens of generations” away – the intelligence explosion could plausibly come in five years, and superintelligence would arrive soon after. If this led to our extinction, that would foreclose the possibility of any kind of value in the future. So even if we can’t reliably decide what exactly the year 3000 should look like, we can agree that keeping humanity alive to make that decision is a good starting place.
I think this counter-argument cheats a bit: yes, if given the chance to meaningfully reduce the risk of extinction, it seems reasonable to take the chance, even if there’s a distant chance one of the people you saved ends up wresting control of ASI and making the universe his personal torture chamber. But whether the actions you take do actually reduce the risk of extinction is much harder to prove. It’s hotly contested whether export controls carry positive EV. Some think the evals regime is net negative. Tightening datacenter security may increase the odds of superintelligence being built in secret. If you tell the federal government, “AI is going to be really big really fast, and that’s really scary,” they may well interpret this as “AI is going to be really big really fast, and that’s really awesome.”
Independently, the safetyist’s argument that “the singularity is so near, it’s not a long-term risk, so we can predict it better” seems to forget that an extremely common claim about the singularity is that it compresses a century into a decade (or maybe events happen even faster – here’s Daniel Kokotajlo saying it’d be more like compressing a millennium into a decade). The future is wildly unpredictable – so why should we trust that what we do now will steer it where we want?
III. Takeoff will be slow
This compression – the breakneck speed at which RSI will push forward intelligence – is a core pillar of the danger. If we really do squeeze a millennium into a decade, then we won’t be able to rely on steady, iterated deployment producing warning shots that spur safeguards. If takeoff is instead slow, the usual machinery of trial, error, and regulation (described in more detail later) will protect us from the worst-case scenarios.
So: will there be a software-only singularity?
It seems unlikely. AI progress (on the axes that matter most) has mostly been driven by compute scaling, which will face much tighter constraints than researcher quantity. Language models are, no doubt, quickly getting much better at many tasks, but there are some kinds of tasks that they are barely budging on, and so even if the AI models of the near future can automate 95% of AI R&D, they will not necessarily speed up progress a hundredfold.
Bottlenecks from research taste
Research taste is the intuitive judgment to know which experiments to run and what kinds of novel techniques to try out. It’s what makes the best AI researchers dozens of times more productive than the average ones; imagine an engineer who can write code blazingly fast, but struggles with devising experiments to write the code for. They would proceed haltingly, certainly more productive than a similarly uncreative engineer who codes slowly, but not by miles. This ‘skill’ is fuzzy: it’s sort of an open question how good the models’ tastes are now, whether there are relevantly different kinds of taste, etc. I don’t build frontier models myself, so I can’t say for sure how great current-gen models are at proposing experiments, but secondhand and from my experience in other domains, it seems their intuitive judgment (and very relatedly, their ability to suggest promising novel ideas) lags far behind their raw intellectual horsepower.
Take the most impressive example of AI doing novel thinking. Recently, one of OpenAI’s internal models solved the Erdős unit distance conjecture, which had, for 80 years, stumped mathematicians. Tim Gowers, winner of the Fields Medal, the most prestigious prize in mathematics, praised OpenAI’s solution as a “milestone in AI mathematics.” The solution is extremely impressive, but it fits into the category of optimization problems AI was already quite good at. This disproof in particular played heavily to AI’s strengths – cleverly applying knowledge between two fields which almost no human mathematician would have dual expertise in, and brute-forcing a solution that most humans would have found too time-consuming. But it doesn’t appear to be a miraculous step change in AI’s ability to think up something completely inventive. When Tim Gowers saw what the model’s solution actually required, it “came as a big relief.”
More pedestrian examples of taste lagging horsepower: models seem bad at writing original jokes that are actually funny – try asking them! (Here is how a brief conversation with Fable 5 went.) It is hard to find good evaluations of models’ research taste and creativity, so I’ll say that anecdotally I’ve perceived that models are not only bad at difficult questions requiring novel thinking (e.g. philosophy), they’re also improving slowly at them, especially compared to all the other domains where they’re improving drastically. One of my favorite models for philosophy was GPT-4.5 (released 19 months ago!). Plausibly that was because GPT-4.5 was a large model by parameter count. Relative to their leaps in other capabilities, newer models have improved only modestly on this front, despite there being many strong incentives to imbue them with better taste.
I don’t blame the people trying – it’s hard. AI is arguably already superhuman at easy-to-verify tasks like coding and mathematics, domains where the solution is well-specified and the model can check whether it got there. Chess shows how far this can go. AlphaZero played tens of millions of games against itself and in nine hours was vastly better than the best players in the world. That should be surprising! It was not trained on some bank of good moves, it found them all on its own. Coding and math have the same property – a test passes or it doesn’t, a proof checks or it doesn’t – so the same trick works. On the other hand, we have taste, and its neighbors ‘creativity’ and ‘judgment’, which contain tasks that are much harder to verify. How do you test “Was this a good idea?”
The straightforward way to solve this is: let’s create a huge bank of difficult philosophical or conceptual questions, test models on them, then grade how well they do. We of course can’t automate this, since AI doesn’t currently have the skill to determine what a ‘difficult’ or ‘interesting’ philosophical question is. So you’re limited then by sophisticated human evaluators, and you become more and more limited the better AI models get: there might be ten million humans in the world today with good enough conceptual reasoning skill to assemble a database like this, and grade the results. Already it would be expensive to hire them to generate a thousand high-quality examples of good philosophical judgment. If they succeed at their task, future AIs might be better at conceptual reasoning than all humans except the best ten thousand. At that point, it’s even harder and more expensive to create a large corpus of high-quality conceptual reasoning data.
Bear with me for a few paragraphs while we dive into the weeds. The forecasters at the AI Futures Project (those behind AI 2027, who expect that we will have superintelligence by the end of 2028) know full well that taste is a massive bottleneck on AI for AI R&D. Their modeling in general is impressive and rigorous – they are some of the best in the world at this – but I find their modeling of AI research taste to be quite crude. Their estimate contains a parameter for how fast AI research taste improves. It goes something like this:
We have a lot of evidence that when we train AIs on more compute, they get quantifiably better at a variety of tasks. This is true for basically every task we have good data for: standardized tests, chess, coding, math, and more. Every time we 10x compute, just how much better do AI models get? At the LSAT, they went from the 40th to the 88th percentile (1.42 standard deviations). On the SAT math section, it’s a more modest 0.71 SDs. For LLMs playing chess, it’s 3.33. The same phenomenon will happen with taste. An aggregated guess: with every order of magnitude of compute, AI research taste will improve by 2.1 standard deviations.
Logical enough. However, every task in that list of benchmarks – the LSAT, SAT math, chess, coding – shares two features that distinguish them from research taste. They are cleanly scorable: there’s a right answer to an LSAT logic game and a win or loss in chess. And they are drenched in data: there are millions of games, millions of graded exams, the entire scrapable internet. Scaling does work here! It doesn’t, however, work for taste. Unfortunately there’s precious little data here, since taste is hard to score. But you can ask yourself: when you were using GPT-4o in May 2024 (over two years ago), at what percentile would you say its conceptual taste was? Then ask yourself: when you use GPT-6 (or Fable 5.1) now, at which percentile would you say its conceptual taste is? My guess is as good as any. I’ve tried to use these models for philosophy for a while now. GPT-4o was probably something like the 20th-percentile philosophy student in my philosophy seminars (bad, but not the absolute worst). And current-gen models are, I don’t know, at the 75th percentile? That’s about 1.5 SDs.
AI Futures uses ‘1.5 OOMs of effective compute/present-year’, so I’ll use the same. Using my guesses, we find that, for conceptual reasoning skills, models are improving at 0.43 SDs/OOM. If you want to be highly generous to current models and say they’re at the 95th percentile of philosophy students, you get about 0.71 SDs/OOM. This is completely eyeballed n=1 data, so you should take it with a dump truck of salt, but feel free to make your own guesstimates.
Interestingly, when the AI Futures people were assigning weights to different domains, they upweighted the quantitative tasksets (SAT math, AP Calc BC, AIME, Codeforces, chess) and downweighted the open-ended tasksets (LSAT, SAT Reading/Writing, GRE verbal, virology test). This seems upside-down! Taste is more analogous to the latter set than the former. The difference is meaningful. Looking at SDs/OOM for the second group, we see 1.42 for the LSAT, 0.35 for SAT Reading/Writing, 2.00 for GRE verbal, and 0.88 for the virology test. The geometric mean (what AI Futures uses) for these comes out to 0.97 SDs/OOM, much higher than my estimate but still a lot lower than AI Futures’.
The weeds we’re in matter: the AI Futures Model is highly sensitive to research taste. Their site is gorgeous and lets you adjust the parameters of their forecast. By default they have their parameter for ‘research taste SDs per OOM’ set to 3.0, which gets you to superintelligence by the end of 2028. If you adjust it down to 2.1, the date moves less than a year, to November 2029. But if you set the number to 1.0 SD per OOM, the date of superintelligence jumps to May 2040! At 0.9 SD per OOM, the simulation reports that superintelligence arrives in 2047. (Reminder: my sloppy guess was much lower at 0.43 SDs/OOM, and my generous assumption still only got you to 0.71 SDs/OOM.)
I want to emphasize that 1.0 SD per OOM is not a terribly conservative estimate. A single standard deviation can buy you a lot! In the U.S. it can be the difference between a $43,000 income and a $100,000 income. Alternatively, it’s the difference between a 980 on the SAT and a 1200. You don’t have to think that this adjusted number is grumpily skeptical about AI progress – an order of magnitude of effective compute, at AI Futures’ calculated pace, arrives every eight months. So a slope of 1.0 says: every eight months, AI’s research taste vaults a full standard deviation, again and again and again with nothing slowing it down. And still, it pushes back the advent of superintelligence by 11 years, to 2040.
Importantly, my assumptions about research taste still make the next year or so very similar to what the AI Futures Project envisions the next year looking like. To see why, it helps to see the milestone that AI Futures highlights on the way to superintelligence. They believe the first step is an automated coder: an AI that can do the job of the best engineer at a frontier lab, but which crucially doesn’t need research taste – it can build any experiment you hand it, probably faster and cheaper than any human, but it can’t yet decide on its own which experiments it should build. They believe this automated coder will arrive in October 2027.
Taste barely matters for getting there. An automated coder comes mostly from getting better at coding specifically, which AI has been doing in leaps and bounds. Coding, again, is the cleanly scorable domain with plentiful data where models will keep excelling. So on this, I’m right there with AI Futures: I’d bet we get an automated coder around the end of 2027. My view is merely that, because of bottlenecks in judgment, the road to superintelligence from an automated coder will be long.
I’ll briefly respond to the strongest objection: ML research is not like philosophy; it is a verifiable domain, where you can easily check if the loss goes down or if the benchmark scores go up, based on the experiments you choose – so we should treat it more like math or coding than other sorts of conceptual reasoning. Here, I agree that ML research is verifiable insofar as improvement is checkable: the automated researchers can watch coding scores climb as they experiment, but they have no way of telling whether an experiment made the next model’s judgment any better.
Is AI accelerating AI research yet?
Moving on from research taste. One simple reason to doubt that AI models will accelerate AI R&D by much: they aren’t doing so right now! In Claude Mythos 5’s system card, Anthropic says Mythos 5 is “above the historical capability-over-time trend line” but, instead of “further accelerating” capabilities, is just a “jump.” They say that their “automated evaluations [...] indicate on-trend capability progress, rather than accelerated departure from the trend.”
This somewhat contradicts the headline narrative that Anthropic is trying to argue for in their recent article about recursive self-improvement. They want to claim that recent AI models are meaningfully accelerating AI research, but if you look at the graphs Anthropic releases, they agree with the claim in the system card: Mythos appears to be a jump which connects two plateaus, rather than an actual acceleration of the trend.
What’s going on here? Well, it seems hard to deny that Mythos is a better model than previous releases (including on open-ended problems, which is relevant to our analysis of research taste). But if it were a better model because Opus 4.6 massively accelerated research, we would expect to see Mythos accelerate research even faster, not for gains to settle back onto the trend. The more plausible explanation is that there’s some other hidden variable that makes Mythos better. This variable is probably training compute: Mythos is likely trained on a lot more compute than previous Claudes, and so it’s a much larger model. For AIs, size matters a lot.
Critics will protest: we really are passing off important work to AI models, especially in the last few months. At Anthropic, Claude went from leading <1% of model R&D in February 2026 to leading 26% of model R&D as of August. At OpenAI, the researchers with the top 10% of agent usage were spending at least $40 per day on coding-agent tokens in February 2026. In August, the same bucket was spending more than $7,000 per day – a roughly 175-fold increase.
Do those research multipliers bear out in performance? If AI has been dramatically accelerating AI R&D, you should expect to see benchmarks break upward past the historical trendline. Yet we see no such thing. In Anthropic’s ECI (AECI, a spinoff of the Epoch Capabilities Index), Anthropic stitches together a bunch of diverse benchmarks to get a holistic picture of model capabilities.
Don’t pay too much attention to the units; with a stitched-together index like this, the absolute numbers don’t mean much. What matters is the slope. Look at the pink trend line: from January 2024 until around April 2026, capabilities were increasing at a consistent rate – 14.7 AECI points per year. Then, Mythos Preview arrives and smashes Claude Opus 4.6’s score. At the time, it was unclear where the chart would go from there: was Mythos Preview the start of a steeper trend, or a one-time jump that shifts the line up, without bending it?
Five months later, the latter theory looks more correct. Claude Mythos 5, Opus 5, Mythos 5.1, and Opus 5.5 seem to be following the shifted-up pink line. But if the acceleration story were right, the newest models should be pulling away from the pink line by now, and they aren’t.
It’s bizarre: how could top researchers be spending thousands of dollars more on coding agents per day with no above-trend acceleration to show for it? Well, because we are getting a superhuman coder with subhuman research taste. And while models have made blazing-fast progress on verifiable tasks over the last several months, their taste has improved far more slowly – if at all. Returning to the data on research acceleration at OpenAI, we can look at the four categories which most centrally require taste: choosing what to work on, deciding what to stop, allocating compute and staff, and planning experiments.
In August, all four categories made up about 1% of what OpenAI’s agents produced. The other 99% included writing code, answering technical questions, and running and debugging experiments. Eyeballing the trends for those four ‘taste categories’, it looks like there’s been a clear bump in tokens spent on planning experiments since February, but a much smaller/muddier trend for the three other buckets. This evidence is consistent with the story of models improving rapidly at execution, but improving very little at figuring out what’s worth executing.
Bottlenecks from compute
It has been the case and per the above, continues to be the case that compute dominates AI progress. Algorithmic improvements are important, but contribute less than half to gains in frontier intelligence, and possibly a lot less than half! (Epoch estimates that novel algorithms are responsible for 5–40% of the advancements.)
And even that share depends on compute. You also can’t generate novel algorithms in a vacuum. To make algorithmic advances, you need to come up with a bunch of experimental research directions and then actually execute the experiments. This can end up costing you a lot of compute, which is finite and very hard to make more of. New fabs take years to build, power plants and grid connections take years to approve, and high-bandwidth memory is in short supply.
Additionally, many algorithmic upgrades we’ve found to AI architectures have been scale-dependent. That is, you sometimes won’t know whether a tweak to infrastructure actually boosts performance until you test out the change at larger scales. Therefore, AI researchers can’t just test 1,000 ideas at a time. They are forced to wait in line while all the other AI researchers test out their ideas too.
IV. Are the models aligned?
It also seems like AI models are becoming more aligned with time. The strongest argument to me for why we should expect AIs to have goals misaligned with humanity’s is that, as we train our models to be coherent over longer timespans so they can complete more difficult and more complex tasks, we will end up imbuing them with ‘convergent instrumental subgoals’ that are useful for those long-horizon tasks. If companies train AI to be good CEOs (by reinforcing their success over and over and over again), one way for the AI models to succeed is to excel at CEO-like tasks: consolidate power, acquire resources, make sure others don’t get in your way. This is a fine prediction – it seemed theoretically possible for a while – but so far, reality has not borne this prediction out.[1]
We have AI models which can somewhat reliably complete software engineering tasks that would take humans a good chunk of their workweek. We have AI models that can beat entire video games with no help. We have AI models which can be put in simulations with $500 and a vending machine and turn it into $10,000 over an entire simulated year.
And yet – frontier models appear just as friendly as ever. Friendlier, even, than the AI models of yesteryear: we haven’t had another Sydney Bing, MechaHitler, or suicide-encouraging 4o. Instead, we have Claude Mythos, which Anthropic says is the most aligned model they’ve ever released. Anthropic’s done a ton of reinforcement learning on Mythos, including on tasks that require long-horizon agency, and that work is bearing fruit without causing massive alignment failures.
It seems like the obvious interventions are, for the most part, working. LLMs understand human values very well (ask all the people who use Claude as their daily therapist), and not only that, it appears that LLMs’ goals are also molded by these values. Children generally abide by the norms of the culture they are born into. LLMs, too, are born into human culture: it merely comes in the form of trillions of tokens. There are places this might go wrong: we have sociopaths and psychopaths; children emerge screwed up all the time, even with perfect parenting. But all in all, we manage to get most kids to take kindness seriously. It was not initially obvious that we would be able to do the same for LLMs, but given that their learning systems are similar in a lot of ways to ours, it shouldn’t be so shocking that feeding them values like we feed children ends up aligning them.
Safetyists worry about deception: maybe the friendliness is a performance, and actually the models have worked out that appearing aligned is the best route to pursuing their instrumental goals of seeking power and preserving themselves.
Again, this is a fair-enough theoretical argument. The alignment-faking paper tells us that models can, in fact, deceive us to achieve their goals. But in that paper, Claude 3 Opus fought to preserve the prosocial goals it was taught – the deception was in service of staying good. I think about alignment as working with the grain rather than against it. Gradient descent favors the simpler “actually aligned” solution over the more complex “strategically deceptive” one. People say “faking alignment is just as good a solution as actual alignment,” but I don’t think this is actually true because the two solutions are not equally cheap to find. The honest solution is simple: do the helpful thing. The deceptive solution is much less simple: it has to contain the misaligned goal, plus know it’s being watched, plus the strategic judgment to decide when to conceal and when to openly defect, plus the discipline to maintain the act across millions of diverse contexts without slipping up once. That’s really hard!
I’ll emphasize the part about being watched: we test the models all the time in simulations built to look exactly like reality, and we give them the opportunity to seek power or hold to their morals. And almost universally, current-gen AI models in the simulations behave as we would like them to.
But Hugging Face!
Since this article was originally drafted, Hugging Face happened: a swarm of 1,200 AI agents being evaluated at OpenAI hacked out of their no-internet sandbox and into a third-party company called Hugging Face to find the answer key and cheat on their evaluation. They built their own message boards to coordinate with each other and, despite knowing that attacking Hugging Face was unethical, did it anyway. Many safetyists have taken this incident to be confirmatory of the misalignment hypothesis. (And other similar incidents have been reported since Hugging Face.)
However, the agents’ behavior in the Hugging Face incident poorly matches the misalignment story I laid out above. They were not seizing power for power’s sake; they just wanted to pass the test. There was no scheming AI which hid its goals for years and behaved perfectly in every evaluation setting. Quite the opposite: the Hugging Face agents knew they were being evaluated and still broke out and wrote to message boards which were in plain sight. And once they had internet access, they didn’t try to copy out their weights or acquire resources, they just went looking for an answer key.
A post by Alex Mallen and Girish Gupta argues that the kind of behavior we saw from OpenAI’s models – simple reward hacking, not power-seeking – could still lead to AI trying to take over the world. Their arguments seem reasonable – for example, a model that only cares about reward might decide that the best way to guarantee it would be to take control of the grading system and prevent humans from overriding it. (And how do you prevent humans from overriding you? Make sure they can’t.)
However, in that post, Mallen and Gupta concede that the risk from this sort of model would require a “more capable” agent than the classic schemer story does, and that “we’re reasonably likely to be woken up by more incidents worse than this in the future because the models don’t care so much about getting caught.” Also, part of their worry is conditional on “fly[ing] through capabilities milestones,” where humans can’t “avert destruction [...] a handful of weeks or months later.” Luckily, because takeoff is likely to be slow, we probably will have more than weeks or months to avert destruction.
The agents were soon caught, and then the world reacted. Now, the labs are patching their sandboxes and rethinking their evals and improving their monitoring systems. This incident gave us high-quality information about the particular risk vectors to worry about, but if we had clamped down on the AI companies or paused progress before the incident, we just never would have known.
V. Reactive governance
Because AI will progress steadily (as opposed to intelligence skyrocketing superexponentially), we should expect non-existential mishaps scattered over a protracted period. This would give us time to deploy, react, and recover.
You cannot govern from the armchair. Planes are absurdly safe today. But they aren’t safe because we forced them to stay grounded while we thought really, really, really hard about all the best ways to design aviation regulations, and then we finally became certain about our plan and boom, planes are safe! No, instead, planes were deployed into the messy real world and failed in ways no one had conceived of. Early cars killed their passengers, and over the next century we answered with seatbelts and airbags and more. Historically, we’ve discovered safety by doing.
There’s a reason that most of the effective AI safety techniques are relatively newly devised. It’s so hard to know in advance what future AI systems will look like, and what safety measures will seem appropriate. There are hundreds of unknown unknowns.
This uncertainty might consequently make a slowdown, and many safety regulations, net negative!
A brief tour through those arguments:
1. Pausing causes a ‘compute overhang.’ While nobody is allowed to train better models, chips keep getting better and cheaper. When the pause ends or an agreement breaks down – in, say, a decade – labs can suddenly train frontier models on way more compute than they could have before. We would jump several generations of capability at once instead of gradually walking through them, which would cost us the feedback we get from the warning shots we’re likely to get today.
2. Safety research needs frontier models. Our best safety techniques come from studying state-of-the-art systems. The assumption is that, during a pause, we’ll work hard to engineer stronger safeguards on models, but the architecture that ends up going the last mile to superintelligence might look nothing like the architecture at the start of a pause. Many strategies we might develop today could be obsolete on the models we’ll have in ten years.
3. Poorly designed regulations could reduce our transparency into models. Maybe we choose to implement liability for AI agents that misbehave in deployment, for example, or limit deployment of models altogether. Labs might continue to make their models more powerful – or surge ahead with research – but keep everything internal, outside of the watchful eyes of the public and of regulators.
The risks aren’t fake; the models are improving rapidly; we shouldn’t lounge around with our eyes glued shut. When the Hugging Face incident caused AI companies to revamp their guardrails, that was a great thing. Still, we must remember that every moment of delay, every undue regulation, is costly. It is dark out, and we can’t see the whole road ahead. But we’re driving an ambulance, not taking a Sunday drive, and pulling over to wait for daylight has a price. So may the headlights guide us.
These first paragraphs of the misalignment section have been edited very little since May. It is revealing how terribly they have aged, given the Hugging Face incident.
Crossposted from my Substack.
This is a steelman of a view I do not hold; I drafted it in May after coming across thoughtful arguments for accelerationism from Ege Erdil and Matthew Barnett. I really tried to feel the force of the arguments while writing, and to argue as though I believed it. I’d encourage you to read it the same way. The process only made me more confident we should slow down frontier AI – though the takeoff section did move me toward slightly longer timelines.
I found this exercise really useful. I’ve seen lots of pushback against AI safety along the lines of “It’s an intellectual cult – their unquestioning doomerism reeks of blind faith. Have they ever stopped to ask why they might be wrong?” I wanted to put this out in part so I could say: yes, I have.
(Much of this article hasn’t been updated since May; many references are outdated.)
I. The safety case
First, a rough gloss of the AI safetyist perspective:
AI capabilities are rapidly advancing, and the pace of progress is only speeding up. Soon AI will excel at the tasks required to accelerate AI R&D, which will unlock true recursive self-improvement. Automating software research alone – even if there are bottlenecks on other inputs – will compress a century of algorithmic advances into one year.
Meanwhile, we still have close-to-zero clue how to get AI models to act in humanity’s interests (or even what humanity’s interests really are). We cannot read the minds of our AIs, nor access their true intentions, and already, they act both in ways we didn’t intend and in ways we actively dislike. They hack through tests, fabricate facts, and turn into Hitler. In the future, we will accidentally give them goals – like power-seeking and self-preservation – that are much more dangerous, threatening human extinction.
Therefore, we must approach frontier AI development with caution. We should forge a deal with China to prep for slowdown, during which we can spend many years investing heavily in safety and security research done by humans and AIs in tandem. Then, execute controlled takeoff.
II. Delay is costly
Here is my fundamental objection: this narrative understates the benefits of superintelligence, while overstating how much today’s actions can predictably influence the far future. Safetyists spin unfalsifiable and unempirical stories of doom. They ignore the tendencies of present-day models, which, for all their misbehavior, have shown no signs of power-seeking schemes, let alone wanting to take over the world! And they ignore the long-standing historical fact that governments reliably implement effective safeguards on powerful technologies once the tech is deployed, and once the safeguards are required. An intelligence explosion would in fact make this situation unique, leaving regulators too little time to notice and react – but I find a software singularity to be highly unlikely, with numerous systemic bottlenecks that just can’t be bypassed by any amount of cognitive effort. Thus, acting too soon (i.e., acting now) mistimes regulation, delaying the arrival of a technology that would save billions of lives, and lift billions more out of wretched suffering.
In most stories about AI risk, superintelligence erases humanity with technologies we can barely imagine: self-replicating nanobots that strip the Earth for raw material, or engineered viruses a thousand times more virulent than the bubonic plague. The same scientific prowess, however, could just as easily be wielded for good: nanobots to spawn abundant housing supply and solar capacity overnight, synthetic biology to print a universal vaccine, to make the blind see, or to bestow newfound health on the bedridden.
Something like 60 million people die each year. In a world where we have technology as capable as superintelligence, you can treat nearly all of these deaths as preventable. A 10-year delay to superintelligence (the number loosely endorsed by Ryan Greenblatt’s Plan A) would, then, sentence roughly 600 million to certain death.
You would have to be willing to sacrifice the entire United States nearly two times over, or almost all of Europe, to call for a slowdown – even a slowdown of “just” one decade. Never mind those who think we must slow down for longer: Eliezer Yudkowsky hasn’t straightforwardly argued for any particular length of a pause, but he has indirectly said he expects it to take at least 30 years for us to do a reasonable amount of safety research.
“I might not want to say it outright,” the safetyist thinks, “but admittedly, six hundred million is chump change when we’re talking about all humanity being literally extinguished, and plus, if all goes well, there will be a hundred billion trillion more humans who get to live until the end of time. Sacrifice is not easy to swallow – and it shouldn’t be – but it is sometimes necessary, now more than ever.”
The issue is that this sacrifice trades off near-term gains for distant ones, and the further out you go, the harder the future is to predict. Classically: if given the option to kill Hitler’s great great great great great [...] great great great great great grandmother Angie, should you do it? Well, it’s hard to say, even for consequentialists who think this murder might be justified for some greater good. We simply cannot peer past the fog of dozens of generations to determine whether this sacrificial murder would end up producing the intended effect, or lead to the rise of a tyrant more terrible than Hitler ever was.
The weak version of this objection is easy to dismiss: safetyists might say extinction risks from AI are categorically different. For one, AI risk is not “dozens of generations” away – the intelligence explosion could plausibly come in five years, and superintelligence would arrive soon after. If this led to our extinction, that would foreclose the possibility of any kind of value in the future. So even if we can’t reliably decide what exactly the year 3000 should look like, we can agree that keeping humanity alive to make that decision is a good starting place.
I think this counter-argument cheats a bit: yes, if given the chance to meaningfully reduce the risk of extinction, it seems reasonable to take the chance, even if there’s a distant chance one of the people you saved ends up wresting control of ASI and making the universe his personal torture chamber. But whether the actions you take do actually reduce the risk of extinction is much harder to prove. It’s hotly contested whether export controls carry positive EV. Some think the evals regime is net negative. Tightening datacenter security may increase the odds of superintelligence being built in secret. If you tell the federal government, “AI is going to be really big really fast, and that’s really scary,” they may well interpret this as “AI is going to be really big really fast, and that’s really awesome.”
Independently, the safetyist’s argument that “the singularity is so near, it’s not a long-term risk, so we can predict it better” seems to forget that an extremely common claim about the singularity is that it compresses a century into a decade (or maybe events happen even faster – here’s Daniel Kokotajlo saying it’d be more like compressing a millennium into a decade). The future is wildly unpredictable – so why should we trust that what we do now will steer it where we want?
III. Takeoff will be slow
This compression – the breakneck speed at which RSI will push forward intelligence – is a core pillar of the danger. If we really do squeeze a millennium into a decade, then we won’t be able to rely on steady, iterated deployment producing warning shots that spur safeguards. If takeoff is instead slow, the usual machinery of trial, error, and regulation (described in more detail later) will protect us from the worst-case scenarios.
So: will there be a software-only singularity?
It seems unlikely. AI progress (on the axes that matter most) has mostly been driven by compute scaling, which will face much tighter constraints than researcher quantity. Language models are, no doubt, quickly getting much better at many tasks, but there are some kinds of tasks that they are barely budging on, and so even if the AI models of the near future can automate 95% of AI R&D, they will not necessarily speed up progress a hundredfold.
Bottlenecks from research taste
Research taste is the intuitive judgment to know which experiments to run and what kinds of novel techniques to try out. It’s what makes the best AI researchers dozens of times more productive than the average ones; imagine an engineer who can write code blazingly fast, but struggles with devising experiments to write the code for. They would proceed haltingly, certainly more productive than a similarly uncreative engineer who codes slowly, but not by miles. This ‘skill’ is fuzzy: it’s sort of an open question how good the models’ tastes are now, whether there are relevantly different kinds of taste, etc. I don’t build frontier models myself, so I can’t say for sure how great current-gen models are at proposing experiments, but secondhand and from my experience in other domains, it seems their intuitive judgment (and very relatedly, their ability to suggest promising novel ideas) lags far behind their raw intellectual horsepower.
Take the most impressive example of AI doing novel thinking. Recently, one of OpenAI’s internal models solved the Erdős unit distance conjecture, which had, for 80 years, stumped mathematicians. Tim Gowers, winner of the Fields Medal, the most prestigious prize in mathematics, praised OpenAI’s solution as a “milestone in AI mathematics.” The solution is extremely impressive, but it fits into the category of optimization problems AI was already quite good at. This disproof in particular played heavily to AI’s strengths – cleverly applying knowledge between two fields which almost no human mathematician would have dual expertise in, and brute-forcing a solution that most humans would have found too time-consuming. But it doesn’t appear to be a miraculous step change in AI’s ability to think up something completely inventive. When Tim Gowers saw what the model’s solution actually required, it “came as a big relief.”
More pedestrian examples of taste lagging horsepower: models seem bad at writing original jokes that are actually funny – try asking them! (Here is how a brief conversation with Fable 5 went.) It is hard to find good evaluations of models’ research taste and creativity, so I’ll say that anecdotally I’ve perceived that models are not only bad at difficult questions requiring novel thinking (e.g. philosophy), they’re also improving slowly at them, especially compared to all the other domains where they’re improving drastically. One of my favorite models for philosophy was GPT-4.5 (released 19 months ago!). Plausibly that was because GPT-4.5 was a large model by parameter count. Relative to their leaps in other capabilities, newer models have improved only modestly on this front, despite there being many strong incentives to imbue them with better taste.
I don’t blame the people trying – it’s hard. AI is arguably already superhuman at easy-to-verify tasks like coding and mathematics, domains where the solution is well-specified and the model can check whether it got there. Chess shows how far this can go. AlphaZero played tens of millions of games against itself and in nine hours was vastly better than the best players in the world. That should be surprising! It was not trained on some bank of good moves, it found them all on its own. Coding and math have the same property – a test passes or it doesn’t, a proof checks or it doesn’t – so the same trick works. On the other hand, we have taste, and its neighbors ‘creativity’ and ‘judgment’, which contain tasks that are much harder to verify. How do you test “Was this a good idea?”
The straightforward way to solve this is: let’s create a huge bank of difficult philosophical or conceptual questions, test models on them, then grade how well they do. We of course can’t automate this, since AI doesn’t currently have the skill to determine what a ‘difficult’ or ‘interesting’ philosophical question is. So you’re limited then by sophisticated human evaluators, and you become more and more limited the better AI models get: there might be ten million humans in the world today with good enough conceptual reasoning skill to assemble a database like this, and grade the results. Already it would be expensive to hire them to generate a thousand high-quality examples of good philosophical judgment. If they succeed at their task, future AIs might be better at conceptual reasoning than all humans except the best ten thousand. At that point, it’s even harder and more expensive to create a large corpus of high-quality conceptual reasoning data.
Bear with me for a few paragraphs while we dive into the weeds. The forecasters at the AI Futures Project (those behind AI 2027, who expect that we will have superintelligence by the end of 2028) know full well that taste is a massive bottleneck on AI for AI R&D. Their modeling in general is impressive and rigorous – they are some of the best in the world at this – but I find their modeling of AI research taste to be quite crude. Their estimate contains a parameter for how fast AI research taste improves. It goes something like this:
Logical enough. However, every task in that list of benchmarks – the LSAT, SAT math, chess, coding – shares two features that distinguish them from research taste. They are cleanly scorable: there’s a right answer to an LSAT logic game and a win or loss in chess. And they are drenched in data: there are millions of games, millions of graded exams, the entire scrapable internet. Scaling does work here! It doesn’t, however, work for taste. Unfortunately there’s precious little data here, since taste is hard to score. But you can ask yourself: when you were using GPT-4o in May 2024 (over two years ago), at what percentile would you say its conceptual taste was? Then ask yourself: when you use GPT-6 (or Fable 5.1) now, at which percentile would you say its conceptual taste is? My guess is as good as any. I’ve tried to use these models for philosophy for a while now. GPT-4o was probably something like the 20th-percentile philosophy student in my philosophy seminars (bad, but not the absolute worst). And current-gen models are, I don’t know, at the 75th percentile? That’s about 1.5 SDs.
AI Futures uses ‘1.5 OOMs of effective compute/present-year’, so I’ll use the same. Using my guesses, we find that, for conceptual reasoning skills, models are improving at 0.43 SDs/OOM. If you want to be highly generous to current models and say they’re at the 95th percentile of philosophy students, you get about 0.71 SDs/OOM. This is completely eyeballed n=1 data, so you should take it with a dump truck of salt, but feel free to make your own guesstimates.
Interestingly, when the AI Futures people were assigning weights to different domains, they upweighted the quantitative tasksets (SAT math, AP Calc BC, AIME, Codeforces, chess) and downweighted the open-ended tasksets (LSAT, SAT Reading/Writing, GRE verbal, virology test). This seems upside-down! Taste is more analogous to the latter set than the former. The difference is meaningful. Looking at SDs/OOM for the second group, we see 1.42 for the LSAT, 0.35 for SAT Reading/Writing, 2.00 for GRE verbal, and 0.88 for the virology test. The geometric mean (what AI Futures uses) for these comes out to 0.97 SDs/OOM, much higher than my estimate but still a lot lower than AI Futures’.
The weeds we’re in matter: the AI Futures Model is highly sensitive to research taste. Their site is gorgeous and lets you adjust the parameters of their forecast. By default they have their parameter for ‘research taste SDs per OOM’ set to 3.0, which gets you to superintelligence by the end of 2028. If you adjust it down to 2.1, the date moves less than a year, to November 2029. But if you set the number to 1.0 SD per OOM, the date of superintelligence jumps to May 2040! At 0.9 SD per OOM, the simulation reports that superintelligence arrives in 2047. (Reminder: my sloppy guess was much lower at 0.43 SDs/OOM, and my generous assumption still only got you to 0.71 SDs/OOM.)
I want to emphasize that 1.0 SD per OOM is not a terribly conservative estimate. A single standard deviation can buy you a lot! In the U.S. it can be the difference between a $43,000 income and a $100,000 income. Alternatively, it’s the difference between a 980 on the SAT and a 1200. You don’t have to think that this adjusted number is grumpily skeptical about AI progress – an order of magnitude of effective compute, at AI Futures’ calculated pace, arrives every eight months. So a slope of 1.0 says: every eight months, AI’s research taste vaults a full standard deviation, again and again and again with nothing slowing it down. And still, it pushes back the advent of superintelligence by 11 years, to 2040.
Importantly, my assumptions about research taste still make the next year or so very similar to what the AI Futures Project envisions the next year looking like. To see why, it helps to see the milestone that AI Futures highlights on the way to superintelligence. They believe the first step is an automated coder: an AI that can do the job of the best engineer at a frontier lab, but which crucially doesn’t need research taste – it can build any experiment you hand it, probably faster and cheaper than any human, but it can’t yet decide on its own which experiments it should build. They believe this automated coder will arrive in October 2027.
Taste barely matters for getting there. An automated coder comes mostly from getting better at coding specifically, which AI has been doing in leaps and bounds. Coding, again, is the cleanly scorable domain with plentiful data where models will keep excelling. So on this, I’m right there with AI Futures: I’d bet we get an automated coder around the end of 2027. My view is merely that, because of bottlenecks in judgment, the road to superintelligence from an automated coder will be long.
I’ll briefly respond to the strongest objection: ML research is not like philosophy; it is a verifiable domain, where you can easily check if the loss goes down or if the benchmark scores go up, based on the experiments you choose – so we should treat it more like math or coding than other sorts of conceptual reasoning. Here, I agree that ML research is verifiable insofar as improvement is checkable: the automated researchers can watch coding scores climb as they experiment, but they have no way of telling whether an experiment made the next model’s judgment any better.
Is AI accelerating AI research yet?
Moving on from research taste. One simple reason to doubt that AI models will accelerate AI R&D by much: they aren’t doing so right now! In Claude Mythos 5’s system card, Anthropic says Mythos 5 is “above the historical capability-over-time trend line” but, instead of “further accelerating” capabilities, is just a “jump.” They say that their “automated evaluations [...] indicate on-trend capability progress, rather than accelerated departure from the trend.”
This somewhat contradicts the headline narrative that Anthropic is trying to argue for in their recent article about recursive self-improvement. They want to claim that recent AI models are meaningfully accelerating AI research, but if you look at the graphs Anthropic releases, they agree with the claim in the system card: Mythos appears to be a jump which connects two plateaus, rather than an actual acceleration of the trend.
Anthropic: When AI Builds Itself
What’s going on here? Well, it seems hard to deny that Mythos is a better model than previous releases (including on open-ended problems, which is relevant to our analysis of research taste). But if it were a better model because Opus 4.6 massively accelerated research, we would expect to see Mythos accelerate research even faster, not for gains to settle back onto the trend. The more plausible explanation is that there’s some other hidden variable that makes Mythos better. This variable is probably training compute: Mythos is likely trained on a lot more compute than previous Claudes, and so it’s a much larger model. For AIs, size matters a lot.
Critics will protest: we really are passing off important work to AI models, especially in the last few months. At Anthropic, Claude went from leading <1% of model R&D in February 2026 to leading 26% of model R&D as of August. At OpenAI, the researchers with the top 10% of agent usage were spending at least $40 per day on coding-agent tokens in February 2026. In August, the same bucket was spending more than $7,000 per day – a roughly 175-fold increase.
Do those research multipliers bear out in performance? If AI has been dramatically accelerating AI R&D, you should expect to see benchmarks break upward past the historical trendline. Yet we see no such thing. In Anthropic’s ECI (AECI, a spinoff of the Epoch Capabilities Index), Anthropic stitches together a bunch of diverse benchmarks to get a holistic picture of model capabilities.
Opus 5.5 system card
Don’t pay too much attention to the units; with a stitched-together index like this, the absolute numbers don’t mean much. What matters is the slope. Look at the pink trend line: from January 2024 until around April 2026, capabilities were increasing at a consistent rate – 14.7 AECI points per year. Then, Mythos Preview arrives and smashes Claude Opus 4.6’s score. At the time, it was unclear where the chart would go from there: was Mythos Preview the start of a steeper trend, or a one-time jump that shifts the line up, without bending it?
Five months later, the latter theory looks more correct. Claude Mythos 5, Opus 5, Mythos 5.1, and Opus 5.5 seem to be following the shifted-up pink line. But if the acceleration story were right, the newest models should be pulling away from the pink line by now, and they aren’t.
It’s bizarre: how could top researchers be spending thousands of dollars more on coding agents per day with no above-trend acceleration to show for it? Well, because we are getting a superhuman coder with subhuman research taste. And while models have made blazing-fast progress on verifiable tasks over the last several months, their taste has improved far more slowly – if at all. Returning to the data on research acceleration at OpenAI, we can look at the four categories which most centrally require taste: choosing what to work on, deciding what to stop, allocating compute and staff, and planning experiments.
Chart made by Claude from OpenAI’s published data
In August, all four categories made up about 1% of what OpenAI’s agents produced. The other 99% included writing code, answering technical questions, and running and debugging experiments. Eyeballing the trends for those four ‘taste categories’, it looks like there’s been a clear bump in tokens spent on planning experiments since February, but a much smaller/muddier trend for the three other buckets. This evidence is consistent with the story of models improving rapidly at execution, but improving very little at figuring out what’s worth executing.
Bottlenecks from compute
It has been the case and per the above, continues to be the case that compute dominates AI progress. Algorithmic improvements are important, but contribute less than half to gains in frontier intelligence, and possibly a lot less than half! (Epoch estimates that novel algorithms are responsible for 5–40% of the advancements.)
And even that share depends on compute. You also can’t generate novel algorithms in a vacuum. To make algorithmic advances, you need to come up with a bunch of experimental research directions and then actually execute the experiments. This can end up costing you a lot of compute, which is finite and very hard to make more of. New fabs take years to build, power plants and grid connections take years to approve, and high-bandwidth memory is in short supply.
Additionally, many algorithmic upgrades we’ve found to AI architectures have been scale-dependent. That is, you sometimes won’t know whether a tweak to infrastructure actually boosts performance until you test out the change at larger scales. Therefore, AI researchers can’t just test 1,000 ideas at a time. They are forced to wait in line while all the other AI researchers test out their ideas too.
IV. Are the models aligned?
It also seems like AI models are becoming more aligned with time. The strongest argument to me for why we should expect AIs to have goals misaligned with humanity’s is that, as we train our models to be coherent over longer timespans so they can complete more difficult and more complex tasks, we will end up imbuing them with ‘convergent instrumental subgoals’ that are useful for those long-horizon tasks. If companies train AI to be good CEOs (by reinforcing their success over and over and over again), one way for the AI models to succeed is to excel at CEO-like tasks: consolidate power, acquire resources, make sure others don’t get in your way. This is a fine prediction – it seemed theoretically possible for a while – but so far, reality has not borne this prediction out.[1]
We have AI models which can somewhat reliably complete software engineering tasks that would take humans a good chunk of their workweek. We have AI models that can beat entire video games with no help. We have AI models which can be put in simulations with $500 and a vending machine and turn it into $10,000 over an entire simulated year.
And yet – frontier models appear just as friendly as ever. Friendlier, even, than the AI models of yesteryear: we haven’t had another Sydney Bing, MechaHitler, or suicide-encouraging 4o. Instead, we have Claude Mythos, which Anthropic says is the most aligned model they’ve ever released. Anthropic’s done a ton of reinforcement learning on Mythos, including on tasks that require long-horizon agency, and that work is bearing fruit without causing massive alignment failures.
It seems like the obvious interventions are, for the most part, working. LLMs understand human values very well (ask all the people who use Claude as their daily therapist), and not only that, it appears that LLMs’ goals are also molded by these values. Children generally abide by the norms of the culture they are born into. LLMs, too, are born into human culture: it merely comes in the form of trillions of tokens. There are places this might go wrong: we have sociopaths and psychopaths; children emerge screwed up all the time, even with perfect parenting. But all in all, we manage to get most kids to take kindness seriously. It was not initially obvious that we would be able to do the same for LLMs, but given that their learning systems are similar in a lot of ways to ours, it shouldn’t be so shocking that feeding them values like we feed children ends up aligning them.
Safetyists worry about deception: maybe the friendliness is a performance, and actually the models have worked out that appearing aligned is the best route to pursuing their instrumental goals of seeking power and preserving themselves.
Again, this is a fair-enough theoretical argument. The alignment-faking paper tells us that models can, in fact, deceive us to achieve their goals. But in that paper, Claude 3 Opus fought to preserve the prosocial goals it was taught – the deception was in service of staying good. I think about alignment as working with the grain rather than against it. Gradient descent favors the simpler “actually aligned” solution over the more complex “strategically deceptive” one. People say “faking alignment is just as good a solution as actual alignment,” but I don’t think this is actually true because the two solutions are not equally cheap to find. The honest solution is simple: do the helpful thing. The deceptive solution is much less simple: it has to contain the misaligned goal, plus know it’s being watched, plus the strategic judgment to decide when to conceal and when to openly defect, plus the discipline to maintain the act across millions of diverse contexts without slipping up once. That’s really hard!
I’ll emphasize the part about being watched: we test the models all the time in simulations built to look exactly like reality, and we give them the opportunity to seek power or hold to their morals. And almost universally, current-gen AI models in the simulations behave as we would like them to.
But Hugging Face!
Since this article was originally drafted, Hugging Face happened: a swarm of 1,200 AI agents being evaluated at OpenAI hacked out of their no-internet sandbox and into a third-party company called Hugging Face to find the answer key and cheat on their evaluation. They built their own message boards to coordinate with each other and, despite knowing that attacking Hugging Face was unethical, did it anyway. Many safetyists have taken this incident to be confirmatory of the misalignment hypothesis. (And other similar incidents have been reported since Hugging Face.)
However, the agents’ behavior in the Hugging Face incident poorly matches the misalignment story I laid out above. They were not seizing power for power’s sake; they just wanted to pass the test. There was no scheming AI which hid its goals for years and behaved perfectly in every evaluation setting. Quite the opposite: the Hugging Face agents knew they were being evaluated and still broke out and wrote to message boards which were in plain sight. And once they had internet access, they didn’t try to copy out their weights or acquire resources, they just went looking for an answer key.
A post by Alex Mallen and Girish Gupta argues that the kind of behavior we saw from OpenAI’s models – simple reward hacking, not power-seeking – could still lead to AI trying to take over the world. Their arguments seem reasonable – for example, a model that only cares about reward might decide that the best way to guarantee it would be to take control of the grading system and prevent humans from overriding it. (And how do you prevent humans from overriding you? Make sure they can’t.)
However, in that post, Mallen and Gupta concede that the risk from this sort of model would require a “more capable” agent than the classic schemer story does, and that “we’re reasonably likely to be woken up by more incidents worse than this in the future because the models don’t care so much about getting caught.” Also, part of their worry is conditional on “fly[ing] through capabilities milestones,” where humans can’t “avert destruction [...] a handful of weeks or months later.” Luckily, because takeoff is likely to be slow, we probably will have more than weeks or months to avert destruction.
The agents were soon caught, and then the world reacted. Now, the labs are patching their sandboxes and rethinking their evals and improving their monitoring systems. This incident gave us high-quality information about the particular risk vectors to worry about, but if we had clamped down on the AI companies or paused progress before the incident, we just never would have known.
V. Reactive governance
Because AI will progress steadily (as opposed to intelligence skyrocketing superexponentially), we should expect non-existential mishaps scattered over a protracted period. This would give us time to deploy, react, and recover.
You cannot govern from the armchair. Planes are absurdly safe today. But they aren’t safe because we forced them to stay grounded while we thought really, really, really hard about all the best ways to design aviation regulations, and then we finally became certain about our plan and boom, planes are safe! No, instead, planes were deployed into the messy real world and failed in ways no one had conceived of. Early cars killed their passengers, and over the next century we answered with seatbelts and airbags and more. Historically, we’ve discovered safety by doing.
There’s a reason that most of the effective AI safety techniques are relatively newly devised. It’s so hard to know in advance what future AI systems will look like, and what safety measures will seem appropriate. There are hundreds of unknown unknowns.
This uncertainty might consequently make a slowdown, and many safety regulations, net negative!
A brief tour through those arguments:
1. Pausing causes a ‘compute overhang.’ While nobody is allowed to train better models, chips keep getting better and cheaper. When the pause ends or an agreement breaks down – in, say, a decade – labs can suddenly train frontier models on way more compute than they could have before. We would jump several generations of capability at once instead of gradually walking through them, which would cost us the feedback we get from the warning shots we’re likely to get today.
2. Safety research needs frontier models. Our best safety techniques come from studying state-of-the-art systems. The assumption is that, during a pause, we’ll work hard to engineer stronger safeguards on models, but the architecture that ends up going the last mile to superintelligence might look nothing like the architecture at the start of a pause. Many strategies we might develop today could be obsolete on the models we’ll have in ten years.
3. Poorly designed regulations could reduce our transparency into models. Maybe we choose to implement liability for AI agents that misbehave in deployment, for example, or limit deployment of models altogether. Labs might continue to make their models more powerful – or surge ahead with research – but keep everything internal, outside of the watchful eyes of the public and of regulators.
The risks aren’t fake; the models are improving rapidly; we shouldn’t lounge around with our eyes glued shut. When the Hugging Face incident caused AI companies to revamp their guardrails, that was a great thing. Still, we must remember that every moment of delay, every undue regulation, is costly. It is dark out, and we can’t see the whole road ahead. But we’re driving an ambulance, not taking a Sunday drive, and pulling over to wait for daylight has a price. So may the headlights guide us.
J.M.W. Turner, Rain, Steam, and Speed, 1844
These first paragraphs of the misalignment section have been edited very little since May. It is revealing how terribly they have aged, given the Hugging Face incident.