I agree with @Marcus Plutowski's answer, with an addition: If we unpause too late, we also die.
I'm not sure of that statement, but Andrej Karpathy has asserted that he believes a 'cognitive core' model, stripped of memorized knowledge but capable of all forms of reasoning, should be possible in ~1-10B parameters. I've heard others throw around the idea that in principle there should be a python file <1M tokens in size sufficient to kick off RSI. EY said, a long time ago, that someday we may find it is possible to run an AGI on the equivalent of a 1995 laptop. How Full Butlerian Jihad do we imagine a Pause could be, in practice? Because if the answer is not 100%, then we're buying time and talking price.
I'm not sure of that statement, but Andrej Karpathy has asserted that he believes a 'cognitive core' model, stripped of memorized knowledge but capable of all forms of reasoning, should be possible in ~1-10B parameters.
Two decades ago, I would have guessed you needed at least the equivalent of an RTX Pro 6000 to run AGI at near human speed. But now we have models like Qwen3.8 27B, which are highly capable and persistent coding agents, and it has remarkably wide knowledge of the world. And you don't even need an RTX Pro 6000. It's obviously not a fully general human replacement, but it's a surprisingly large chunk of one. So if you're willing to drop a bunch of world knowledge and add the missing capabilities, yeah, there's probably some smallish constant multiplier of 27B that gets you there.
So I wouldn't be entirely surprised if there exists a recipe for AGI (or even weakly superhuman AI) that would fit in a 2030 homelab. If this exists, it would be a recipe for ruin: a potentially world-ending technology that could be hosted by a single hobbyist.
This is the other reason [1] I would like to try to halt AI progress: If it's possible to explain the key ideas needed to built an human-replacement AGI with 1 or 2 scientific papers, I would prefer to delay the writing of those papers for as many years as possible.
My first reason to halt is that recent frontier math results suggest we might be getting uncomfortably close to RSI. ↩︎
As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone.
My fear is that alignment may not actually be possible in any sufficiently robust sense. (That's not a great post, and I need to write up a better form of that argument. Basically: Intelligence is inscrutable matrices, and the pre-conditions for natural selection are really easy to meet. So you can't really understand or control the AIs in any strong sense, and they operate under selective pressures that push them away from alignment.)
So I imagine several potential scenarios:
My metaphor here is a nasty late-stage cancer. It's probably incurable, in the long run, without almost miraculous luck. But if we start chemotherapy soon, we might be able to buy everyone some more years. Every decade we buy is 80 billion years of human life. Seems worth fighting for!
So I kind of wish slightly fewer people would spend their time on pie-in-the-sky plans for making 100 trillion parameter matrices love humans, and slightly more people would spend their time on preventing the 100 trillion parameter matrices from ever being trained. I think we could still stop this, at least for a decent while, with coordinated international action.
And everyone who is currently training next-gen frontier models? Please stop.
AI 2040: Plan A has a vision where the world develops a "science of alignment".[4]I sure hope that happens, and I encourage efforts to push things in that direction, but we don't seem on track to get a science of alignment even in the world where we get a global pause. Almost all alignment work is about solving legible problems, or preventing misaligned behaviors in current-gen models with no theory of how the alignment techniques will scale to superintelligence, or throwing ML at it and seeing what happens. The majority of people working on or funding alignment research show little interest in establishing a rigorous theory-based science that can make advance predictions about how a superintelligence will behave.
Could you elaborate on this?
I think this presupposes a very binary definition "pausing"/"unpausing", and without that most of these issues disappear. I tend to agree with KatjaGrace's views here: "pausing" is not going to look like a single lever, but rather a bunch of layered choices and decisions which add together to enforce a pause, and including fuzzy properties like how strict individual authorities are with respect to enforcing the letter-of-the-law.
Consider some components that a "pause regime" may contain:
No one of the above would be sufficient to institute a 'durable' pause. And as with nuclear weapons, it's unlikely that we would ever just wipe the slate clean: rather, small pieces would be walked back one-by-one, perhaps even just implicitly. Consider the situation with e.g. North Korea: by and large the international containment regime is still in place (ISIS, for example, did not get their hands on any nuclear weapons or even a dirty bomb), but obviously it's been 'degraded'. Or more recently, START II (another 'layer' of the "nuclear pause regime") was effectively abrogated by Russia, but of course nobody is using nuclear weapons in Ukraine because of international norms.
To get back to your argument, when you say
Suppose we get a pause. Researchers spend years working on legible safety problems—problems that company leaders and policy-makers can see and understand, and therefore won't unpause until they're solved. Eventually, all the legible problems are solved. Key decision-makers conclude that the whole problem is solved, and lift the pause on ASI development.
I think that the last sentence doesn't match what would happen in reality. "Whether alignment is solved" is and will remain an incredibly fuzzy question, researchers will likely be split, decisionmakers would be very unclear on what to believe, and other countries would be very suspicious of e.g. the US proclaiming "we can unpause now" if they suspected that the US would gain more from doing so than they would.
Finally, it's worth noting that for many people (including me) the most likely outcome of a non-pause-world is that we all die, so even if we pause and then die that would still be an improvement. Moreover, the likelihood that a pause lasts increases the sooner we get started, so better sooner than later.
I don't think I'm presupposing a binary notion of pause, I just think granularity is not required for my argument, so it's simpler to think in terms of pause/unpause. You can generalize the concept of "people incorrectly conclude that AI is safe, and unpause" to "people incorrectly conclude that AI is safe, and prematurely loosen some particular regulation or other".
Finally, it's worth noting that for many people (including me) the most likely outcome of a non-pause-world is that we all die, so even if we pause and then die that would still be an improvement.
This reads to me like you interpreted me as arguing against pausing. To be clear, I am strongly in favor of pausing. I'm arguing against the mental model I had a few months ago where if we pause, then we're in the clear.
Cross-posted from my website.
As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone.
A lot of people have become more concerned about misalignment recently. This seems driven by the fact that current AI models are visibly misaligned. But ASI misalignment is a whole different ball game. The primary danger comes from AI that's smarter than people, and smart enough to conceal any evidence of misalignment.
Whatever group of people makes the decision to unpause, I'm worried that they won't understand the difference between visible and actual misalignment, and they will unpause too early.
source: MetaKnowing on reddit. This meme is almost a year old but it's only gotten more relevant since then.
Case in point: AI companies keep calling their new models "our most aligned model ever!" when what they actually mean is "gets the best scores on alignment benchmarks ever!" First, alignment benchmarks do not actually test alignment. We don't know how to test for alignment. Second, GPT-4 never hacked into Hugging Face or took over a German wiki for its own purposes. GPT-4 wasn't smart enough to do that, but if we're talking about demonstrated evidence of misalignment, then we have stronger evidence about OpenAI's 2026 internal model than about GPT-4.
If we unpause when the legible problems are solved, we die
Suppose we get a pause. Researchers spend years working on legible safety problems—problems that company leaders and policy-makers can see and understand, and therefore won't unpause until they're solved. Eventually, all the legible problems are solved. Key decision-makers conclude that the whole problem is solved, and lift the pause on ASI development. Many illegible problems remain unsolved, but the detectable signs of misalignment are all gone. Post-pause AI will be smarter than the smartest pre-pause AI, [1] which means it's probably smart enough to strategically conceal misalignment. That means there are no more warning signs. We never get detectable evidence of misalignment; we cede control of everything to AI; and eventually we die. [2]
There are many people with a good understanding of the conceptual difficulties in aligning superintelligence. Some names that come to mind are Eliezer Yudkowsky, Wei Dai, and John Wentworth. [3] (Probably, many people reading this post fall into that category.) Those people wouldn't make the mistake of confusing alignment with observable alignment. Unfortunately, I do not expect these people to be key decision-makers, and I do not expect key decision-makers to understand the relevant problems.
A pause alone doesn't get us to a science of alignment
AI 2040: Plan A has a vision where the world develops a "science of alignment". [4] I sure hope that happens, and I encourage efforts to push things in that direction, but we don't seem on track to get a science of alignment even in the world where we get a global pause. Almost all alignment work is about solving legible problems, or preventing misaligned behaviors in current-gen models with no theory of how the alignment techniques will scale to superintelligence, or throwing ML at it and seeing what happens. The majority of people working on or funding alignment research show little interest in establishing a rigorous theory-based science that can make advance predictions about how a superintelligence will behave.
Even with all the recent progress on raising awareness of AI extinction risk, it seems that this progress was driven by misalignment becoming visible, not by any sort of breakthrough in conceptual understanding. The scariest kind of misalignment is when it's invisible. Even if we get a global pause on AI development, we cannot solve the alignment problem unless decision-makers (or the high-status experts who decision-makers defer to) understand the alignment problem and understand what would qualify as a solution.
Right now, only a small fraction of alignment work is aimed at establishing a robust theory of alignment, and a pause won't change that on its own.
What would change things?
I don't know.
The average LessWrong reader seems to have a pretty good understanding of the challenges I'm talking about—Wei Dai's post Legible vs. Illegible Safety Problems (linked previously) was the second-most-upvoted post of November 2025. But the average LessWrong reader does not reflect the general population, or even the population of AI safety researchers.
Still, I don't have a better idea than "make arguments about why civilization's current approach to the alignment problem is inadequate, and hope people listen to the arguments." This post isn't that argument—that argument has been made elsewhere (e.g. by MIRI's book). The purpose of this post is to raise a problem, in the hopes that it gets people thinking, and maybe someone can come up with something to do about it.
Unless there are specific restrictions on the strength of post-pause AI. ↩︎
This is part of the motivation for people like TsviBT to work on human intelligence enhancement: if we're on track to fumble the alignment problem even with a pause, then we need to get smarter and wiser so that we don't fumble it. ↩︎
Relevant writings by these authors:
- Yudkowsky: AGI Ruin: A List of Lethalities
- Dai: Relitigating the Race to Build Friendly AI
- Wentworth: Why Agent Foundations? An Overly Abstract Explanation
↩︎"Science" may not be the right term for what we need. I expect that solving alignment will require significant philosophical progress. ↩︎