As a fellow Greenblattologist, thank you! I think Ryan is one of our top few alignment thinkers right now. Helping the rest of us keep up on his thinking is a valuable service. I like your compression here.
I'd like to condense the point on conceptual capabilities: perhaps we should improve conceptual capabilities faster, while we're still in the domain of models that don't have the capability to succeed at scheming because they still have CoT and inadequate capacity to do adequate thinking and steganography to evade detection.
Like you, I find this logic compelling but with so many caveats/dependencies that I'm not at all sure about it.
One positive side effect you haven't mentioned is making AI better for epistemics. That's a large upside. If anyone who asked an AI got the answer "for god's sake slow down and coordinate", that alone would help a lot, along with much better advice about how to do this. (This is included as "exogenous risks" in your summary, but it has broader benefits if the conceptual improvements are widespread across AIs; I don't remember if Ryan pulled that point out, probably he did somewhere.)
But it would also accelerate capabilities toward being takeover-capable, and increase the odds of misalignment by making LLM agents more capable of and prone to reason about its goals and discover misalignments by default.
So I'm torn. But we're in a bit of a pickle here, and just not doing dangerous things doesn't seem like the actual ethical option (while of course remaining aware of the unilaterists curse and the strength of Motivated reasoning on these topics).
I had read How do we (more) safely defer to AIs? but not AI 2040: Plan A, Alignment Roadmap; I hadn't caught its link in Plan A. I'm reading it now. I think it's extremely valuable in making explicit the default plan at this point.
It does seem pretty bad to accept trying to solve alignment alignment during an all-out race against others developing takeover-capable AI. That's the world we live in now, but we should be doing our utmost to change it as much as possible. I think everyone developing takeover-capable AGI will be quite concerned about alignment as they get closer, but they may very well not be concerned enough, or under such political pressure (or outright government control) that they can't really avoid racing.

These are not the remaining possibilities. There are two more possibilities. Goodhartburg, where the AI is misaligned in a way that only comes apart from human goodness under enough optimization pressure, and Incoheristead, where the AI is ultimately diametrically misaligned relative to humans, but doesn't realize it, because it is not that coherent, reflective or far-sighted.
In these worlds scheming propensities are pretty much directly proportional to conceptual reasoning, and what is to have as low conceptual reasoning levels as possible at any given level of capabilities.
Now, we might worry that the conceptual boost will partially transfer to scheming capabilities and propensities, because the AIs are better able to reflect on their values, strategise about how to undermine our evaluations, etc. Even if so, for a fixed conceptual capability, the AIs would have lower scheming capabilities/propensities after the boost.
Why is there only partial transfer? Well, scheming capabilities will be bottlenecked on opaque reasoning, steganography, obfuscating internal activations, etc.
This is true for any given thing? I don't see why it matters unless you're trying to just get the AIs to basically solve the alignment problem in one go.
Having a more jagged AI capability profile means it will be less scary at any level of capability of the spikes.
This article is about:
If you've read them, I'm impressed, they're both very long. If you haven't read them, you might be confused about:
For the sake of my fellow Greenblattologists, I have tried to boil down the arguments to a simple diagram.
Motivating scenario.
Responsible Leader. Let's assume that we're advising a reasonable AI company, with a 1-12 month lead over its competitors. The company will have poor incentives, it's managed by humans with typical flaws. However, the company has broadly good intentions, and isn't wildly mistaken about the strategic situation.
Conceptual workload. The reasonable AI company faces exogenous risks, e.g. a reckless competitor, or a rogue misaligned AI about to hit a software-only singularity. Managing these exogenous risks would require a sizable load of conceptual work, which is fuzzy, philosophically-loaded, and hard-to-verify. This includes:
There aren't enough humans around to do this in time.
The AIs seem capable and aligned. The company has AIs which look capable, aligned, and well-elicited on easy-to-verify tasks like programming, proving theorems, ML research with short feedback loops, etc. Behavioural testing hasn't found egregious misalignment, so the remaining possibilities are:
In Greenblatt's taxonomy, the live possibilities are Slopolis, Lurkville, or Easyland — not Hackistan or Schemeria.[1]
Our optimisation problem.
We're advising the company on how to minimise total risk. Simplifying somewhat, this is the sum of two things:
The company's strategy consists of three phases:
Overall, company's strategy is characterised by three levers:
Let's walk through how these levers determine the overall risk.
Overall capabilities. The longer the company spends scaling to larger models, the more capable the models will be. This includes:
Note that scaling to larger models also reduces the remaining lead time that the company can spend on evaluation and deployment. At the limit, the company could spend their entire lead on building the most capable model they can. This would have the maximal conceptual abilities, but the AIs would have little time to actually deal with the exogenous risks.
Scheming propensities. The company wants to ensure that the AIs aren't covertly egregiously misaligned. If the AIs are more generally capable, then it's more likely that our finetuning and evaluations have been undermined (see above). Therefore, not only are bigger AIs more capable of scheming, but they are also more likely to actually be scheming when we deploy them.
The company can burn more of its lead time on evaluations and finetunes — this reduces the scheming propensities, but not the scheming capabilities.
Loss-of-control. When the company deploys the AIs, it can choose how much humans (and weak trusted AIs) should check the AIs thought processes, outputs, and general activity. If AIs have scheming propensities, and their scheming capabilities are sufficient for subverting this oversight, then this leads to a loss-of-control catastrophe. However, if the safety of the oversight exceeds the scheming capabilities, then the AIs cannot take over even if they are scheming.
Exogenous risk. Unfortunately, oversight reduces how much work achieved the AIs can achieve to manage the exogenous risks — the AIs must constantly wait for humans to get "clued up" on what's happening. Overall, the total work achieved by the AIs is determined by:
The case for early handoff
You can see, by meditating on the first diagram, that for some choices of parameters, an early handoff reduces overall risk. I'll describe the modal region of parameter space.
Workload exceeds lead time by 100-500x. Maybe the rival is 3 months from a software-only singularity. To deal with this, the reasonable company must build extremely powerful AIs. But doing this safely would require breakthroughs in machine learning, agency, philosophy, etc — this might be decades of intellectual progress. If we say 25-125 years, then that's a ratio of 100-500x.
Human-level conceptual reasoning suffices for workload. Greenblatt thinks "epistemics comparable to Rohin Shah, Jared Kaplan, Daniel Kokotajlo, and Nate Soares would probably all be OK despite the fact that these people disagree a huge amount."
Why is this enough for a 500x speedup? Well, AIs will think 50x faster than humans. Expanding from 1000 humans to 100,000 should add another 10-20x speedup. There's also a 3x from working 24/7, and other multipliers from superhuman coordination (e.g. being able to be instantly redirected towards the most promising tasks) and better epistemic tracking (e.g. AIs can track which heuristics and epistemic processes actually lead to good outcomes or accurate predictions).
Below human-level conceptual reasoning won't suffice. If the AIs are noticeably worse than humans at conceptual reasoning, then they lack the core skills required to "stay on track" even with a long deployment and low oversight. This is because conceptual tasks can't be easily decomposed into subtasks that sub-par AIs can complete.
Pessimism about (scalable) oversight. The AIs must deliver a 100-500x speedup on conceptual work — is this feasible with an oversight protocol with decent safety? Probably not. Suppose we wanted to only 5% probability of loss-of-control, conditional on scheming, to 5%. This would require humans to check the work, monitor their reasoning, read the AIs research, learn about the new techniques, build up new intuitions, etc. But this is roughly as hard as doing the work itself. Maybe with scalable oversight like debate we can achieve a 10x uplift, but not a 100x uplift.
Possibly, if the AIs were superhuman at conceptual reasoning, then they could perform useful work despite our safety-bounded oversight protocols. However, this would require training bigger models, who have more chance of scheming, and this might increase the loss-of-control risk, compared with deploying weaker models with less oversight.
So we should "hand off" to the AIs, deferring to them about most strategic and research questions, without making them wait until we can check their work. This buys maximum usefulness, but has no robustness to scheming.
Handoff early in capability space. Because oversight won't be robust to scheming, it's important we don't deploy schemers. If we grant schemers control of the leading datacenter, and aren't checking how they train successors, they could straightforwardly disempower humanity. Hence, we should train weaker models. This limits their scheming capabilities, so we can trust our evaluations more. It also directly reduces their scheming propensities, because these will probably grow with scale.[3]
Unfortunately, the AI company might have a bias against this, because deploying stronger models with greater oversight might seem less scary than deploying weaker models with less oversight.
Handoff early in calendar time. A bonus of handing off to weaker AIs means that we have more lead time remaining for evaluation (which again reduces scheming propensities) and deployment (so the AIs can manage the exogenous risks).
Once we've trained the weaker models, we shouldn't burn too much of our remaining lead time evaluating them. Although this evaluation reduces the chance of handing off to schemers, it means that the AIs will have less time to manage the exogenous risks.
There's a runaway regime we must avoid: You delayed deployment for too long, so your AIs can't meet the deadline at their current capability, so you're forced to scale to larger models, reducing their lead time even further, so you're forced to scale even further, etc.
The case for improving conceptual reasoning
Suppose the safety community delivers a one-time boost to conceptual capabilities — benchmarks, datasets, training environments. How does this change overall risk?
Conceptual uplift reduces exogenous risk. If the AI company keeps the other levers fixed (train-lead ratio, eval-deploy split, oversight) then better conceptual capabilities would straightforwardly reduce exogenous risks, because the AIs can achieve more work during deployment. They can move faster, think more creatively, and make better decisions.
Conceptual uplift reduces loss-of-control risk if the company adapts its levers in response to the boost. This is because the company can train smaller models to achieve the same conceptual capabilities — which means less scheming capabilities and propensities, and more time for evaluation. The reasonable company isn't forced to scale to larger models.
Now, we might worry that the conceptual boost will partially transfer to scheming capabilities and propensities, because the AIs are better able to reflect on their values, strategise about how to undermine our evaluations, etc. Even if so, for a fixed conceptual capability, the AIs would have lower scheming capabilities/propensities after the boost.
Why is there only partial transfer? Well, scheming capabilities will be bottlenecked on opaque reasoning, steganography, obfuscating internal activations, etc.
Companies will neglect conceptual reasoning. The default trajectory through the capability profile is determined by:
For example, AI R&D capabilities will go fast, because there are commercial incentives and it is relatively easy to build environments to train for this.
So companies will underinvest in two things:
Overall, this suggests that the reasonable company will neglect conceptual reasoning, compared with the other requirements for handoff (e.g. AI R&D). Hence, safety-minded outsiders can expect some returns for investing in improving conceptual capabilities.
Specific flaws/cruxes/limitations
Here are some obvious ways the argument could fail:
Reasonable company. We assume that the leading company is reasonable, and willing to slow down in order to reduce loss-of-control. However, if the level of political will at the leading company is substantially lower, then this doesn't make sense. The cynical prediction is that companies will race as fast as possible.
Miracles. This model assumes that handing off to schemers after one month is no worse than handing off after six. But it isn't: There are other routes for mitigating exogenous risk which don't route through the reasonable company racing ahead, such as ordinary coordination, drastic government action, the competitor simply stumbling, or random miracles. Therefore, an early catastrophe is worse, because it forecloses these other hopes.
Coup risk. If you're worried about an illegitimate powergrab from the company CEOs or the domestic government, then that complicates the model. In particular, boosting conceptual capabilities will have significant transfer to coup-enabling capabilities — both involve strategic reasoning, forecasting, threat modelling, etc. Unlike loss-of-control risks, these can't be mitigated with strong oversight, because the threat actor presumably has the corporate or political power to change the safeguards.
Exfiltration. If you're worried about an adversary (e.g. the rival AI company, or country) stealing the weights, then this is an additional case for early handoff. If the adversary stole the weights of an AI with greater AI R&D capabilities and cyberoffensive capabilities, then that severely cuts the lead time.
Scale might reduce scheming. This model assumes that larger models are more likely to have scheming propensities. This might be false. In particular, many scheming propensities seem to result from imperfect labelling or poorly designed RL environments. But if the larger models are far more sample-efficient, then maybe we can train them on fewer environments of higher quality, and we can put more human effort into grading their actions.
Handoff determined by date, not conceptual capabilities. The AI company might continue training until some fixed date, rather than adapting the train-lead ratio to some level of conceptual capabilities. If so, then the conceptual boost looks worse — the boost has some transfer on scheming capabilities and propensities, and this increases loss-of-control.
No handoff. If companies never actually hand off to AIs, then boosting conceptual capabilities looks worse. This is because the bottleneck on the work achieved will be the humans checking the AIs work, not the conceptual capabilities of the AIs. In this case, the main effect of the conceptual boost is increasing loss-of-control risk, via transfer on scheming capabilities/propensities.
Deeper worries
Beyond the specific limitations, you should have a healthy suspicion of my model and any arguments based on it.
I'm also worried that the whole approach is flawed in a deeper sense. At best, the conclusions will be conjunctive, sensitive to unknown parameters, contingent on the changing strategic landscape, and miss crucial dynamics. Instead, maybe we should act on robust heuristics, like "transparency is good" or "refrain from actions you think will directly cause catastrophic harm". I don’t want to tacitly endorse backchaining through a sprawling network of boxes.
My best guess is that the arguments in How do we (more) safely defer to AIs? and AI 2040: Plan A, Alignment Roadmap are broadly correct. But nothing in this article should be interpreted as adding any rigour to what was already there.
This is not a wise approach to macrostrategy.
From AI 2040: Plan A, Alignment Roadmap:
I call this "loss-of-control" instead of endogenous risk, because it excludes other endogenous risks like coups.
It's an open question how scheming propensities (not scheming capabilities) change with scale, but it probably increases.