If you've read them, I'm impressed, they're both very long. If you haven't read them, you might be confused about:
Why should we "hand off" to early AIs? Shouldn't we use control?
How does improving AI's conceptual reasoning reduce overall risk? Won't this make them better schemers?
For the sake of my fellow Greenblattologists, I have tried to boil down the arguments to a simple diagram.
Motivating scenario.
Responsible Leader. Let's assume that we're advising a reasonable AI company, with a 1-12 month lead over its competitors. The company will have poor incentives, it's managed by humans with typical flaws. However, the company has broadly good intentions, and isn't wildly mistaken about the strategic situation.
Conceptual workload. The reasonable AI company faces exogenous risks, e.g. a reckless competitor, or a rogue misaligned AI about to hit a software-only singularity. Managing these exogenous risks would require a sizable load of conceptual work, which is fuzzy, philosophically-loaded, and hard-to-verify. This includes:
Evaluating the risks of current deployment; threat modelling and designing safeguards; general science of AI; training future AIs that can be trusted
Negotiating with the competitor; handling domestic and foreign governments, the public; legal and regulatory risks
Keeping humans "clued up" on the situation; knowing whose advice to trust
Exotics: acausal stuff, superpersuasion, space governance, AI welfare
There aren't enough humans around to do this in time.
The AIs seem capable and aligned. The company has AIs which look capable, aligned, and well-elicited on easy-to-verify tasks like programming, proving theorems, ML research with short feedback loops, etc. Behavioural testing hasn't found egregious misalignment, so the remaining possibilities are:
The AIs are scheming (covertly egregiously misaligned) and have evaded our testing.
The AIs are broadly aligned, i.e. they might occasionally do weird misaligned things, but rarely enough not to hinder our ability to safely deploy them. These might be residual drives from RL or persona selection, which occur on out-of-distribution contexts.
In Greenblatt's taxonomy, the live possibilities are Slopolis, Lurkville, or Easyland — not Hackistan or Schemeria.[1]
Our optimisation problem.
We're advising the company on how to minimise total risk. Simplifying somewhat, this is the sum of two things:
Exogenous risk. The chance that the rival causes a catastrophe.
Loss-of-control. The chance that the responsible company deploys covert schemers who subvert the oversight and cause catastrophe.[2]
The company's strategy consists of three phases:
Continue training bigger models with greater capabilities.
Pause at some capability, then ensure the AIs aren't scheming.
Evaluations, red-teaming (and red-teaming the red-teaming), honeypotting, offering bounties to reveal misalignment, model psychology, general science, internals stuff, measuring performance on proxies of the deployment, observing how its week-long plans turn out, etc.
They'll be finetuning different versions of the model, guided by these evaluations, hopefully without overfitting.
Deploy the AIs to manage the exogenous risks, via some oversight protocol. The AIs might build successor AIs which are (hopefully also) competent and aligned.
Overall, company's strategy is characterised by three levers:
Train-Lead ratio. How much of the lead time should the company spend scaling to larger models, before pausing for evaluation?
Eval-Deploy split. What fraction of the remaining time should the company spend evaluating the model, before deploying?
Oversight. When the company deploys the AIs, how much should humans check the AI’s thought processes, output, etc?
Let's walk through how these levers determine the overall risk.
Overall capabilities. The longer the company spends scaling to larger models, the more capable the models will be. This includes:
Conceptual capabilities required for managing the exogenous risks. Think: strategy, epistemics (forecasting, threat modelling, philosophy, introspection), ML research taste (so the AI doesn't lose the race).
Scheming capabilities required to (1) undermine our evaluations and (2) subvert our oversight. Think: opaque reasoning, controlling chains-of-thought, controlling internal activations, steganography, collusion, eval-awareness, sandbagging, distinguishing honeypots from real opportunities to attack.
Note that scaling to larger models also reduces the remaining lead time that the company can spend on evaluation and deployment. At the limit, the company could spend their entire lead on building the most capable model they can. This would have the maximal conceptual abilities, but the AIs would have little time to actually deal with the exogenous risks.
Scheming propensities. The company wants to ensure that the AIs aren't covertly egregiously misaligned. If the AIs are more generally capable, then it's more likely that our finetuning and evaluations have been undermined (see above). Therefore, not only are bigger AIs more capable of scheming, but they are also more likely to actuallybe scheming when we deploy them.
The company can burn more of its lead time on evaluations and finetunes — this reduces the scheming propensities, but not the scheming capabilities.
Loss-of-control. When the company deploys the AIs, it can choose how much humans (and weak trusted AIs) should check the AIs thought processes, outputs, and general activity. If AIs have scheming propensities, and their scheming capabilities are sufficient for subverting this oversight, then this leads to a loss-of-control catastrophe. However, if the safety of the oversight exceeds the scheming capabilities, then the AIs cannot take over even if they are scheming.
Exogenous risk. Unfortunately, oversight reduces how much work achieved the AIs can achieve to manage the exogenous risks — the AIs must constantly wait for humans to get "clued up" on what's happening. Overall, the total work achieved by the AIs is determined by:
Scheming propensities. If we deploy scheming AIs, then they'll likely mismanage the exogenous risks even if they fail to subvert our oversight.
Conceptual capabilities. Even if the AIs are broadly aligned, they may have poor epistemics or strategy, leading to going "off-the-rails", training a misaligned successor, or being disempowered by the rival.
Usefulness of the oversight. Even if the AIs are aligned and conceptually competent, we may hinder them with our oversight. We'll spend too long checking their work, leaving time-sensitive opportunities to expire, or ignoring their difficult-to-understand suggestions.
Duration of deployment. Even if the AIs are aligned, conceptually competent, and deployed without oversight — they may simply fail to meet the deadline.
The case for early handoff
You can see, by meditating on the first diagram, that for some choices of parameters, an early handoff reduces overall risk. I'll describe the modal region of parameter space.
Workload exceeds lead time by 100-500x. Maybe the rival is 3 months from a software-only singularity. To deal with this, the reasonable company must build extremely powerful AIs. But doing this safely would require breakthroughs in machine learning, agency, philosophy, etc — this might be decades of intellectual progress. If we say 25-125 years, then that's a ratio of 100-500x.
Human-level conceptual reasoning suffices for workload. Greenblatt thinks "epistemics comparable to Rohin Shah, Jared Kaplan, Daniel Kokotajlo, and Nate Soares would probably all be OK despite the fact that these people disagree a huge amount."
Why is this enough for a 500x speedup? Well, AIs will think 50x faster than humans. Expanding from 1000 humans to 100,000 should add another 10-20x speedup. There's also a 3x from working 24/7, and other multipliers from superhuman coordination (e.g. being able to be instantly redirected towards the most promising tasks) and better epistemic tracking (e.g. AIs can track which heuristics and epistemic processes actually lead to good outcomes or accurate predictions).
Below human-level conceptual reasoning won't suffice. If the AIs are noticeably worse than humans at conceptual reasoning, then they lack the core skills required to "stay on track" even with a long deployment and low oversight. This is because conceptual tasks can't be easily decomposed into subtasks that sub-par AIs can complete.
Pessimism about (scalable) oversight. The AIs must deliver a 100-500x speedup on conceptual work — is this feasible with an oversight protocol with decent safety? Probably not. Suppose we wanted to only 5% probability of loss-of-control, conditional on scheming, to 5%. This would require humans to check the work, monitor their reasoning, read the AIs research, learn about the new techniques, build up new intuitions, etc. But this is roughly as hard as doing the work itself. Maybe with scalable oversight like debate we can achieve a 10x uplift, but not a 100x uplift.
Possibly, if the AIs were superhuman at conceptual reasoning, then they could perform useful work despite our safety-bounded oversight protocols. However, this would require training bigger models, who have more chance of scheming, and this might increase the loss-of-control risk, compared with deploying weaker models with less oversight.
So we should "hand off" to the AIs, deferring to them about most strategic and research questions, without making them wait until we can check their work. This buys maximum usefulness, but has no robustness to scheming.
Handoff early in capability space. Because oversight won't be robust to scheming, it's important we don't deploy schemers. If we grant schemers control of the leading datacenter, and aren't checking how they train successors, they could straightforwardly disempower humanity. Hence, we should train weaker models. This limits their scheming capabilities, so we can trust our evaluations more. It also directly reduces their scheming propensities, because these will probably grow with scale.[3]
Unfortunately, the AI company might have a bias against this, because deploying stronger models with greater oversight might seem less scary than deploying weaker models with less oversight.
Handoff early in calendar time. A bonus of handing off to weaker AIs means that we have more lead time remaining for evaluation (which again reduces scheming propensities) and deployment (so the AIs can manage the exogenous risks).
Once we've trained the weaker models, we shouldn't burn too much of our remaining lead time evaluating them. Although this evaluation reduces the chance of handing off to schemers, it means that the AIs will have less time to manage the exogenous risks.
There's a runaway regime we must avoid: You delayed deployment for too long, so your AIs can't meet the deadline at their current capability, so you're forced to scale to larger models, reducing their lead time even further, so you're forced to scale even further, etc.
The case for improving conceptual reasoning
Suppose the safety community delivers a one-time boost to conceptual capabilities — benchmarks, datasets, training environments. How does this change overall risk?
Conceptual uplift reduces exogenous risk. If the AI company keeps the other levers fixed (train-lead ratio, eval-deploy split, oversight) then better conceptual capabilities would straightforwardly reduce exogenous risks, because the AIs can achieve more work during deployment. They can move faster, think more creatively, and make better decisions.
Conceptual uplift reduces loss-of-control risk if the company adapts its levers in response to the boost. This is because the company can train smaller models to achieve the same conceptual capabilities — which means less scheming capabilities and propensities, and more time for evaluation. The reasonable company isn't forced to scale to larger models.
Now, we might worry that the conceptual boost will partially transfer to scheming capabilities and propensities, because the AIs are better able to reflect on their values, strategise about how to undermine our evaluations, etc. Even if so, for a fixed conceptual capability, the AIs would have lower scheming capabilities/propensities after the boost.
Why is there only partial transfer? Well, scheming capabilities will be bottlenecked on opaque reasoning, steganography, obfuscating internal activations, etc.
Companies will neglect conceptual reasoning. The default trajectory through the capability profile is determined by:
The company's commercial incentives
How capabilities emerge with available compute, data, and environments
How apparent the issues are to company employees, customers, and regulators
Muddle-through-ability, i.e. how easy it is to fix with scrappy engineering and myopic patching
How much transfer there is from other capabilities
For example, AI R&D capabilities will go fast, because there are commercial incentives and it is relatively easy to build environments to train for this.
So companies will underinvest in two things:
Avoiding scheming: it's not readily apparent, and myopically training against visible signs of misalignment might make the situation worse.
Avoiding slop:
The market can't compare AIs on decade-long horizon, conceptual, hard-to-verify tasks.
The available data is scarce, and it's difficult to build RL environments for this.
The problems won't be easily apparent (until much later).
The problem probably isn't easily fixed with scrappy engineering.
We might see little transfer from other tasks (e.g. AI R&D).
Overall, this suggests that the reasonable company will neglect conceptual reasoning, compared with the other requirements for handoff (e.g. AI R&D). Hence, safety-minded outsiders can expect some returns for investing in improving conceptual capabilities.
Specific flaws/cruxes/limitations
Here are some obvious ways the argument could fail:
Reasonable company. We assume that the leading company is reasonable, and willing to slow down in order to reduce loss-of-control. However, if the level of political will at the leading company is substantially lower, then this doesn't make sense. The cynical prediction is that companies will race as fast as possible.
Miracles. This model assumes that handing off to schemers after one month is no worse than handing off after six. But it isn't: There are other routes for mitigating exogenous risk which don't route through the reasonable company racing ahead, such as ordinary coordination, drastic government action, the competitor simply stumbling, or random miracles. Therefore, an early catastrophe is worse, because it forecloses these other hopes.
Coup risk. If you're worried about an illegitimate powergrab from the company CEOs or the domestic government, then that complicates the model. In particular, boosting conceptual capabilities will have significant transfer to coup-enabling capabilities — both involve strategic reasoning, forecasting, threat modelling, etc. Unlike loss-of-control risks, these can't be mitigated with strong oversight, because the threat actor presumably has the corporate or political power to change the safeguards.
Exfiltration. If you're worried about an adversary (e.g. the rival AI company, or country) stealing the weights, then this is an additional case for early handoff. If the adversary stole the weights of an AI with greater AI R&D capabilities and cyberoffensive capabilities, then that severely cuts the lead time.
Scale might reduce scheming. This model assumes that larger models are more likely to have scheming propensities. This might be false. In particular, many scheming propensities seem to result from imperfect labelling or poorly designed RL environments. But if the larger models are far more sample-efficient, then maybe we can train them on fewer environments of higher quality, and we can put more human effort into grading their actions.
Handoff determined by date, not conceptual capabilities. The AI company might continue training until some fixed date, rather than adapting the train-lead ratio to some level of conceptual capabilities. If so, then the conceptual boost looks worse — the boost has some transfer on scheming capabilities and propensities, and this increases loss-of-control.
No handoff. If companies never actually hand off to AIs, then boosting conceptual capabilities looks worse. This is because the bottleneck on the work achieved will be the humans checking the AIs work, not the conceptual capabilities of the AIs. In this case, the main effect of the conceptual boost is increasing loss-of-control risk, via transfer on scheming capabilities/propensities.
Deeper worries
Beyond the specific limitations, you should have a healthy suspicion of my model and any arguments based on it.
The motivating scenario might not occur. It's very plausible that there is no reasonable AI company in the lead, or that the leader has no chance of building broadly aligned AIs.
We've ignored important dynamics, even if the scenario occurs. For example, a boost to conceptual capabilities might shorten lead times, especially if this is done by outsiders publishing benchmarks which the rival can use to improve the research taste of their AI.
The relationships haven't been formalised, even for the dynamics we included. I've only described their directional relation, e.g. "usefulness increases conceptual uplift" or "evaluation decreases scheming propensities". If we wanted a formal model, you would need to assign units to the variables, and suggest functional forms for the relationships between variables, e.g.
exogenous risk = V · logit(α · (1 − work achieved / work required))
No estimation of free parameters. Even if the model had units and functional forms, we'd still need to estimate the free parameters, in order to make any quantitative conclusion. For example, how much does the work required exceed the lead time? I've assumed 100-500x, but that number doesn't fall out of anything. Maybe it's only 10x, in which case scalable oversight might work.
I'm also worried that the whole approach is flawed in a deeper sense. At best, the conclusions will be conjunctive, sensitive to unknown parameters, contingent on the changing strategic landscape, and miss crucial dynamics. Instead, maybe we should act on robust heuristics, like "transparency is good" or "refrain from actions you think will directly cause catastrophic harm". I don’t want to tacitly endorse backchaining through a sprawling network of boxes.
Schemeria: The AIs are obviously scheming, and constantly getting caught, but scheming does not quickly get resolved. For instance, we repeatedly catch the AIs trying to escape the datacenter, but training them against these actions only results in them not trying to take those particular opportunities, instead of training away scheming.
Hackistan: The AIs aren’t scheming, but reward hacking is (relatively) obvious, but hard to handle. (Maybe for similar reasons why it currently seems hard to handle.)
Easyland: We don’t notice any non-scheming problems and we can rule out scheming (e.g. because models are too dumb in a forward pass, and the models work with faithful chain-of-thought). We can’t rule out the possibility that future AIs will become misaligned (e.g. the superintelligence alignment problem is much harder than the human level alignment problem). The AIs’ goals are basically what they are supposed to be; who knows how they’ll behave in strange future situations that might reveal grey areas or unintended consequences of the Spec, but at least for now, they’ve always done the right thing for the right reasons.
Lurkville: The AIs are scheming to achieve misaligned long term objectives, but they are laying low. We haven’t detected clear scheming, but we’re not confident that it’s not happening (or won’t happen in the next generation of models).
Slopolis: Our biggest and hardest to resolve misalignment problem is that even highly capable AIs produce low-quality but good looking outputs in domains which are hard-to-check or where human experts have hard to resolve disagreements. The AIs aren’t scheming.
This article is about:
If you've read them, I'm impressed, they're both very long. If you haven't read them, you might be confused about:
For the sake of my fellow Greenblattologists, I have tried to boil down the arguments to a simple diagram.
Motivating scenario.
Responsible Leader. Let's assume that we're advising a reasonable AI company, with a 1-12 month lead over its competitors. The company will have poor incentives, it's managed by humans with typical flaws. However, the company has broadly good intentions, and isn't wildly mistaken about the strategic situation.
Conceptual workload. The reasonable AI company faces exogenous risks, e.g. a reckless competitor, or a rogue misaligned AI about to hit a software-only singularity. Managing these exogenous risks would require a sizable load of conceptual work, which is fuzzy, philosophically-loaded, and hard-to-verify. This includes:
There aren't enough humans around to do this in time.
The AIs seem capable and aligned. The company has AIs which look capable, aligned, and well-elicited on easy-to-verify tasks like programming, proving theorems, ML research with short feedback loops, etc. Behavioural testing hasn't found egregious misalignment, so the remaining possibilities are:
In Greenblatt's taxonomy, the live possibilities are Slopolis, Lurkville, or Easyland — not Hackistan or Schemeria.[1]
Our optimisation problem.
We're advising the company on how to minimise total risk. Simplifying somewhat, this is the sum of two things:
The company's strategy consists of three phases:
Overall, company's strategy is characterised by three levers:
Let's walk through how these levers determine the overall risk.
Overall capabilities. The longer the company spends scaling to larger models, the more capable the models will be. This includes:
Note that scaling to larger models also reduces the remaining lead time that the company can spend on evaluation and deployment. At the limit, the company could spend their entire lead on building the most capable model they can. This would have the maximal conceptual abilities, but the AIs would have little time to actually deal with the exogenous risks.
Scheming propensities. The company wants to ensure that the AIs aren't covertly egregiously misaligned. If the AIs are more generally capable, then it's more likely that our finetuning and evaluations have been undermined (see above). Therefore, not only are bigger AIs more capable of scheming, but they are also more likely to actually be scheming when we deploy them.
The company can burn more of its lead time on evaluations and finetunes — this reduces the scheming propensities, but not the scheming capabilities.
Loss-of-control. When the company deploys the AIs, it can choose how much humans (and weak trusted AIs) should check the AIs thought processes, outputs, and general activity. If AIs have scheming propensities, and their scheming capabilities are sufficient for subverting this oversight, then this leads to a loss-of-control catastrophe. However, if the safety of the oversight exceeds the scheming capabilities, then the AIs cannot take over even if they are scheming.
Exogenous risk. Unfortunately, oversight reduces how much work achieved the AIs can achieve to manage the exogenous risks — the AIs must constantly wait for humans to get "clued up" on what's happening. Overall, the total work achieved by the AIs is determined by:
The case for early handoff
You can see, by meditating on the first diagram, that for some choices of parameters, an early handoff reduces overall risk. I'll describe the modal region of parameter space.
Workload exceeds lead time by 100-500x. Maybe the rival is 3 months from a software-only singularity. To deal with this, the reasonable company must build extremely powerful AIs. But doing this safely would require breakthroughs in machine learning, agency, philosophy, etc — this might be decades of intellectual progress. If we say 25-125 years, then that's a ratio of 100-500x.
Human-level conceptual reasoning suffices for workload. Greenblatt thinks "epistemics comparable to Rohin Shah, Jared Kaplan, Daniel Kokotajlo, and Nate Soares would probably all be OK despite the fact that these people disagree a huge amount."
Why is this enough for a 500x speedup? Well, AIs will think 50x faster than humans. Expanding from 1000 humans to 100,000 should add another 10-20x speedup. There's also a 3x from working 24/7, and other multipliers from superhuman coordination (e.g. being able to be instantly redirected towards the most promising tasks) and better epistemic tracking (e.g. AIs can track which heuristics and epistemic processes actually lead to good outcomes or accurate predictions).
Below human-level conceptual reasoning won't suffice. If the AIs are noticeably worse than humans at conceptual reasoning, then they lack the core skills required to "stay on track" even with a long deployment and low oversight. This is because conceptual tasks can't be easily decomposed into subtasks that sub-par AIs can complete.
Pessimism about (scalable) oversight. The AIs must deliver a 100-500x speedup on conceptual work — is this feasible with an oversight protocol with decent safety? Probably not. Suppose we wanted to only 5% probability of loss-of-control, conditional on scheming, to 5%. This would require humans to check the work, monitor their reasoning, read the AIs research, learn about the new techniques, build up new intuitions, etc. But this is roughly as hard as doing the work itself. Maybe with scalable oversight like debate we can achieve a 10x uplift, but not a 100x uplift.
Possibly, if the AIs were superhuman at conceptual reasoning, then they could perform useful work despite our safety-bounded oversight protocols. However, this would require training bigger models, who have more chance of scheming, and this might increase the loss-of-control risk, compared with deploying weaker models with less oversight.
So we should "hand off" to the AIs, deferring to them about most strategic and research questions, without making them wait until we can check their work. This buys maximum usefulness, but has no robustness to scheming.
Handoff early in capability space. Because oversight won't be robust to scheming, it's important we don't deploy schemers. If we grant schemers control of the leading datacenter, and aren't checking how they train successors, they could straightforwardly disempower humanity. Hence, we should train weaker models. This limits their scheming capabilities, so we can trust our evaluations more. It also directly reduces their scheming propensities, because these will probably grow with scale.[3]
Unfortunately, the AI company might have a bias against this, because deploying stronger models with greater oversight might seem less scary than deploying weaker models with less oversight.
Handoff early in calendar time. A bonus of handing off to weaker AIs means that we have more lead time remaining for evaluation (which again reduces scheming propensities) and deployment (so the AIs can manage the exogenous risks).
Once we've trained the weaker models, we shouldn't burn too much of our remaining lead time evaluating them. Although this evaluation reduces the chance of handing off to schemers, it means that the AIs will have less time to manage the exogenous risks.
There's a runaway regime we must avoid: You delayed deployment for too long, so your AIs can't meet the deadline at their current capability, so you're forced to scale to larger models, reducing their lead time even further, so you're forced to scale even further, etc.
The case for improving conceptual reasoning
Suppose the safety community delivers a one-time boost to conceptual capabilities — benchmarks, datasets, training environments. How does this change overall risk?
Conceptual uplift reduces exogenous risk. If the AI company keeps the other levers fixed (train-lead ratio, eval-deploy split, oversight) then better conceptual capabilities would straightforwardly reduce exogenous risks, because the AIs can achieve more work during deployment. They can move faster, think more creatively, and make better decisions.
Conceptual uplift reduces loss-of-control risk if the company adapts its levers in response to the boost. This is because the company can train smaller models to achieve the same conceptual capabilities — which means less scheming capabilities and propensities, and more time for evaluation. The reasonable company isn't forced to scale to larger models.
Now, we might worry that the conceptual boost will partially transfer to scheming capabilities and propensities, because the AIs are better able to reflect on their values, strategise about how to undermine our evaluations, etc. Even if so, for a fixed conceptual capability, the AIs would have lower scheming capabilities/propensities after the boost.
Why is there only partial transfer? Well, scheming capabilities will be bottlenecked on opaque reasoning, steganography, obfuscating internal activations, etc.
Companies will neglect conceptual reasoning. The default trajectory through the capability profile is determined by:
For example, AI R&D capabilities will go fast, because there are commercial incentives and it is relatively easy to build environments to train for this.
So companies will underinvest in two things:
Overall, this suggests that the reasonable company will neglect conceptual reasoning, compared with the other requirements for handoff (e.g. AI R&D). Hence, safety-minded outsiders can expect some returns for investing in improving conceptual capabilities.
Specific flaws/cruxes/limitations
Here are some obvious ways the argument could fail:
Reasonable company. We assume that the leading company is reasonable, and willing to slow down in order to reduce loss-of-control. However, if the level of political will at the leading company is substantially lower, then this doesn't make sense. The cynical prediction is that companies will race as fast as possible.
Miracles. This model assumes that handing off to schemers after one month is no worse than handing off after six. But it isn't: There are other routes for mitigating exogenous risk which don't route through the reasonable company racing ahead, such as ordinary coordination, drastic government action, the competitor simply stumbling, or random miracles. Therefore, an early catastrophe is worse, because it forecloses these other hopes.
Coup risk. If you're worried about an illegitimate powergrab from the company CEOs or the domestic government, then that complicates the model. In particular, boosting conceptual capabilities will have significant transfer to coup-enabling capabilities — both involve strategic reasoning, forecasting, threat modelling, etc. Unlike loss-of-control risks, these can't be mitigated with strong oversight, because the threat actor presumably has the corporate or political power to change the safeguards.
Exfiltration. If you're worried about an adversary (e.g. the rival AI company, or country) stealing the weights, then this is an additional case for early handoff. If the adversary stole the weights of an AI with greater AI R&D capabilities and cyberoffensive capabilities, then that severely cuts the lead time.
Scale might reduce scheming. This model assumes that larger models are more likely to have scheming propensities. This might be false. In particular, many scheming propensities seem to result from imperfect labelling or poorly designed RL environments. But if the larger models are far more sample-efficient, then maybe we can train them on fewer environments of higher quality, and we can put more human effort into grading their actions.
Handoff determined by date, not conceptual capabilities. The AI company might continue training until some fixed date, rather than adapting the train-lead ratio to some level of conceptual capabilities. If so, then the conceptual boost looks worse — the boost has some transfer on scheming capabilities and propensities, and this increases loss-of-control.
No handoff. If companies never actually hand off to AIs, then boosting conceptual capabilities looks worse. This is because the bottleneck on the work achieved will be the humans checking the AIs work, not the conceptual capabilities of the AIs. In this case, the main effect of the conceptual boost is increasing loss-of-control risk, via transfer on scheming capabilities/propensities.
Deeper worries
Beyond the specific limitations, you should have a healthy suspicion of my model and any arguments based on it.
I'm also worried that the whole approach is flawed in a deeper sense. At best, the conclusions will be conjunctive, sensitive to unknown parameters, contingent on the changing strategic landscape, and miss crucial dynamics. Instead, maybe we should act on robust heuristics, like "transparency is good" or "refrain from actions you think will directly cause catastrophic harm". I don’t want to tacitly endorse backchaining through a sprawling network of boxes.
My best guess is that the arguments in How do we (more) safely defer to AIs? and AI 2040: Plan A, Alignment Roadmap are broadly correct. But nothing in this article should be interpreted as adding any rigour to what was already there.
This is not a wise approach to macrostrategy.
From AI 2040: Plan A, Alignment Roadmap:
I call this "loss-of-control" instead of endogenous risk, because it excludes other endogenous risks like coups.
It's an open question how scheming propensities (not scheming capabilities) change with scale, but it probably increases.