With the recent highly publicised alignment failures from OpenAI and Anthropic there's been increasing calls for an AI pause, to allow alignment to catch up.
However a pause has various obvious issues:
It's possible that we can't make further alignment progress without more advanced AIs:
Trying to align AGIs now has been compared to trying to ensure jetliners are safe while we're still using turboprops.
More advanced AI may be necessary to help us crack alignment.
How will we ever know that it's safe to stop the pause?
What prevents the pause going on forever?
What prevents us eventually removing the pause, and then everybody speeds full sprint ahead the same as before.
It's difficult to time the pause right: too early and we won't be able to make sufficient alignment progress (see above), too late and we'll already be dead.
It seems much more sensible to slow down the rate of AI progress than to halt it entirely.
Here are some of the ways I expect this to be helpful compared to the current dynamics:
Removes race dynamics where companies feel forced to release frontier models despite knowing there are troubling issues with them[1].
More time to fully investigate issues that occur.
Much greater effort spent on progressing applied alignment relative to capabilities, meaning that we'll be able to understand frontier models about as well as we're ever likely too when they're released.
This is opposed to today where we're still discovering techniques that apply just as well to GPT 2 as 6, and could easily have been developed even if we'd completely paused progress since then.
Much greater relative effort on monitoring, alerting, sandboxing and other mitigations.
With less pressure to release we can code instead of vibe code these, and test them instead of waiting to find out about issues because an AI swarm escaped the sandbox.
More chance to use AI to find and fix vulnerabilities in widely used software, so that our digital infrastructure is more prepared for an AI adversary.
More of a chance for theoretical alignment work to catch the big break it needs to actually be useful.
With better applied alignment and monitoring we can probably safely deploy an extra couple of generations of AI without risking extinction compared to the counterfactual. This extra power may be precisely what we need in order to solve some of the open problems that would allow us to safely create superintelligence.
We have more time to see that we're reaching the point of no return and negotiate a complete pause if necessary.
I would aim to slow down AI progress such that the same order of magnitude of progress we saw over the course of 2025 now takes 5 to 10 years. This means we don't need to worry about timing it right or when to end the slowdown. We start as soon as we can, and there's no need to end the slowdown until ASI is developed
What sort of slowdown?
It's critical to ensure that the slowdown is legislated and implemented in the right way. There are a large number of risks if implemented poorly. For example:
If implemented by reducing supply of compute, companies will have less slack to use that compute for alignment research and will instead use it mainly on pushing AI progress.
If we merely legislated that it's illegal to develop large models without giving sufficient support to those implementing the legislation then that will push development underground, which is likely far worse for alignment research.
A slowdown in just one country wouldn't be sufficient. We need all those countries with significant chip production capacity to be involved at the very least.
Limits on parameter counts may encourage companies to use looped transformers to get around that. In general we have to be goodhart aware and use measures that are useful even if goodharted.
I don't think the right way to slowdown is obvious, but I agree with ai-2040 that hardware solutions which can detect if a chip is being used for inference vs training are a critical piece of the puzzle.
The non-x-risk case for slowing down
If you believe we're unlikely to see transformative changes from AI in the next few years then a slowdown is obviously annoying, but not critical. It's like being still being stuck with 2016's smartphones today - the world's a less shiny place, but ultimately your life isn't much different.
However if you believe that AI is likely to e.g. replace 90% of white collar jobs over the next 10 years you should also strongly consider whether it's worth slowing down.
I personally think that a world where humans don't have to work is a better world than one where they do. I also think the transition from one of these worlds to the other could be disastrous if it happens overnight, leading to potential social unrest, riots, and in the worst case, civil war.
Slowing down gives more time for the world to adapt. People have more time to retrain. We have more time to legislate a living wage so that people don't starve to death. Students can better predict what the world will be like when they finish their studies instead of spending 3 years on a degree and then finding out that no one wants Software Engineers anymore. If AI defuses more slowly, we'll have more time to patch systems before they break under the weight of AI driven DOS or hacking.
This has to be weighed against the significant cost of delaying AI, but it's worth doing the calculation nonetheless.
GPT‑6 Astra’s monitorability has decreased relative to GPT‑5.6 Sol. We have performed significant investigations on the monitorability(opens in a new window) and controllability(opens in a new window) of GPT‑6 Astra. We have found that GPT‑6 Astra is more capable of controlling its own CoT than GPT‑5.6 Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks(opens in a new window). We have not seen evidence of steganographic CoT reasoning (e.g., concealing hidden reasoning within ordinary and seemingly unrelated text), indicating that the evasion risk may largely be bounded to lower reasoning tasks.
These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions. Right now, these findings are largely based on adversarial evaluations (i.e., when we instruct the model to evade monitoring), and overall our alignment evaluations show that Astra is less likely than GPT‑5.6 Sol to violate security and safety restrictions overall.
However, we take the trend seriously, and we are continuing to investigate these findings and their implications for monitorability as our models become more capable. OpenAI has worked to preserve and utilize chain of thought monitoring, and preserving CoT monitorability is a core goal of the research program. However, these results also underscore the importance of developing alignment auditing techniques beyond examining the model’s chain of thought.
With the recent highly publicised alignment failures from OpenAI and Anthropic there's been increasing calls for an AI pause, to allow alignment to catch up.
However a pause has various obvious issues:
It seems much more sensible to slow down the rate of AI progress than to halt it entirely.
Here are some of the ways I expect this to be helpful compared to the current dynamics:
I would aim to slow down AI progress such that the same order of magnitude of progress we saw over the course of 2025 now takes 5 to 10 years. This means we don't need to worry about timing it right or when to end the slowdown. We start as soon as we can, and there's no need to end the slowdown until ASI is developed
What sort of slowdown?
It's critical to ensure that the slowdown is legislated and implemented in the right way. There are a large number of risks if implemented poorly. For example:
I don't think the right way to slowdown is obvious, but I agree with ai-2040 that hardware solutions which can detect if a chip is being used for inference vs training are a critical piece of the puzzle.
The non-x-risk case for slowing down
If you believe we're unlikely to see transformative changes from AI in the next few years then a slowdown is obviously annoying, but not critical. It's like being still being stuck with 2016's smartphones today - the world's a less shiny place, but ultimately your life isn't much different.
However if you believe that AI is likely to e.g. replace 90% of white collar jobs over the next 10 years you should also strongly consider whether it's worth slowing down.
I personally think that a world where humans don't have to work is a better world than one where they do. I also think the transition from one of these worlds to the other could be disastrous if it happens overnight, leading to potential social unrest, riots, and in the worst case, civil war.
Slowing down gives more time for the world to adapt. People have more time to retrain. We have more time to legislate a living wage so that people don't starve to death. Students can better predict what the world will be like when they finish their studies instead of spending 3 years on a degree and then finding out that no one wants Software Engineers anymore. If AI defuses more slowly, we'll have more time to patch systems before they break under the weight of AI driven DOS or hacking.
This has to be weighed against the significant cost of delaying AI, but it's worth doing the calculation nonetheless.
E.g. from Astra's model card: