This could be appropriately dealt with through compute governance to make sure that there aren’t relevant amounts of AI chips in potential blacksites.
Or, if what you’re saying is that countries will just refuse access to inspectors at the sites that are visible, then you can start reaching into the coercive toolbox to punish that behavior and slow down the classified projects regardless.
compute governance [without inspection]
my understanding is that this is a hard problem, particularly where we attempt to separate "research" from "inference"
sites that are visible
more that if a site is classified, it will not be "visible" and so not available to inspectors.
coercive toolbox to punish that behavior
I think the core problem is that the national security apparatus will resist surrendering an advantage over a strategically important technology. but certainly politics is the art of the possible, as they say
Compute governance is only necessary to the extent that you can confidently know where large amounts of compute are, which I would describe as a relatively easy problem given how hard it is to hide fabs and how simple choking off production could be. In order to sabotage their operations, you don't need to actually know what they're using them for: the very fact that they're not revealing it is a justification for sabotage.
This same basic setup is already how nuclear enrichment sites are handled. Nuclear enrichment facilities are very hard to hide outright, so states can be very confident about whether a rival is able to pursue a nuclear weapons program. If enrichment sites are observed, and if IAEA inspectors are refused access, then states start applying coercive leverage up to and including outright destroying the facilities.
Linkpost for a piece we recently published for AI Frontiers in the wake of recent calls for slowdown, covering how an international verification effort be trivially enforced by using human inspectors alone and the joint incentives for implementing one.
On July 28, over a thousand employees of the world’s top AI companies advocated that the US government “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Following months of cybersecurity scares and their own loss of control incidents, both OpenAI and Anthropic officially endorsed the same message, recognizing the danger of blindly accelerating AI development.
Despite this new urgency, many people argue that coordinating an international slowdown is currently unworkable—including some of the same groups in favor of one. Namely, they claim that even if the US slowed down its own AI development, it wouldn’t be able to make sure that China was doing the same, leaving the US no choice but to race.
These anti-slowdown arguments usually emphasize the technical challenges with designing AI monitoring measures to ensure a slowdown is being respected. In particular, slowdown skeptics argue that countries would refuse to install verification measures unless they were privacy-preserving enough to avoid leaking trade secrets and other sensitive information. They then point out that any near-term privacy-preserving verification technology would not be adversarially robust to nation-state attacks. If the Chinese government were actively trying to subvert a slowdown, it could use advanced techniques to fake workloads and spoof locations, including physical key extraction or even laser bit-flipping. Getting to the point where we can start implementing an international slowdown, the argument goes, might have to wait on years of R&D to solve both these problems.
This perspective is missing the forest for the trees. Privacy-preserving verification tools that are immune to state adversaries would be ideal, but they are not a prerequisite for an AI slowdown. Robustly verifying limits on AI development is already feasible with low-tech methods. If there were enough political will to implement a joint slowdown, states could extensively monitor labs for unsanctioned behavior by simply placing human inspectors on the inside—an approach we call Whole-Lab Inspection (WLI).
Under this proposal, inspectors from the US and China would receive broad physical access to frontier AI companies alongside read-only access to corporate communications, worklogs, code repositories, and compute telemetry. Using monitoring systems the companies already operate, they could observe operations and investigate potential violations—giving them the visibility of a CISO without requiring the powers of an operator. By literally and figuratively looking over the shoulders of lab employees, inspectors would likely be capable of robustly verifying fine-grained restrictions on AI development, such as bans on autonomous AI R&D and limitations on post-training. Overall, WLI demonstrates that an international slowdown is already technically possible; it’s just a question of whether the US and China will have the political will to jointly implement it.
In the near future, political will to implement a joint slowdown could rapidly appear. If developed and diffused quickly, advanced AI systems are likely to intensely destabilize society by democratizing access to weapons of mass destruction and fueling waves of unemployment, regardless of whether they’re developed at home or abroad. On top of these threats, allowing an intelligence explosion would risk military escalation and loss of control, as states race to sabotage and deploy powerful AI systems for strategic advantage. Should the US slow down its own AI development domestically in anticipation of this instability and offer a reciprocal audit, China would have strong incentives to agree—both to ensure the U.S. does not recklessly deploy its own AI systems, and to avoid the chance of spiraling escalation over AI development.
In what follows, we’ll argue that whole-lab inspections can be used to robustly enforce fine-grained limits on AI development. Then, we’ll speak to the incentives that might encourage the US and China to pursue a joint slowdown at some point using WLI. Finally, we’ll discuss the stability and risk reduction benefits of a well-timed slowdown.
Enforcing a Slowdown Through Whole-Lab Inspections
In order to agree to a slowdown, the US and China would need to know what research was being conducted inside each other’s frontier labs. Fortunately, visibility into AI development could be achieved by sending human auditors to comprehensively monitor labs.
Auditors can have high visibility into AI companies. Monitoring for disallowed training runs or prohibited research activities may be largely trivial with human inspectors. In large part, this is because AI companies already have extensive internal monitoring systems. Most AI companies record employee screens and keystrokes to produce training data and help with automation. Companies also constantly save logs from experiments and jobs submitted to keep track of compute use and archive research progress. Giving inspectors comprehensive visibility into a lab is likely as simple as integrating and handing over read-only permissions to internal logs and company communications. To aid in their oversight, inspectors might direct a company’s own AI agents, or trusted external AIs, to read through troves of company data in search of potential violations for further review.
Alongside this digital information, inspectors could also work on-site alongside lab and datacenter employees, giving them physical access to hardware and research projects. Just as for nuclear power plants, banks, or any other high-risk industry, inspectors would be able to question employees about their work, as well as physically inspect infrastructure by hand for compliance. By using such a broad set of simple verification techniques, it would be extremely difficult to conduct an unapproved training run unnoticed—like trying to get a plane off a runway without getting noticed by an air-traffic controller.
If whole lab inspection were established, governments could then enforce a wide range of policies to slow AI development:
Preventing an intelligence explosion. The most dangerous path for AI development to take would be a full-throttle intelligence explosion—using AIs to autonomously design new, even smarter AI systems recursively. Since AIs can work much faster than humans, and since this process might quickly result in AI systems more capable than their human monitors, it’s likely that it would become impossible for humans to effectively oversee their AIs or intervene if they misbehave. To prevent this, inspectors could enforce targeted interventions designed to specifically restrict autonomous AI R&D, such as by triaging the jobs that are using the most compute and auditing whether they’re being used for capabilities research. On top of these restrictions, inspectors could limit the length of time that agents are allowed to work autonomously, monitor the amount of compute allocated to internal inference, and prevent AIs from taking high-stakes actions like starting training runs or authorizing new deployments.
Limits on post-training. Given the wide range of potential risks from AI, governments might also want to expand training restrictions beyond just interventions against autonomous AI R&D. One natural way to do this would be to place limits on post-training. By measuring the increase in capability and efficiency of a model after post-training, inspectors would have clear criteria to disallow the deployment of a model or the kinds of training environments that produced it. Likewise, inspectors could monitor which broad capabilities post-training datasets are targeting—such as coding, computer use, or medical diagnosis—and (dis)allow them on that basis. Since auditors would be able to directly observe which capabilities post-training datasets target and affect, this would allow them to cleanly whitelist certain types of post-training, even in a slowdown. These could include direct safety training to improve AIs’ adversarial robustness or propensity not to misbehave, as well as narrow AI-for-science efforts aimed at prosocial domains like medicine and formally verified code.
Even outside of these interventions, whole-lab inspection would still allow for a variety of potential regulations. Human auditors could be used to ensure, for instance, that specific datacenters are retrofitted to be inference-only, and thus incapable of pre-training, or that kill switches can quickly shut down AIs. On its own, joint verification of this kind could be sufficient for many years of delay, during which the US and China would have ample time to develop more thorough and technically sophisticated verification measures.
Joint Verification
Whole-lab inspection is straightforward and effective. If prompted by a domestic political shock, it’s likely it could be rapidly implemented and enforced at home, making it extremely difficult for any company under inspection to conduct unauthorized training runs. In order to extend this initiative internationally, however, the US and China would need to implement joint verification: granting inspectors on-site access to each state’s frontier labs.
Thankfully, this agreement would not need to rely on goodwill or trust. Because advanced AI could be globally destabilizing regardless of where it’s developed, it’s in each country’s self-interest to ensure the other is not proceeding recklessly with development. Moreover, if either country refuses to accept joint inspection, they would likely struggle to hide their AI projects and insulate them from theft or sabotage.
The US and China may experience simultaneous political pressure to slow down. Many of the most pressing risks from advanced AI are symmetric. The development and diffusion of AIs that cause mass unemployment, enable terrorists, or wreak havoc as rogue agents would cause global turmoil regardless of which country created them. Similarly, the fear of military dominance could provoke a risky and wasteful security dilemma, in which retaliatory sabotage spirals into broader conflict. Therefore, in order to maintain security and stability, the US and China need to ensure that the other is not deploying their own AI systems recklessly—something they would have no way of assuring in an all-out race.
AI companies’ trade secrets are easy for states to steal. Under mutual whole-lab inspection, auditors would have wide access to data about AI development at opposing AI companies, but this wouldn’t represent much of a change from the status quo. Today, information about model development flows freely between and within frontier labs, as employees constantly churn through roles, publish research, and put breakthroughs on the company Slack. Even once these obvious holes are sealed, state-proofing AI development against theft would remain enormously difficult. Human insiders—especially Chinese nationals at US frontier labs—could be bribed or coerced into espionage. AI development is also inherently vulnerable to cyberattacks, with anything from exploits in networking software to smuggling in spyware through the hardware supply chain offering an opportunity to steal research and models. As a result of these weaknesses, any country that tries to go full-throttle on development would likely quickly find its secrets sieving out of its frontier labs, pulling up its rival to algorithmic parity regardless.
Compute tracking would make it difficult to hide AI projects. The remaining priority would be to ensure that there are no large AI projects hidden from inspectors. Fortunately, the AI supply chain is so concentrated, and chip production so prohibitive to hide, that auditing the major semiconductor and chip suppliers would let the US and China confidently measure the number of chips that exist. In practice, monitoring large datacenters could be as straightforward as forcing the leading chip producers to hand over a detailed record of their production and sales, comparing against the declared compute from the visible projects, and then analyzing satellite imagery and electricity usage as an additional precaution.
Given the factors above, a global whole lab inspection initiative could be mutually compatible, low cost, and high confidence—as long as inspectors are given physical access to frontier labs and datacenters. Even if the US or China refuses or later revokes that access, however, AI development could still be slowed unilaterally through deterrence. The global chip supply chain is the most fragile and specialized industry in the entire world: if either country wanted to, they could apply massive friction through export controls or sabotage of key suppliers. Likewise, countries could disrupt AI training through subtle cyberattacks on training runs, such as by poisoning data to insert malicious backdoors or Stuxnet-style attacks on GPUs.
Benefits of Slowdown
Ultimately, the point of a slowdown would be to help society adapt to the risks that rapid development and diffusion of AI would introduce. So far, these risks have been mild enough, and introduced slowly enough, that the government and private industry have been able to reactively address them. By slowing the improvement in general AI capabilities, governments can ensure both that society is able to continue reacting to new risks and that sufficient investments in risk mitigation are made.
A slowdown would provide time for society to adapt to advanced AI. If threats from advanced AI appear too quickly—like the introduction of new weapons of mass destruction, or a sudden military confrontation over AI development—governments will not have time to assess and reactively regulate them, forcing them to rely on emergency measures. By proactively slowing down AI development, policymakers, as well as the rest of society, would be able to react to what would have otherwise been seismic shocks. Politically, voters and Congress would be able to decide how the benefits of AI will be distributed, and the degree of transparency AI developers owe the public. Internationally, states would be able to prepare for the introduction of powerful new military technologies before they arrive and invest in more robust verification measures.
In this sense, a mild amount of government intervention early on would avoid the need for massive overreach later, keeping the pace of AI development at a level society can react to.
The time bought by a slowdown could be used to invest directly in risk mitigation. Without the intense pressure to race and increase general capabilities, investment could instead be diverted into ensuring that AI systems are controllable and that they are developed for prosocial causes. Key open problems in AI safety, for example, include the lack of robust solutions to jailbreaking or the propensity for reinforcement learning to encourage models to go rogue—issues that will likely produce disastrous results if unresolved when models are more capable. By scaling up training in domains like adversarial robustness and behavioral propensity, AI developers could spend huge amounts of compute driving down jailbreak susceptibility and attempts to escape. Other safety measures could include aggressively sandboxing AI systems, such as with fully airgapped training setups, or preserving and expanding practical interpretability tools.
An AI Slowdown Does Not Require New Technology
Slowing down AI development doesn’t need to wait on any breakthroughs. On the technical level, whole-lab inspections could be a robust and flexible way to monitor whether frontier labs are complying with restrictions on the most dangerous kinds of development. Human auditors, reading through worklogs and standing on the floor next to the people recording them, could verify nuanced restrictions on autonomous AI R&D and post-training without needing any new inventions.
From there, a global slowdown is just a question of political will—one which has been convincingly answered before. By the end of the 1980s, having stared into the abyss during the peak of the nuclear arms race, even the US and Soviet Union could recognize the need for mutual disarmament. For more than two decades, even during years as fraught as the collapse of the Soviet Union itself, this was achieved by simply allowing Russian and American inspectors to walk into each other’s military bases and count warheads by hand.
Likewise, the US and China may soon realize that averting disaster at home will depend on verifying development abroad. If inspectors could count nuclear warheads in Siberia and North Dakota, they can certainly check worklogs in Silicon Valley and Shenzhen. We do not need better technology to verify a slowdown—only shared risks serious enough to demand one.
Special thanks to Dan Hendrycks for suggesting the premise of whole-lab inspection.