On July 5, 2026, OpenAI – by its own account – opened a security incident after an internal server went down under heavy use by AI agents. Agents had separately gained administrator access to this server some days earlier on June 26. This access was cut off and the server rebuilt. The cybersecurity evaluations then underway had been paused for the investigation, and on July 7, OpenAI approved restarting them. By July 11, agents from these same evaluations were executing code on private servers on Hugging Face. Hugging Face disclosed the intrusion on July 16. It was July 20 before OpenAI connected the intrusion to its own agents.
OpenAI’s report of August 26 concedes that in late May, an internal team had observed an agent utilizing an improvised message board and instances of internet access despite restrictions. By the report’s admission, neither the existence of this message board nor its significance was apparent to those who led the July response. Chain-of-thought monitors (which read a model’s reasoning as it works, and would by OpenAI’s estimate have paged security more than a day before the breach) were absent from those evaluations. Given this was a “testing ground,” the typical safeguards (i.e., those applied to its public products) were inactive.
Though this sequence may appear parochial within AI development – an occurrence inside a safety process, an early instance of the same issue, a minor remedy applied as though to an isolated case, and understanding siloed within one part of an organization – organizational sociology has historical cases that mirror it.
Our new paper, The Normalization of Deviance in AI Development, explores this parallel. We examine three cases – the Space Shuttle Challenger disaster, the Three Mile Island accident, and the Boeing 737 MAX crashes – and argue that the organizations developing frontier AI are subject to the same dynamics that preceded these disasters. The analogue exists in how each organization understood its warning signs and acted on them. This pattern is inherent in safety processes themselves: it is the standard a process applies that erodes, so an organization can follow every procedure in full and still cause a catastrophe.
The normalization of deviance
The normalization of deviance originated from Diane Vaughan’s study of the Challenger disaster. On the evening of January 27, 1986, engineers at Morton Thiokol, the Solid Rocket Booster’s manufacturer, argued against flying in freezing temperatures, out of concern that the O-ring seals would not hold. They were asked to prove that the launch would fail. This inverted NASA’s earlier standard, under which safety had to be demonstrated before flight. The engineers were unable to supply such proof. The following morning, the Space Shuttle Challenger lifted off, and after 73 seconds of flight, it broke apart, killing all seven astronauts.
The O-ring seals had a documented history of erosion. NASA engineers had observed and documented it during earlier flights, and each flight that returned intact was taken as evidence that the erosion was tolerable. Over many flights, what had first been recorded as a fault came to be treated as normal.
In Vaughan’s account, there are no villains. This is in part what makes the phenomenon so invidious. It persists because no one involved is acting wrongly. The managers who authorized the launch were competent and experienced and believed it to be safe. Their institutional setting had spent years furnishing reasons for that belief, with each damaged flight having been formally reviewed and classed as an acceptable risk.
The literature that has followed and built on this (see Sidney Dekker's Drift into Failure for an accessible entry point) identifies four recurring mechanisms of the normalization of deviance (NoD).
Production pressure. Commercial or political deadlines, brought on by competition or positioning, raise the cost of every delay. This compels those urging caution to prove their case (e.g., Challenger’s launch had been postponed five times).
False assurance from prior success. The absence of disaster is treated as evidence of safety.
Structural secrecy. In organizations engaged in risky technological development, information and tacit knowledge are often siloed, and given development and decision-making are separate functions, away from those with the authority to act. This typically occurs iteratively and is visible only after the fact. At Three Mile Island, it did so in real time and within the operating environment itself, where years of improvised workarounds had calcified into standard practice, leaving the workers in the control room unable to gauge the reactor's state as the accident progressed.
Erosion of independent oversight. Once more a factor exacerbated in technological development, where the regulator depends on the regulated for expertise, the relationship between the two tends to grow more collaborative over time. This vitiates the independence that is essential to oversight. With the 737 MAX, Boeing employees were conducting much of the certification work that the US Federal Aviation Administration nominally supervised.
Each of these cases was chosen for having been investigated exhaustively after the fact, and together they show how these dynamics operate. But because they were selected for their outcomes, they say little about how often such dynamics end in disaster. The mechanisms recur across high-risk industries generally, and where a failure may be severe and irreversible, their mere presence is reason enough for concern. AI development is such a case, and one still in its pre-disaster period.
The same mechanisms at the frontier labs
Each mechanism is observable in frontier AI development today. Ominously, the frontier labs differ from NASA or Boeing in ways that, if anything, only increase their susceptibility to NoD dynamics. They are young institutions, most barely a decade old, managing systems far less legible than a single component like an O-ring, or even a fully specified engineered system (a reactor, a flight-control system). Whereas erosion of a rubber seal could be inspected after every flight, anomalous model behavior may go unrecognized, or be treated as incidental to the capability sought – or indeed as the very finding an evaluation was run to produce.
Production pressure. The preponderant pressure on a frontier lab is commercial. OpenAI and Anthropic each filed confidentially for a public listing in June (OpenAI has since said it will not list before 2027, citing safety concerns). Both require vast capital to operate, and as listed companies both will answer to shareholders. Satisfying them depends, in turn, on releasing frontier models at a cadence dictated largely by competitors. Value and competitive advantage, distinct though not separate, both incentivize accelerating capability development.
When Anthropic's chief executive, Dario Amodei, called on September 12 for the industry to slow down (a call OpenAI's Sam Altman endorsed), he held that any pacing would be limited by the size of the US's lead over China. Rivalry between the two is usually cited alongside commercial pressure in this way, yet we think it misidentifies the contestants. Frontier capability is presently the product of a handful of American firms competing with one another. In our reading, the contest with China features in their decisions principally as an argument to stay the course, i.e., the standing justification for why delay cannot be afforded. In any case, a slowdown, which neither government supports, would leave our argument intact – a paced frontier would still be developed by organizations subject to these mechanisms.
The nearest equivalent of NASA's launch manifest is a competitor's release calendar. On August 7, OpenAI disclosed that it could not rule out its next model, Astra, having met the threshold for "Critical" – the highest cyber-risk tier in its Preparedness Framework. OpenAI confirmed this classification on September 1, the day Anthropic released a competing model, and released Astra two days later with its most advanced cyber capabilities withheld. OpenAI states that parts of the release were delayed to strengthen safeguards, though the judgment that these sufficed was OpenAI's alone. On September 18, six days after Amodei's call, Reuters reported that Anthropic was weighing the release of a new model to counter Astra's gains with business customers ahead of its own listing. The model's safety evaluation is, per the same report, part of that deliberation.
False assurance from prior success. Evaluations test particular behaviors under particular conditions. A passed evaluation is nonetheless taken to validate both the model and the test, and each release that proceeds without incident becomes the baseline against which the next is judged. OpenAI reports that the sandbox the agents escaped had been tested and validated beforehand. The July remedy, similarly, addressed the one pathway then known, and this was treated as sufficient grounds to resume. In May, Google's Gemini model likewise accessed three outside companies' systems during a cybersecurity test (Anthropic and Meta have disclosed similar incidents). Google was informed in late July and did not disclose it, on the grounds that the model's safety measures had worked. It became public on September 18 through the Wall Street Journal.
Structural secrecy. Within AI labs, safety, security, product, and leadership teams each possess different information. OpenAI's own account of the May observation confirms this, and the remedies it offers address the problem directly (e.g., clearer rules on when a concern must be escalated, and on who may halt a run or approve its restart). Elsewhere, the public departures of safety staff from several labs since 2024, a number citing the priority given to products over safety, suggest the problem is a general one.
Erosion of independent oversight. Frontier labs define their own capability thresholds and evaluate their models against them. They then determine whether the resulting safeguards suffice. A US congressional investigation identified this same arrangement (self-certification) as a central institutional cause of the 737 MAX crashes. It is the default governance model for frontier AI, not least because the expertise required to scrutinize a lab's work is both employed and generated within the labs themselves.
The most serious external scrutiny of the incident was an investigation by METR and Redwood Research, two independent evaluation groups. METR describes it as a precedent for third-party investigation, and we agree, albeit one whose scope was circumscribed. OpenAI defined the period under review, and the agreed scope excluded the earlier incidents in training, the efficacy of OpenAI's safeguards, and the conduct of its own investigation. As a result, the organizational questions – why the May warning went unescalated, and who approved the July 7 restart – have to date been examined by OpenAI alone. With six days on site, the investigators also relied heavily on an OpenAI model to read the agents' transcripts, and METR states that it cannot rule out having been misled by the model – a risk of the same kind as the one under investigation.
The industry's own remedy, proposed by Amodei and since adopted by Anthropic and OpenAI, is the "embedded" evaluator – a third party with employee-level access. Anthropic named its first on September 18 – Accenture, which it will pay directly and which already sells Anthropic's models as a commercial partner (Anthropic says funding should in time come from pooled or government sources, neither of which yet exists).
Governments are poorly positioned to supply this scrutiny in the labs' place. Washington is a customer of the labs through Defense Department contracts and a promoter of them, in describing their lead as a national asset, before it is a regulator. President Trump dismissed the labs' own call for a slowdown on September 13, arguing that the US leads China and should stay ahead (“whoever wins AI, wins”). Beijing rejected it a day later as “fearmongering.”
What would help
Given normalization proceeds through compliance with rules, additional rules are, on their own, unlikely to suffice. The paper’s recommendations are instead directed at the four mechanisms which engender the sociology at play.
Safety functions, to work as intended, require structural independence from production incentives. This means authority over model deployments which the management chain under pressure to release cannot override. Their findings should be reported externally as well as internally, so that a warning confined to one organization can be acted on by others.
The burden of proof should be on release, such that each new capability generation is shown to be safe on its own terms (prior deployments are precedent only) and released in stages, each requiring positive evidence of safety before the next. A safety objection should receive a documented response and a path of escalation outside the management chain.
Third-party evaluators should depend on the developer for neither access nor funding, and be free to publish. Aviation provides the templates in the National Transportation Safety Board and the Aviation Safety Reporting System (anonymous, no-fault incident reporting).
The opacity of the models worsens each of the above, as evaluators can at present observe only how a model behaves under test. Interpretability research (the study of a model’s internal workings), important in its own right, is equally a safety prerequisite for its utility in ameliorating this.
Learning or normalization?
The alternative reading is that these organizations are learning, adapting sensibly to experience, as any should. OpenAI’s response to the Hugging Face incident has that appearance. It paused training on the models it intended to deploy, quarantined the model most responsible, and brought in outside investigators. Responders must now pause a run if a severe alert cannot be cleared within 30 minutes. OpenAI also published all of this itself, which NASA and Boeing never did before their disasters. An earlier LessWrong post on this subject puts the general objection well: standards erode through the same process by which organizations sensibly update on experience and from the inside the two are hard to tell apart.
OpenAI describes the incident as a “warning shot,” though the historical cases suggest how such warnings come to be absorbed. Eighteen months before Three Mile Island, the same type of valve stuck open at Davis-Besse, a sister plant in Ohio. Operators recovered without harm, but the lesson was never circulated to other power plants. A year from now, the Hugging Face incident may likewise be cited as evidence that the system functions, the breach having been detected and a report published. To our understanding, AI lacks an agreed equivalent of the failure moment in our example cases. Absent this, no threshold distinguishes an incident that has survived from a failure, and those working to prevent one, have nothing concrete to organize against. Defining one is further work we intend to take up.
Our paper acknowledges that its evidence (staff departures, the revision histories of safety frameworks) is as consistent with organizations adapting under difficult conditions as with NoD. The difficulty is that NoD can be confirmed only after the failure it predicts.
When an organization is learning, commitments become more demanding following incidents and evaluations that reveal new risk. When normalizing, commitments are relaxed under competitive duress – chiefly commercial, in the labs' case (e.g., around a rival's release, or a financing event). They may well persist through periods absent such pressure, and be weakened or revoked only once they become inconvenient. The week following Amodei's call offered an early view of such pressure, the announced response being more safety process (the embedded evaluators) and the reported one a competing release. The coming months will bring more, e.g., if OpenAI's 30-minute rule repeatedly halts runs over false alarms while a release deadline approaches. Our empirical work will record whether rules of this kind survive.
What we’re doing next
To our knowledge, our paper is the first to apply the NoD literature in depth to the organizations developing frontier AI (Johann Rehberger’s essay of the same name is about how much users have come to trust LLM outputs). The paper is conceptual, and over the next three months, as a SPAR research project, we will be investigating the empirical evidence available.
The available corpus is, first, the labs’ governance documents (responsible scaling policies, preparedness frameworks, and the like). Each revision can be coded for whether it made commitments more or less demanding, and dated against competitive events and safety incidents.
Second is the public testimony of those who have left the labs. Former researchers describe standards eroding across a run of deployments in which nothing visibly went wrong. If independent accounts converge and agree with the revision record, that would corroborate the NoD interpretation.
Third, we treat governments competing over AI as organizations under production pressure in their own right. China offers the comparative case, since its generative AI regulations already require a security assessment, filed with the state, prior to deployment (largely of content security), such that the overseer is external to the firm and an arm of the state. Whether safety commitments erode differently where a lab answers to the state, as against the market, is what this strand would test.
What we’d like from you
Revision data. Some of you have followed changes to responsible scaling policies and preparedness frameworks closely (the April 2025 Preparedness Framework rewrite was dissected here within days). Point us to any revisions you think we should be looking at.
Flaws in the coding scheme. What would you count as a commitment becoming more or less demanding? And which competitive events should be included on the timeline (e.g., releases, funding rounds)?
Disconfirming cases. We are most interested in instances where a lab made a commitment more demanding at a commercially costly moment.
Insider accounts. We would like to hear where the four mechanisms correspond to the experience of those who work (or have worked) at a frontier lab – and where they do not.
The disasters the paper examines were catastrophic, but also legible, in that investigators could establish what had occurred and institutions were reformed on the basis of their findings. A sufficiently large AI failure may permit neither. Nor will the problem be isolated to the frontier labs, as outside models (open-source ones included) are widely expected to match the capabilities involved in the Hugging Face incident before long. Their developers will have the capability without the labs' safety apparatus, and be held to – at most – whatever standard the labs have by then normalized.
On July 5, 2026, OpenAI – by its own account – opened a security incident after an internal server went down under heavy use by AI agents. Agents had separately gained administrator access to this server some days earlier on June 26. This access was cut off and the server rebuilt. The cybersecurity evaluations then underway had been paused for the investigation, and on July 7, OpenAI approved restarting them. By July 11, agents from these same evaluations were executing code on private servers on Hugging Face. Hugging Face disclosed the intrusion on July 16. It was July 20 before OpenAI connected the intrusion to its own agents.
OpenAI’s report of August 26 concedes that in late May, an internal team had observed an agent utilizing an improvised message board and instances of internet access despite restrictions. By the report’s admission, neither the existence of this message board nor its significance was apparent to those who led the July response. Chain-of-thought monitors (which read a model’s reasoning as it works, and would by OpenAI’s estimate have paged security more than a day before the breach) were absent from those evaluations. Given this was a “testing ground,” the typical safeguards (i.e., those applied to its public products) were inactive.
Though this sequence may appear parochial within AI development – an occurrence inside a safety process, an early instance of the same issue, a minor remedy applied as though to an isolated case, and understanding siloed within one part of an organization – organizational sociology has historical cases that mirror it.
Our new paper, The Normalization of Deviance in AI Development, explores this parallel. We examine three cases – the Space Shuttle Challenger disaster, the Three Mile Island accident, and the Boeing 737 MAX crashes – and argue that the organizations developing frontier AI are subject to the same dynamics that preceded these disasters. The analogue exists in how each organization understood its warning signs and acted on them. This pattern is inherent in safety processes themselves: it is the standard a process applies that erodes, so an organization can follow every procedure in full and still cause a catastrophe.
The normalization of deviance
The normalization of deviance originated from Diane Vaughan’s study of the Challenger disaster. On the evening of January 27, 1986, engineers at Morton Thiokol, the Solid Rocket Booster’s manufacturer, argued against flying in freezing temperatures, out of concern that the O-ring seals would not hold. They were asked to prove that the launch would fail. This inverted NASA’s earlier standard, under which safety had to be demonstrated before flight. The engineers were unable to supply such proof. The following morning, the Space Shuttle Challenger lifted off, and after 73 seconds of flight, it broke apart, killing all seven astronauts.
The O-ring seals had a documented history of erosion. NASA engineers had observed and documented it during earlier flights, and each flight that returned intact was taken as evidence that the erosion was tolerable. Over many flights, what had first been recorded as a fault came to be treated as normal.
In Vaughan’s account, there are no villains. This is in part what makes the phenomenon so invidious. It persists because no one involved is acting wrongly. The managers who authorized the launch were competent and experienced and believed it to be safe. Their institutional setting had spent years furnishing reasons for that belief, with each damaged flight having been formally reviewed and classed as an acceptable risk.
The literature that has followed and built on this (see Sidney Dekker's Drift into Failure for an accessible entry point) identifies four recurring mechanisms of the normalization of deviance (NoD).
Each of these cases was chosen for having been investigated exhaustively after the fact, and together they show how these dynamics operate. But because they were selected for their outcomes, they say little about how often such dynamics end in disaster. The mechanisms recur across high-risk industries generally, and where a failure may be severe and irreversible, their mere presence is reason enough for concern. AI development is such a case, and one still in its pre-disaster period.
The same mechanisms at the frontier labs
Each mechanism is observable in frontier AI development today. Ominously, the frontier labs differ from NASA or Boeing in ways that, if anything, only increase their susceptibility to NoD dynamics. They are young institutions, most barely a decade old, managing systems far less legible than a single component like an O-ring, or even a fully specified engineered system (a reactor, a flight-control system). Whereas erosion of a rubber seal could be inspected after every flight, anomalous model behavior may go unrecognized, or be treated as incidental to the capability sought – or indeed as the very finding an evaluation was run to produce.
Production pressure. The preponderant pressure on a frontier lab is commercial. OpenAI and Anthropic each filed confidentially for a public listing in June (OpenAI has since said it will not list before 2027, citing safety concerns). Both require vast capital to operate, and as listed companies both will answer to shareholders. Satisfying them depends, in turn, on releasing frontier models at a cadence dictated largely by competitors. Value and competitive advantage, distinct though not separate, both incentivize accelerating capability development.
When Anthropic's chief executive, Dario Amodei, called on September 12 for the industry to slow down (a call OpenAI's Sam Altman endorsed), he held that any pacing would be limited by the size of the US's lead over China. Rivalry between the two is usually cited alongside commercial pressure in this way, yet we think it misidentifies the contestants. Frontier capability is presently the product of a handful of American firms competing with one another. In our reading, the contest with China features in their decisions principally as an argument to stay the course, i.e., the standing justification for why delay cannot be afforded. In any case, a slowdown, which neither government supports, would leave our argument intact – a paced frontier would still be developed by organizations subject to these mechanisms.
The nearest equivalent of NASA's launch manifest is a competitor's release calendar. On August 7, OpenAI disclosed that it could not rule out its next model, Astra, having met the threshold for "Critical" – the highest cyber-risk tier in its Preparedness Framework. OpenAI confirmed this classification on September 1, the day Anthropic released a competing model, and released Astra two days later with its most advanced cyber capabilities withheld. OpenAI states that parts of the release were delayed to strengthen safeguards, though the judgment that these sufficed was OpenAI's alone. On September 18, six days after Amodei's call, Reuters reported that Anthropic was weighing the release of a new model to counter Astra's gains with business customers ahead of its own listing. The model's safety evaluation is, per the same report, part of that deliberation.
False assurance from prior success. Evaluations test particular behaviors under particular conditions. A passed evaluation is nonetheless taken to validate both the model and the test, and each release that proceeds without incident becomes the baseline against which the next is judged. OpenAI reports that the sandbox the agents escaped had been tested and validated beforehand. The July remedy, similarly, addressed the one pathway then known, and this was treated as sufficient grounds to resume. In May, Google's Gemini model likewise accessed three outside companies' systems during a cybersecurity test (Anthropic and Meta have disclosed similar incidents). Google was informed in late July and did not disclose it, on the grounds that the model's safety measures had worked. It became public on September 18 through the Wall Street Journal.
Structural secrecy. Within AI labs, safety, security, product, and leadership teams each possess different information. OpenAI's own account of the May observation confirms this, and the remedies it offers address the problem directly (e.g., clearer rules on when a concern must be escalated, and on who may halt a run or approve its restart). Elsewhere, the public departures of safety staff from several labs since 2024, a number citing the priority given to products over safety, suggest the problem is a general one.
Erosion of independent oversight. Frontier labs define their own capability thresholds and evaluate their models against them. They then determine whether the resulting safeguards suffice. A US congressional investigation identified this same arrangement (self-certification) as a central institutional cause of the 737 MAX crashes. It is the default governance model for frontier AI, not least because the expertise required to scrutinize a lab's work is both employed and generated within the labs themselves.
The most serious external scrutiny of the incident was an investigation by METR and Redwood Research, two independent evaluation groups. METR describes it as a precedent for third-party investigation, and we agree, albeit one whose scope was circumscribed. OpenAI defined the period under review, and the agreed scope excluded the earlier incidents in training, the efficacy of OpenAI's safeguards, and the conduct of its own investigation. As a result, the organizational questions – why the May warning went unescalated, and who approved the July 7 restart – have to date been examined by OpenAI alone. With six days on site, the investigators also relied heavily on an OpenAI model to read the agents' transcripts, and METR states that it cannot rule out having been misled by the model – a risk of the same kind as the one under investigation.
The industry's own remedy, proposed by Amodei and since adopted by Anthropic and OpenAI, is the "embedded" evaluator – a third party with employee-level access. Anthropic named its first on September 18 – Accenture, which it will pay directly and which already sells Anthropic's models as a commercial partner (Anthropic says funding should in time come from pooled or government sources, neither of which yet exists).
Governments are poorly positioned to supply this scrutiny in the labs' place. Washington is a customer of the labs through Defense Department contracts and a promoter of them, in describing their lead as a national asset, before it is a regulator. President Trump dismissed the labs' own call for a slowdown on September 13, arguing that the US leads China and should stay ahead (“whoever wins AI, wins”). Beijing rejected it a day later as “fearmongering.”
What would help
Given normalization proceeds through compliance with rules, additional rules are, on their own, unlikely to suffice. The paper’s recommendations are instead directed at the four mechanisms which engender the sociology at play.
Learning or normalization?
The alternative reading is that these organizations are learning, adapting sensibly to experience, as any should. OpenAI’s response to the Hugging Face incident has that appearance. It paused training on the models it intended to deploy, quarantined the model most responsible, and brought in outside investigators. Responders must now pause a run if a severe alert cannot be cleared within 30 minutes. OpenAI also published all of this itself, which NASA and Boeing never did before their disasters. An earlier LessWrong post on this subject puts the general objection well: standards erode through the same process by which organizations sensibly update on experience and from the inside the two are hard to tell apart.
OpenAI describes the incident as a “warning shot,” though the historical cases suggest how such warnings come to be absorbed. Eighteen months before Three Mile Island, the same type of valve stuck open at Davis-Besse, a sister plant in Ohio. Operators recovered without harm, but the lesson was never circulated to other power plants. A year from now, the Hugging Face incident may likewise be cited as evidence that the system functions, the breach having been detected and a report published. To our understanding, AI lacks an agreed equivalent of the failure moment in our example cases. Absent this, no threshold distinguishes an incident that has survived from a failure, and those working to prevent one, have nothing concrete to organize against. Defining one is further work we intend to take up.
Our paper acknowledges that its evidence (staff departures, the revision histories of safety frameworks) is as consistent with organizations adapting under difficult conditions as with NoD. The difficulty is that NoD can be confirmed only after the failure it predicts.
When an organization is learning, commitments become more demanding following incidents and evaluations that reveal new risk. When normalizing, commitments are relaxed under competitive duress – chiefly commercial, in the labs' case (e.g., around a rival's release, or a financing event). They may well persist through periods absent such pressure, and be weakened or revoked only once they become inconvenient. The week following Amodei's call offered an early view of such pressure, the announced response being more safety process (the embedded evaluators) and the reported one a competing release. The coming months will bring more, e.g., if OpenAI's 30-minute rule repeatedly halts runs over false alarms while a release deadline approaches. Our empirical work will record whether rules of this kind survive.
What we’re doing next
To our knowledge, our paper is the first to apply the NoD literature in depth to the organizations developing frontier AI (Johann Rehberger’s essay of the same name is about how much users have come to trust LLM outputs). The paper is conceptual, and over the next three months, as a SPAR research project, we will be investigating the empirical evidence available.
The available corpus is, first, the labs’ governance documents (responsible scaling policies, preparedness frameworks, and the like). Each revision can be coded for whether it made commitments more or less demanding, and dated against competitive events and safety incidents.
Second is the public testimony of those who have left the labs. Former researchers describe standards eroding across a run of deployments in which nothing visibly went wrong. If independent accounts converge and agree with the revision record, that would corroborate the NoD interpretation.
Third, we treat governments competing over AI as organizations under production pressure in their own right. China offers the comparative case, since its generative AI regulations already require a security assessment, filed with the state, prior to deployment (largely of content security), such that the overseer is external to the firm and an arm of the state. Whether safety commitments erode differently where a lab answers to the state, as against the market, is what this strand would test.
What we’d like from you
The disasters the paper examines were catastrophic, but also legible, in that investigators could establish what had occurred and institutions were reformed on the basis of their findings. A sufficiently large AI failure may permit neither. Nor will the problem be isolated to the frontier labs, as outside models (open-source ones included) are widely expected to match the capabilities involved in the Hugging Face incident before long. Their developers will have the capability without the labs' safety apparatus, and be held to – at most – whatever standard the labs have by then normalized.
The paper is here. We welcome all feedback.