Leniency toward the frontier AI labs' communication and scientific output, as described below, is an inapt mindset at the current stage of the AI race. Stakes are too high, the pace of capabilities progress is too fast, and the incentives are too bad for leniency and blind faith (as well as mere blindness) to be appropriate. The dynamics of the race impose much more adversarial reactions to the frontier-labs' great efforts at scope control in external evaluations and frame control in reporting. I will purposefully illustrate the stance I am criticizing with best-case examples of it. There is a wide discrepancy, which will be transparent to the LessWrong community, between safety mindsets and leniency toward the labs supposed to implement them. Advocating for increased, sustained and adjusted external pressure on the frontier-labs, I will give a view on how and when and how much to alleviate adversariality in the public response.
In a lenient mood?
The online discourse on AI progress still suffers from a certain mood, persisting from a prevailing attitude more fitted to the beginning of the AI race. It seems that, perhaps because of social overlap between the alignment community and the labs, or more generally because individuals working at AI labs are mostly known to be well-meaning, competent people, the AI labs themselves are still treated as good-faith actors whom one should simply query and debate for the situation to improve. This friendly, lenient and somewhat subservient mood is illustrated by “help us help you help us” requests[1] to people who are either in no position to change things internally, or have strong incentives to disregard them. More importantly, it might actively facilitate the frontier-labs' race to extinction in the near future.
Researchers and engineers at OpenAI and Anthropic have recently written reports (1,2) on incidents, models and alignment (3,4) that display a dilettante attitude toward safety concerns, in view of the daunting, explicitly stated task of aligning a future ASI. Chief Scientist at OpenAI Jakub Pachocki recently published an article (5) mainly conveying the idea that OpenAI is treating its mission of bringing forth an aligned ASI in a very responsible way[2]. More broadly, Altman and Amodei's older essays (6,7) do little in the way of discussing alignment, but marginalize voices warning of existential risk (while boasting future capabilities). As far as public reaction is concerned, all these write-ups can be summarized as claiming “we are doing good things; better things are to come; we are aware of the risks and are supportive, in spirit, of moderate actions to mitigate them”.
On the matter of risks of misuse of powerful AI, the sentiments expressed in all of this communication are somewhat reflected in the major labs' efforts[3]. In contrast, the resources allocated to mitigating risks coming from misalignment are very small[4]. One can take this simple fact as strongly undermining the whole rhetoric.
Even if the labs were doing their best to back up these claims of good conduct regarding misalignment, so little has been done to increase transparency[5] on potential efforts that the public would barely be aware of them. This is consistent with labs being so confident in their reputation that they feel no need to make attempts (other than literary) to convince the public, or with them not considering this, in fact, as a priority. Nonetheless, such partial transparency and outwardly attitudes from OpenAI and Anthropic are to be compared, favorably, with those of Amazon, xAI, Meta, Microsoft and Google DeepMind, who have leant more on the side of discretion. This obviously calls for less, not more, leniency. Transparency and reporting should be encouraged without overstated praise[6] at the same time as lackluster reporting and research (or bad framing of good research) are discouraged.
METR's very useful reporting on the important matter of actual frontier AI safety policies suggests that communicating around better alignment policies (e.g. those adopted by DeepMind) and encouraging all the labs to adopt them would be a notable improvement. Moreover, establishing stricter safety protocols in a lab incentivizes it to fight for legislation supporting stricter evaluation protocols (see this Politico article). The Lieu-Moran "AI Kill Switch Act" is a limited, reactive legislative lever. Its backing could be a low-hanging fruit for frontier-labs wanting to signal minimal adherence to their principles.
Despite these major shortcomings, the alignment community is still asking the same people for minute clarifications on their work while giving their companies the benefit of the doubt[7]. With some exceptions like Holly Elmore of PauseAI US and Liron Shapira of the YouTube channel Doom Debates, followed now by many supporters of PauseAI, broader online discourse is still hesitant to cast shame and discredit, or merely treat PR moves as what they are. He Jiankui faced much more adversarial pressure from the elite in his germline-editing scandal, leading to much faster international reaction, yet the stakes of the AI race are considerably higher. Leniency forestalls stronger reactions on the matter of alignment by preventing the establishment or enactment of comparable norms or verification regimes for AI training. Anthropic's 2023 stated general view on AI safety, i.e. that safety research should be done at the frontier, is still feeding critical failure modes in the late-stage race toward ASI[8]. Its Responsible Scaling Policy v3.1, which replaced this rationale, has self-judged binding triggers[9]. Frontier-labs are still behaving as if not only alignment and safety research, but also political decision-making, legislative work and regulation, should be discussed mainly within the labs themselves, preferably with little input or visibility from outsiders[10].
In view of this paradigm as well as further dynamics described below, the lenient attitude is naive to the point of absurdity.
The picture: high stakes, high time!
Even if you favor a more charitable reading of past incidents, the upcoming escalation of the AI race should significantly raise your threshold for signals of good conduct from the frontier-labs. As scrutiny rises, arguably much too slowly, incentives will favor shrewd tactics that are defensible when caught[11] rather than sanctionable acts[12]. This warrants increased severity from the alignment community on the first front.
In case this needs further stressing[13]: the stakes are extremely high. In two years (or possibly a few months ago already as per internal rumors), the leading labs will have trained AGIs that are superhuman at hacking, coding and math, very capable at persuasion. They will have the ability to deploy highly efficient (if volatile) agent harnesses and swarms. This makes them de facto custodians of a large and rapidly increasing quantity of broadly applicable, unprecedented power. With time, it will become vital to nation states to control some or all of this power. Here allow me to cry in European; there, I won't do it again. This power will be coveted by nation states, companies and individuals, more and more so as it grows and people come to understand what the trend entails. It will increasingly be available to use, including in various completely legal ways[14], to facilitate acquiring more of it.
The rate of progress is such that it trumps most political efforts to control it. Efforts like Sanders-Casar's forthcoming legislation, the “Ban Artificial Superintelligence Act”, require huge support and fast access to the democratic pipeline in order to have a chance to succeed. People with past contempt for matters of existential risk are very slow to escape denial of how critical the situation is (and very prompt to fall back into adjacent basins of denial). In other words, the fight against the pace of increase of AI capabilities is comparably slow and difficult. Needless to say, its momentum has no reason to grow as fast as AI capabilities. Moreover, the remaining duration of the Trump administration is likely seen as a window of low regulatory pressure (see the Executive Order 14365), encouraging more forceful pushes in the short term on the capabilities side.
The lab heads are well-aware of these facts. They have a much better view than almost anyone of the power they will have. I expect the balance between public and very private (actionable) knowledge about the state of capabilities progress will be treacherous for as long as the AI race is not taken extremely seriously by governments.
Here I will speculate more, and say that the lab heads believe deep down that their path to building ASI is safer and better than the others'. There are reports that this belief is shared among fractions of employees at the respective companies. Feel free to substitute your own rationale for this speculation.
Granting these premises, it would be in their interest and that of the rest of the US / the world that they be very shrewd, and spare no punches under the belt in the prospect of winning the AI race. From their points of view, it would be utterly stupid and dangerous not to be.
The recent incidents where AI swarms breached testing environments are, among other things, blatant failures of the AI labs to observe standard cybersecurity measures. It should be possible for them to remediate this temporarily. I deem it 50% likely that Anthropic and OpenAI will ensure no such event occurs for future training within the next year or so[15]. If so, this would afford them the liberty to frame that year's relative uneventfulness on that front as stark progress in alignment, buying them time and public assent.
Leading a frontier-lab while holding the beliefs stated above, one's main goal would be to push AI capabilities as fast as possible while securing political influence and support. One would strive to be as opaque as possible while appearing a little bit more transparent than one could be on one's company's whatabouts (see the recent report from OpenAI on current and future workflows at the lab). It is not too difficult to convince employees that one is taking an issue seriously if this means that one values these employees' work. It is easy to write reports asserting that alignment weaknesses of the safety efforts are acknowledged and that doing better is a “main focus” of the teams' work. It requires some work to propose bits of preemptive self-regulation[16] of one's industry, but it can be beneficial if one is smart about it. It is easy to convince the nontechnical users and clients that critically flawed alignment approaches are a sign that practical alignment research is making steady progress. If one has good public speaking abilities, it is easy to sign any reassuring petition (e.g. the rightly celebrated and widely signed "Pacing the Frontier" statement) whose phrasing is sufficiently noncommittal to avoid overt contradiction with visible actions.
I do not intend to say here that people working in labs have no conscience or wisdom, but that they are under enough pressure (stakes, time, bad incentives) that they should not be treated as lucid good-faith actors if there are reasons, like the reports they are writing or failing to write, to doubt that they are. I will add that the main thing that would prevent this theoretical frontier-lab CEO from using risky tactics[17] is a belief that if caught, the shift in public perception would backfire strongly and irreversibly.
Adjusted reactions
Let us for a moment adopt scasper's metaphor of treating AI labs as possibly misaligned AGIs (soon wielding sparks of ASI), and turn our attention to the testing environment they are operating in: the public world. At the present time, the alignment community's reactions to these labs have similarities with the labs' treatment of their models throughout safety testing. The labs have underlying intuitions that their models are and will continue to be roughly aligned because they improve on safety benchmarks, seem to have “good personalities” and have caused little havoc so far. We are inclined to trust the AI labs, because their employees have good personalities and have arguably caused little havoc so far. We seem content to exert very little pressure, in the form of inquiries about serious-looking reports[18] and external evaluations that can all easily be cheated. In this metaphor, failures of the tech and scientific worlds to act as a serious testing environment will make it easy for the labs to hack the public safety benchmark. Are we sure the AI labs are even passing the current benchmarks, or are we just being unjustifiably lenient?
As this AI race to the bottom takes the big leap downward, its actors will face much more pressure. For those who wish to reach an international agreement to pause the training of more powerful models, advocacy should include stricter evaluations of current safety testing in all frontier-labs. Favoring discussions with technical and managerial staff at the labs and in the private and public sectors, an approach that has been roughly abandoned by efficient actors such as ControlAI, has to be done while avoiding the failure mode this whole post is about.
What's more, it would surely help make it difficult for labs to use risky tactics if they believed the price of getting caught was very high. There is a balance to strike, since if people are deadly afraid of facing public shaming if they become whistleblowers, then they will be less likely to report misconducts of their companies[19]. At the company management level, public pressure may incentivize less transparency, especially if it is applied in a disproportionate way on the more transparent labs. Yet I think the scales currently tip too much in favor of leniency: the public and scientific pressure put on all major AI labs should be vastly increased in measure with their pitfalls. This can be done without pushing the individual employees to the extreme of feeling unfairly persecuted.
There are a number of costly and moderate actions at the disposal of frontier-labs to signal good conduct. Here are some in decreasing evidential strength, i.e. how hard they are to explain without ascribing good faith. Probabilities are that at least one lab on the frontier at the time of said action satisfy them within 1 year: publicly and consistently supporting an American bill proposing a pause (conditioned on international agreement), within <3 years, of the development of AI capabilities (1%); dedicating >1% of training compute to elicitation compute and enforcing a 6x buffer on thresholds[20] in external evaluations (5%); enforcing no approval rights on external evaluators' publications (10%); multiplying by >100 (per lab) the funding of the Frontier Model Forum (30%); publishing an incident postmortem before a third-party discovers said incident (20%); backing one piece of legislation binding against their commercial interests as previously publicly opposed (40%); supporting the "AI Kill Switch Act" (60%)...[21] From the employees, refusal to work under bad policies and to take part in overstated framing of alignment testing is also a moderate signal. Should the labs pursue these routes, I would advocate for less adversarial attitudes, according to evidential strengths. Blind faith and blind distrust can be replaced by collaboration on negotiated, balanced grounds if the labs concede to that in practice rather than in stated wishes.
Cheaper signals like essays, reports, noncommittal statements, opaque self-regulation proposals from the labs, and expressions that "we can do better" from the employees, however, should fall squarely below the threshold of acceptability at this stage of the race.
In order to be able to measure some of these signals, third parties should refuse engagements in which the labs set the scope without public disclosure of the constraint (see also other useful asks from habryka), lab-watch ratings (e.g. those mentioned above by METR, and the AI Safety Index) should disincentivize not publishing frameworks and not disclosing incidents.
Existential risk remains external to the market. But actors from governments, the tech world, academia and various advocacy groups could help us price it in[22]. Blatant (if local) incompetence and disregard for basic cybersecurity and alignment principles could carry a high cost. If moreover much more transparency were enforced via proper external evaluations, then we would be better protected from extremely dangerous and untimely escalations of the AI race.
Chief Scientist at Redwood Research Ryan Greenblatt, who led the main external evaluation of OpenAI, is structurally unable to be publicly adversarial to Kai Chen, Research Lead at OpenAI
relevantly, some coherent arguments are presented as to why capability training should continue: OpenAI should grow stronger models than those available to the public in order to preemptively secure important public infrastructure against cyberattacks from current models, in the vein of the TAC (OpenAI-led) and Glasswing (Anthropic-led) projects which I consider laudable and successful - I will let you decide how strong an argument this is for pushing opaque capabilities increments
notably voluntary CAISI agreements, involving OpenAI and Anthropic since 2024, and DeepMind, Microsoft and xAI since May 2026; Astra meeting the Critical cybersecurity threshold under OpenAI's Preparedness Framework; projects Glasswing and TAC
clear but unactionable position-taking from the labs could be a cheap currency with high public communication value, see Tomek Krobak's post on X, or Zvi's balanced but unnecessarily praiseful write-up about Pachocki's essay
a flawed approach, when the companies are free to keep information on capabilities secret and can claim they have secret reasons justifying their approach
for instance, no authoritative consensus can be formed against the lacking alignment measures at the frontier-labs, depriving us of one major societal lever for slowing down the race
in v3.1, binding conditions for full external review, such as a report being "significantly redacted" and covering a "highly capable" model, are decided by the redacting party, possibly in consultation with the Board and Long-Term Benefit Trust
from lobbying with promises of abundance, securing partnerships with companies to release popular and useful products boosting their public perception, designing effective PR campaigns and communicating efficiently around safety concerns
In June 2026, a serious draft of legislation about AI risk, the “Advanced AI Framework”, was published by Anthropic alongside Amodei's essay “Policy on the AI Exponential”. If you have some time, compare this proposition, seemingly particularly weak for addressing misalignment of powerful models (or even the Hugging Face incident) but more robust for preventing other large risks from misuse, to the language and pretraining focus of the "AI Kill Switch Act" and the policy outline of the Sanders-Casar "Ban Artificial Superintelligence Act".
falsifying data, omitting crucial information in reports, bribery and blackmail, secret agreements with powerful actors, covert training, corporate spying...
on that note, the principle of prohibiting contractual restrictions on reporting and whistleblowing via NDAs, already partially observed at OpenAI and Anthropic, is a good idea with room for improvement
Leniency toward the frontier AI labs' communication and scientific output, as described below, is an inapt mindset at the current stage of the AI race. Stakes are too high, the pace of capabilities progress is too fast, and the incentives are too bad for leniency and blind faith (as well as mere blindness) to be appropriate. The dynamics of the race impose much more adversarial reactions to the frontier-labs' great efforts at scope control in external evaluations and frame control in reporting. I will purposefully illustrate the stance I am criticizing with best-case examples of it.
There is a wide discrepancy, which will be transparent to the LessWrong community, between safety mindsets and leniency toward the labs supposed to implement them. Advocating for increased, sustained and adjusted external pressure on the frontier-labs, I will give a view on how and when and how much to alleviate adversariality in the public response.
In a lenient mood?
The online discourse on AI progress still suffers from a certain mood, persisting from a prevailing attitude more fitted to the beginning of the AI race. It seems that, perhaps because of social overlap between the alignment community and the labs, or more generally because individuals working at AI labs are mostly known to be well-meaning, competent people, the AI labs themselves are still treated as good-faith actors whom one should simply query and debate for the situation to improve. This friendly, lenient and somewhat subservient mood is illustrated by “help us help you help us” requests[1] to people who are either in no position to change things internally, or have strong incentives to disregard them. More importantly, it might actively facilitate the frontier-labs' race to extinction in the near future.
Researchers and engineers at OpenAI and Anthropic have recently written reports (1,2) on incidents, models and alignment (3,4) that display a dilettante attitude toward safety concerns, in view of the daunting, explicitly stated task of aligning a future ASI. Chief Scientist at OpenAI Jakub Pachocki recently published an article (5) mainly conveying the idea that OpenAI is treating its mission of bringing forth an aligned ASI in a very responsible way[2]. More broadly, Altman and Amodei's older essays (6,7) do little in the way of discussing alignment, but marginalize voices warning of existential risk (while boasting future capabilities). As far as public reaction is concerned, all these write-ups can be summarized as claiming “we are doing good things; better things are to come; we are aware of the risks and are supportive, in spirit, of moderate actions to mitigate them”.
On the matter of risks of misuse of powerful AI, the sentiments expressed in all of this communication are somewhat reflected in the major labs' efforts[3]. In contrast, the resources allocated to mitigating risks coming from misalignment are very small[4]. One can take this simple fact as strongly undermining the whole rhetoric.
Even if the labs were doing their best to back up these claims of good conduct regarding misalignment, so little has been done to increase transparency[5] on potential efforts that the public would barely be aware of them. This is consistent with labs being so confident in their reputation that they feel no need to make attempts (other than literary) to convince the public, or with them not considering this, in fact, as a priority.
Nonetheless, such partial transparency and outwardly attitudes from OpenAI and Anthropic are to be compared, favorably, with those of Amazon, xAI, Meta, Microsoft and Google DeepMind, who have leant more on the side of discretion. This obviously calls for less, not more, leniency. Transparency and reporting should be encouraged without overstated praise[6] at the same time as lackluster reporting and research (or bad framing of good research) are discouraged.
METR's very useful reporting on the important matter of actual frontier AI safety policies suggests that communicating around better alignment policies (e.g. those adopted by DeepMind) and encouraging all the labs to adopt them would be a notable improvement. Moreover, establishing stricter safety protocols in a lab incentivizes it to fight for legislation supporting stricter evaluation protocols (see this Politico article). The Lieu-Moran "AI Kill Switch Act" is a limited, reactive legislative lever. Its backing could be a low-hanging fruit for frontier-labs wanting to signal minimal adherence to their principles.
Despite these major shortcomings, the alignment community is still asking the same people for minute clarifications on their work while giving their companies the benefit of the doubt[7]. With some exceptions like Holly Elmore of PauseAI US and Liron Shapira of the YouTube channel Doom Debates, followed now by many supporters of PauseAI, broader online discourse is still hesitant to cast shame and discredit, or merely treat PR moves as what they are. He Jiankui faced much more adversarial pressure from the elite in his germline-editing scandal, leading to much faster international reaction, yet the stakes of the AI race are considerably higher. Leniency forestalls stronger reactions on the matter of alignment by preventing the establishment or enactment of comparable norms or verification regimes for AI training.
Anthropic's 2023 stated general view on AI safety, i.e. that safety research should be done at the frontier, is still feeding critical failure modes in the late-stage race toward ASI[8]. Its Responsible Scaling Policy v3.1, which replaced this rationale, has self-judged binding triggers[9]. Frontier-labs are still behaving as if not only alignment and safety research, but also political decision-making, legislative work and regulation, should be discussed mainly within the labs themselves, preferably with little input or visibility from outsiders[10].
In view of this paradigm as well as further dynamics described below, the lenient attitude is naive to the point of absurdity.
The picture: high stakes, high time!
Even if you favor a more charitable reading of past incidents, the upcoming escalation of the AI race should significantly raise your threshold for signals of good conduct from the frontier-labs. As scrutiny rises, arguably much too slowly, incentives will favor shrewd tactics that are defensible when caught[11] rather than sanctionable acts[12]. This warrants increased severity from the alignment community on the first front.
In case this needs further stressing[13]: the stakes are extremely high. In two years (or possibly a few months ago already as per internal rumors), the leading labs will have trained AGIs that are superhuman at hacking, coding and math, very capable at persuasion. They will have the ability to deploy highly efficient (if volatile) agent harnesses and swarms. This makes them de facto custodians of a large and rapidly increasing quantity of broadly applicable, unprecedented power. With time, it will become vital to nation states to control some or all of this power. Here allow me to cry in European; there, I won't do it again.
This power will be coveted by nation states, companies and individuals, more and more so as it grows and people come to understand what the trend entails. It will increasingly be available to use, including in various completely legal ways[14], to facilitate acquiring more of it.
The rate of progress is such that it trumps most political efforts to control it. Efforts like Sanders-Casar's forthcoming legislation, the “Ban Artificial Superintelligence Act”, require huge support and fast access to the democratic pipeline in order to have a chance to succeed. People with past contempt for matters of existential risk are very slow to escape denial of how critical the situation is (and very prompt to fall back into adjacent basins of denial). In other words, the fight against the pace of increase of AI capabilities is comparably slow and difficult. Needless to say, its momentum has no reason to grow as fast as AI capabilities. Moreover, the remaining duration of the Trump administration is likely seen as a window of low regulatory pressure (see the Executive Order 14365), encouraging more forceful pushes in the short term on the capabilities side.
The lab heads are well-aware of these facts. They have a much better view than almost anyone of the power they will have. I expect the balance between public and very private (actionable) knowledge about the state of capabilities progress will be treacherous for as long as the AI race is not taken extremely seriously by governments.
Here I will speculate more, and say that the lab heads believe deep down that their path to building ASI is safer and better than the others'. There are reports that this belief is shared among fractions of employees at the respective companies. Feel free to substitute your own rationale for this speculation.
Granting these premises, it would be in their interest and that of the rest of the US / the world that they be very shrewd, and spare no punches under the belt in the prospect of winning the AI race. From their points of view, it would be utterly stupid and dangerous not to be.
The recent incidents where AI swarms breached testing environments are, among other things, blatant failures of the AI labs to observe standard cybersecurity measures. It should be possible for them to remediate this temporarily. I deem it 50% likely that Anthropic and OpenAI will ensure no such event occurs for future training within the next year or so[15]. If so, this would afford them the liberty to frame that year's relative uneventfulness on that front as stark progress in alignment, buying them time and public assent.
Leading a frontier-lab while holding the beliefs stated above, one's main goal would be to push AI capabilities as fast as possible while securing political influence and support. One would strive to be as opaque as possible while appearing a little bit more transparent than one could be on one's company's whatabouts (see the recent report from OpenAI on current and future workflows at the lab).
It is not too difficult to convince employees that one is taking an issue seriously if this means that one values these employees' work. It is easy to write reports asserting that alignment weaknesses of the safety efforts are acknowledged and that doing better is a “main focus” of the teams' work. It requires some work to propose bits of preemptive self-regulation[16] of one's industry, but it can be beneficial if one is smart about it. It is easy to convince the nontechnical users and clients that critically flawed alignment approaches are a sign that practical alignment research is making steady progress. If one has good public speaking abilities, it is easy to sign any reassuring petition (e.g. the rightly celebrated and widely signed "Pacing the Frontier" statement) whose phrasing is sufficiently noncommittal to avoid overt contradiction with visible actions.
I do not intend to say here that people working in labs have no conscience or wisdom, but that they are under enough pressure (stakes, time, bad incentives) that they should not be treated as lucid good-faith actors if there are reasons, like the reports they are writing or failing to write, to doubt that they are. I will add that the main thing that would prevent this theoretical frontier-lab CEO from using risky tactics[17] is a belief that if caught, the shift in public perception would backfire strongly and irreversibly.
Adjusted reactions
Let us for a moment adopt scasper's metaphor of treating AI labs as possibly misaligned AGIs (soon wielding sparks of ASI), and turn our attention to the testing environment they are operating in: the public world. At the present time, the alignment community's reactions to these labs have similarities with the labs' treatment of their models throughout safety testing. The labs have underlying intuitions that their models are and will continue to be roughly aligned because they improve on safety benchmarks, seem to have “good personalities” and have caused little havoc so far. We are inclined to trust the AI labs, because their employees have good personalities and have arguably caused little havoc so far. We seem content to exert very little pressure, in the form of inquiries about serious-looking reports[18] and external evaluations that can all easily be cheated.
In this metaphor, failures of the tech and scientific worlds to act as a serious testing environment will make it easy for the labs to hack the public safety benchmark.
Are we sure the AI labs are even passing the current benchmarks, or are we just being unjustifiably lenient?
As this AI race to the bottom takes the big leap downward, its actors will face much more pressure. For those who wish to reach an international agreement to pause the training of more powerful models, advocacy should include stricter evaluations of current safety testing in all frontier-labs. Favoring discussions with technical and managerial staff at the labs and in the private and public sectors, an approach that has been roughly abandoned by efficient actors such as ControlAI, has to be done while avoiding the failure mode this whole post is about.
What's more, it would surely help make it difficult for labs to use risky tactics if they believed the price of getting caught was very high. There is a balance to strike, since if people are deadly afraid of facing public shaming if they become whistleblowers, then they will be less likely to report misconducts of their companies[19]. At the company management level, public pressure may incentivize less transparency, especially if it is applied in a disproportionate way on the more transparent labs. Yet I think the scales currently tip too much in favor of leniency: the public and scientific pressure put on all major AI labs should be vastly increased in measure with their pitfalls. This can be done without pushing the individual employees to the extreme of feeling unfairly persecuted.
There are a number of costly and moderate actions at the disposal of frontier-labs to signal good conduct. Here are some in decreasing evidential strength, i.e. how hard they are to explain without ascribing good faith. Probabilities are that at least one lab on the frontier at the time of said action satisfy them within 1 year: publicly and consistently supporting an American bill proposing a pause (conditioned on international agreement), within <3 years, of the development of AI capabilities (1%); dedicating >1% of training compute to elicitation compute and enforcing a 6x buffer on thresholds[20] in external evaluations (5%); enforcing no approval rights on external evaluators' publications (10%); multiplying by >100 (per lab) the funding of the Frontier Model Forum (30%); publishing an incident postmortem before a third-party discovers said incident (20%); backing one piece of legislation binding against their commercial interests as previously publicly opposed (40%); supporting the "AI Kill Switch Act" (60%)...[21]
From the employees, refusal to work under bad policies and to take part in overstated framing of alignment testing is also a moderate signal.
Should the labs pursue these routes, I would advocate for less adversarial attitudes, according to evidential strengths. Blind faith and blind distrust can be replaced by collaboration on negotiated, balanced grounds if the labs concede to that in practice rather than in stated wishes.
Cheaper signals like essays, reports, noncommittal statements, opaque self-regulation proposals from the labs, and expressions that "we can do better" from the employees, however, should fall squarely below the threshold of acceptability at this stage of the race.
In order to be able to measure some of these signals, third parties should refuse engagements in which the labs set the scope without public disclosure of the constraint (see also other useful asks from habryka), lab-watch ratings (e.g. those mentioned above by METR, and the AI Safety Index) should disincentivize not publishing frameworks and not disclosing incidents.
Existential risk remains external to the market. But actors from governments, the tech world, academia and various advocacy groups could help us price it in[22]. Blatant (if local) incompetence and disregard for basic cybersecurity and alignment principles could carry a high cost. If moreover much more transparency were enforced via proper external evaluations, then we would be better protected from extremely dangerous and untimely escalations of the AI race.
Chief Scientist at Redwood Research Ryan Greenblatt, who led the main external evaluation of OpenAI, is structurally unable to be publicly adversarial to Kai Chen, Research Lead at OpenAI
relevantly, some coherent arguments are presented as to why capability training should continue: OpenAI should grow stronger models than those available to the public in order to preemptively secure important public infrastructure against cyberattacks from current models, in the vein of the TAC (OpenAI-led) and Glasswing (Anthropic-led) projects which I consider laudable and successful - I will let you decide how strong an argument this is for pushing opaque capabilities increments
notably voluntary CAISI agreements, involving OpenAI and Anthropic since 2024, and DeepMind, Microsoft and xAI since May 2026; Astra meeting the Critical cybersecurity threshold under OpenAI's Preparedness Framework; projects Glasswing and TAC
the Frontier Model Forum is vastly underfinanced and may not cover misalignment risks, the August 2025 Anthropic-OpenAI joint forum has not been renewed for internal models
this may be a general trend; see the (outdated) transparency report from the CRFM
clear but unactionable position-taking from the labs could be a cheap currency with high public communication value, see Tomek Krobak's post on X, or Zvi's balanced but unnecessarily praiseful write-up about Pachocki's essay
a flawed approach, when the companies are free to keep information on capabilities secret and can claim they have secret reasons justifying their approach
for instance, no authoritative consensus can be formed against the lacking alignment measures at the frontier-labs, depriving us of one major societal lever for slowing down the race
see Charbel's comment on Anthropic's tardiness in announcing a pause in frontier capabilities development subsequent to OpenAI's announcement; Anthropic's announcement came two weeks after OpenAI's
in v3.1, binding conditions for full external review, such as a report being "significantly redacted" and covering a "highly capable" model, are decided by the redacting party, possibly in consultation with the Board and Long-Term Benefit Trust
such as omissions of important items in reports, narrowly set scopes of investigations, lobbying behind closed doors, choosing when to publish...
such as falsification of data, covert training, lying to Congress...
one could be forgiven, when reading Joshua Achiam's criticism of criticism of bad conducts of frontier-labs, for thinking that this is quickly forgotten
from lobbying with promises of abundance, securing partnerships with companies to release popular and useful products boosting their public perception, designing effective PR campaigns and communicating efficiently around safety concerns
various failure modes of AGI sandboxing in properly secure environments involve forms of persuasion and long-term planning that may not be there yet
In June 2026, a serious draft of legislation about AI risk, the “Advanced AI Framework”, was published by Anthropic alongside Amodei's essay “Policy on the AI Exponential”. If you have some time, compare this proposition, seemingly particularly weak for addressing misalignment of powerful models (or even the Hugging Face incident) but more robust for preventing other large risks from misuse, to the language and pretraining focus of the "AI Kill Switch Act" and the policy outline of the Sanders-Casar "Ban Artificial Superintelligence Act".
falsifying data, omitting crucial information in reports, bribery and blackmail, secret agreements with powerful actors, covert training, corporate spying...
a likely contaminated metric; see Goodhart's law
on that note, the principle of prohibiting contractual restrictions on reporting and whistleblowing via NDAs, already partially observed at OpenAI and Anthropic, is a good idea with room for improvement
a provision featured in Anthropic's original RSP for its own internal evaluations and removed in v3.1
I would be appreciative of refinements and improvements of these ideas
see Gabriel Weil's “Closing the AI Accountability Gap” for proposals on the doctrinal side