When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected.
We argue there should be a dedicated effort to
Replicate alignment experiments from frontier labs.
Scrutinize the experiments by stress-testing the methodology.
Open-source replications to encourage external researchers to validate our work, build on the experiment, and further audit the lab’s methods.
The case to replicate safety research from labs
CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why, Beneficial RL)[1].
There’s good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter provider or LoRA alpha. Some researchers have told us directly that they think there may exist some arbitrary methodological choices in their own research that could plausibly change results. A replication is a basic test for this. An independent team will inevitably make many choices differently and is unlikely to share the same bug, lucky seed, or prompting failure.
If an experiment doesn't replicate, this is strong evidence that the finding won’t generalize well in the future. An alternative explanation for a failed replication is that the replicators themselves did something wrong. But if a skilled team[2] works for an extended period of time, asks the authors for help, and still can't reproduce the result, some of the blame falls on the lab.[3] It tells us the lab is not releasing safety research that is easily externally verifiable, which is a kind of result within itself.
Model cards may be an especially high-leverage subject for replication. When a lab claims that they are releasing their “most aligned model ever,” we could replicate and stress-test their evidence.[4] These replications could also be applied to frontier models with less comprehensive model cards to effectively fill in what the lab left out.
Replications are not shiny, but that’s precisely what makes them counterfactually useful
The case for replications is somewhat self-evident. It is not an exaggeration to say that most AI safety researchers would agree that someone should do this work, and yet, little effort is paid in its direction.
Across scientific disciplines, replication is systematically neglected because the work is unsexy and novel research is exciting. It may also be unattractive because it feels less likely to result in publication or land you a job offer. Regardless of the explanation, the work is not currently happening at the scale it should, and there isn’t strong reason to believe this will change without targeted effort.
The case to stress test
Labs present safety research as progress toward aligning frontier models and, perhaps eventually, superintelligence. If we take this at face value, the claims deserve auditing effort proportional to those stakes.
In many cases, stress testing looks like replicating the experiment across a wide variety of settings to identify effect size, trends, and alternative hypotheses. This could include obvious things like hyperparameter sweeps, testing different model families, or running additional evals. In other cases, it may be necessary to perform robustness checks that are more specific to the experiment at hand:
Looking for cases where the baseline was weak or misconfigured
Testing if a causal claim is not instead downstream of a non-obvious correlation
Reading through transcripts to see if the descriptions from the paper qualitatively match model behavior
One limitation of stress-testing is that it can't tell you if a result is relevant to alignment of superintelligence or if the paper communicated in a misleading fashion. Raising issues related to high-level methodology or communication is often better suited to position-paper-like blog posts, which we also think the field could use more of. However, stress-testing many papers could reveal patterns that bear on these questions: if clusters of research don’t hold up to scrutiny, this could help build consensus on alignment being hard or certain techniques not being particularly promising.
In most cases, we would probably do both empirical stress testing along with some limited non-experimental analysis. Even when stress-testing claims, the goal is to be balanced, and not overemphasize flaws (see more of our thoughts here).
The case to open source
It is currently difficult for third-party alignment researchers to interface with lab-released work. Before building on an experiment, you have to replicate it, often from a paper which may include few methodological details. Open sourcing replications would allow external AI safety researchers to participate: they could use the code to further stress-test a lab’s research, build on it, or understand it more deeply. The hope is that this could make the AI safety community more coherent, self-correcting, and likely to build consensus on important facts.
Equally important, open sourcing replications allows other external researchers to check the replication’s own claims and methods.
Replicating frontier lab work is difficult but tractable
Replicating at a frontier scale is time-consuming, requires some taste, and can be expensive. But these bottlenecks can be overcome, and we think it's worth the effort.
First, stress testing any method is messy. We expect claims from labs to replicate in some ways but be flawed in others. Trying to disentangle what replicates cleanly, what generalizes poorly, and what the high-level story is will require good researchers and time. But intentional recruiting and competitive salaries could solve this. Likewise, compute for replicating at a frontier scale can be expensive, but fundraising is tractable.
What cannot be unbottlenecked is access to frontier model weights. That said, open-weight models trail the frontier by only ~4 months, so if an experiment doesn’t replicate on them, this still represents a relevant datapoint.
Conclusion
History would tell us we are playing with fire. Psychology and preclinical cancer biology have indexed on misleading or false claims that sensible meta-science could have uncovered. Machine learning has its own versions: a lot of deep RL results from the mid-2010s turned out to be difficult to reproduce and adversarial machine learning results often fail to hold up to scrutiny. Similar structural issues may exist in AI safety today.
Meta-science at a frontier scale is unglamorous and unincentivized. But it’s also the cheapest insurance we have against repeating history.
Thank you to Anastasia Wei, Coby Kassner, and Rhea Kanuparthi for helpful feedback on this post.
One caveat is that it’s possible results don’t replicate because the models being tested are sufficiently different. For example, if OpenAI’s goblins results didn’t replicate on other models, that of course would be unsurprising.
vital work. so much folk lore is presented as established fact. an actual legitimate use for open weight models, and the Allen Institute's Olmo is true open source
When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected.
We argue there should be a dedicated effort to
The case to replicate safety research from labs
CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why, Beneficial RL)[1].
There’s good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter provider or LoRA alpha. Some researchers have told us directly that they think there may exist some arbitrary methodological choices in their own research that could plausibly change results. A replication is a basic test for this. An independent team will inevitably make many choices differently and is unlikely to share the same bug, lucky seed, or prompting failure.
If an experiment doesn't replicate, this is strong evidence that the finding won’t generalize well in the future. An alternative explanation for a failed replication is that the replicators themselves did something wrong. But if a skilled team[2] works for an extended period of time, asks the authors for help, and still can't reproduce the result, some of the blame falls on the lab.[3] It tells us the lab is not releasing safety research that is easily externally verifiable, which is a kind of result within itself.
Model cards may be an especially high-leverage subject for replication. When a lab claims that they are releasing their “most aligned model ever,” we could replicate and stress-test their evidence.[4] These replications could also be applied to frontier models with less comprehensive model cards to effectively fill in what the lab left out.
Replications are not shiny, but that’s precisely what makes them counterfactually useful
The case for replications is somewhat self-evident. It is not an exaggeration to say that most AI safety researchers would agree that someone should do this work, and yet, little effort is paid in its direction.
Across scientific disciplines, replication is systematically neglected because the work is unsexy and novel research is exciting. It may also be unattractive because it feels less likely to result in publication or land you a job offer. Regardless of the explanation, the work is not currently happening at the scale it should, and there isn’t strong reason to believe this will change without targeted effort.
The case to stress test
Labs present safety research as progress toward aligning frontier models and, perhaps eventually, superintelligence. If we take this at face value, the claims deserve auditing effort proportional to those stakes.
In many cases, stress testing looks like replicating the experiment across a wide variety of settings to identify effect size, trends, and alternative hypotheses. This could include obvious things like hyperparameter sweeps, testing different model families, or running additional evals. In other cases, it may be necessary to perform robustness checks that are more specific to the experiment at hand:
One limitation of stress-testing is that it can't tell you if a result is relevant to alignment of superintelligence or if the paper communicated in a misleading fashion. Raising issues related to high-level methodology or communication is often better suited to position-paper-like blog posts, which we also think the field could use more of. However, stress-testing many papers could reveal patterns that bear on these questions: if clusters of research don’t hold up to scrutiny, this could help build consensus on alignment being hard or certain techniques not being particularly promising.
In most cases, we would probably do both empirical stress testing along with some limited non-experimental analysis. Even when stress-testing claims, the goal is to be balanced, and not overemphasize flaws (see more of our thoughts here).
The case to open source
It is currently difficult for third-party alignment researchers to interface with lab-released work. Before building on an experiment, you have to replicate it, often from a paper which may include few methodological details. Open sourcing replications would allow external AI safety researchers to participate: they could use the code to further stress-test a lab’s research, build on it, or understand it more deeply. The hope is that this could make the AI safety community more coherent, self-correcting, and likely to build consensus on important facts.
Equally important, open sourcing replications allows other external researchers to check the replication’s own claims and methods.
Replicating frontier lab work is difficult but tractable
Replicating at a frontier scale is time-consuming, requires some taste, and can be expensive. But these bottlenecks can be overcome, and we think it's worth the effort.
First, stress testing any method is messy. We expect claims from labs to replicate in some ways but be flawed in others. Trying to disentangle what replicates cleanly, what generalizes poorly, and what the high-level story is will require good researchers and time. But intentional recruiting and competitive salaries could solve this. Likewise, compute for replicating at a frontier scale can be expensive, but fundraising is tractable.
What cannot be unbottlenecked is access to frontier model weights. That said, open-weight models trail the frontier by only ~4 months, so if an experiment doesn’t replicate on them, this still represents a relevant datapoint.
Conclusion
History would tell us we are playing with fire. Psychology and preclinical cancer biology have indexed on misleading or false claims that sensible meta-science could have uncovered. Machine learning has its own versions: a lot of deep RL results from the mid-2010s turned out to be difficult to reproduce and adversarial machine learning results often fail to hold up to scrutiny. Similar structural issues may exist in AI safety today.
Meta-science at a frontier scale is unglamorous and unincentivized. But it’s also the cheapest insurance we have against repeating history.
Thank you to Anastasia Wei, Coby Kassner, and Rhea Kanuparthi for helpful feedback on this post.
We are currently working on replicating both of these.
The bar is probably at or above the level of a MATS mentee or Anthropic Fellow.
One caveat is that it’s possible results don’t replicate because the models being tested are sufficiently different. For example, if OpenAI’s goblins results didn’t replicate on other models, that of course would be unsurprising.
Because some experiments require visibility into the model's CoT, there are inherent limitations to how comprehensive this could be.