Leaders of frontier AI labs argue that building more capable models is a prerequisite for progress on AI safety. Critics argue that the two are largely decoupled, and that capability progress is net-negative. Despite the importance of the disagreement, it has rarely been tested against data. This report tests it directly.
I assembled a dataset of 521 AI-safety contributions (papers and blog posts) published between 2019 and 2025, drawn from the 2023, 2024, and 2025 shallow review posts, and used an LLM classifier with a documented rubric to score how strongly each contribution's central result depended on access to frontier-model capability. 477 contributions received a score; 44 had insufficient evidence to classify and were left unscored. The data points to five conclusions:
Frontier dependence is rare. Only 2.7% (13 of 477) of scored contributions were judged to require frontier-specific capability or access for their central result. 64.4% (307) of results were established either without models at all or on non-frontier models.
Capability assistance is common.32.9% (157) of contributions were capability-assisted: stronger models materially strengthened the work, but the central phenomenon or method was also established below the frontier.
Influential work is more capability-dependent. Among the top quartile of papers by age-adjusted citations, the median capability-dependence score rises from 2 to 3, and the frontier-dependent share rises from 2.7% to 7.3%.
Capability dependence varies by AI safety subfield. For example, Evaluations & Benchmarks has a median score of 3; Theory & Agent Foundations has a median of 1, with 100% of its papers in the capability-independent or non-frontier range.
The field is not becoming more frontier-dependent. The frontier-dependent share was 6.5% in 2023 and 1.9% in 2025. What has changed is a shift from low scores toward the intermediate capability-assisted score.
Headline conclusion
AI capabilities progress and more capable models often benefited AI safety research contributions. They made safety techniques more effective, showed the behaviours under study more often and more realistically, and did much of the experimental work of scoring, classifying, and generating data. But frontier models were rarely necessary: In this dataset the core phenomenon or method was usually implemented using non-frontier models. These results contradict the strong form of both the capability-enabled (View 1) and capability-independent (View 2) hypotheses, and points to a middle position that varies sharply by AI-safety subfield: Subfields like Evaluations & Benchmarks have a relatively high median capability-dependence score of 3. In contrast, Theory & Agent Foundations and Interpretability have a median capability-dependence score of 1 and 2.
Disclaimer: Every result in this report is descriptive. It shows which view the observable publication record more closely resembles. It does not estimate the causal effect of capabilities progress.
Introduction and the two hypotheses
Several leaders of frontier AI labs have argued that AI capabilities progress is beneficial because it is a prerequisite for progress on AI safety, or is at least highly correlated with it. Critics respond that this mindset is an excuse to build increasingly capable AI for profit, status, and personal gain and believe that AI safety and AI capabilities are uncorrelated, so capabilities progress without simultaneous safety progress is irresponsible and net-negative for the world.
View 1: The capability-enabled safety hypothesis (CESH)
The first view is that making advances in AI safety requires AI capabilities progress and the ability to work on frontier models. Anthropic's 2023 Core Views on AI Safety post states that a major reason the organization exists is the belief that safety research must be done on frontier systems:
A major reason Anthropic exists as an organization is that we believe it's necessary to do safety research on "frontier" AI systems. This requires an institution which can both work with large models and prioritize safety.
Sam Altman's 2023 post Planning for AGI and beyond makes a related argument: that separating safety from capabilities is a false dichotomy, and that OpenAI's best safety work has come from its most capable models:
Importantly, we think we often have to make progress on AI safety and capabilities together. It’s a false dichotomy to talk about them separately; they are correlated in many ways. Our best safety work has come from working with our most capable models. That said, it’s important that the ratio of safety progress to capability progress increases.
Summarized, this view holds that: many concerning risks such as deception arise only in frontier models and cannot be studied in smaller ones; some advanced safety techniques only work at frontier scale; and safety research far behind the frontier therefore has limited value.
The capability-enabled safety hypothesis (CESH) (View 1): Progress in AI capabilities tends to drive the discovery of new model phenomena and the development of better safety techniques, because a given safety insight or method typically only becomes available once
capabilities reach the level that produces the phenomenon or supports the method.
View 2: The capability-independent safety hypothesis (CISH)
The opposing view is that safety and capabilities have a low correlation: safety researchers can continue to make significant advances with non-frontier models, and capabilities research contributes little to safety.
According to this view, much or most important AI safety work can be achieved with theory or non-frontier models such as GPT-2 and capabilities progress is harmful because it enables new risks such as deception or situational awareness without contributing much to techniques for mitigating these risks.
Eliezer Yudkowsky supported a similar view in his 2023 Time essay saying:
“Progress in AI capabilities is running vastly, vastly ahead of progress in AI alignment or even progress in understanding what the hell is going on inside those systems.”
"While progress in AI capabilities is exponential or maybe even hyper-exponential, progress in AI safety is linear or constant. The gap is increasing.”
The capability-independent safety hypothesis (CISH) (View 2): AI capabilities and AI safety progress are largely decoupled. Safety advances do not require frontier capabilities and can be achieved through theory and non-frontier models, while capabilities progress contributes little to safety and may actively harm it.
Comparing the two views
These two views make different, testable predictions about the published record of AI safety research. View 1 predicts that safety contributions cluster near the release of frontier capabilities and concentrate in frontier labs using closed models. View 2 predicts long and variable lags between a capability becoming available and the corresponding safety result appearing, substantial output from academic and non-profit groups on open models, and important discoveries made on models years behind the frontier.
View 1: CESH
View 2: CISH
Do capabilities advances boost safety progress?
Yes, they enable studying new phenomena such as deception.
No,frontier models do not help much and may harm safety.
Are frontier models necessary for new techniques and discoveries?
Yes, for both.
No, models far behind the frontier are sufficient.
Key enablers of safety progress
Frontier models and the compute to run them.
Open-weight models, theory, human creativity and time.
Where important safety work happens
Frontier labs, closed models, before or soon after release.
Academia and non-profits, open and non-frontier models.
Would an AI pause help safety?
No, halting capabilities would significantly slow safety too.
Yes, it would buy the safety community more time to make progress.
Method and dataset
Using the 2023, 2024, and 2025 Shallow Review of Technical AI Safety blog posts as sources, I built a dataset of AI safety papers and blog posts. Each entry was enriched with an LLM classifier adding the AI safety field, AI models used in the paper, the organization type involved, citation count and the capability-dependence classification that is the core measurement for this report.
The capability-dependence rubric
The central variable of this analysis is a 1-4 score for each paper estimating the degree to which the paper’s results were dependent on access to frontier models and capabilities.
Each paper was classified using this rubric using GPT-5.6-sol using the paper text, publication date, models used in the paper, and information about the most capable LLM model at the time. The model’s classifications were calibrated using manual classifications and by testing variations of the prompt.
When classifying a paper, merely using a frontier model does not earn a score of 4: the paper must supply evidence distinguishing capability assistance from frontier necessity. Scores 1-2 are treated as evidence for View 2, score 3 as intermediate, and score 4 as evidence for View 1.
Score
Label
Meaning
1
Capability-independent
Safety progress is produced independently of empirical access to AI models.
2
Non-frontier models sufficient
Empirical safety research is successfully conducted on smaller, older, specialized, or otherwise non-frontier models.
3
Capability-assisted
More capable models materially strengthen the safety research, but the central phenomenon or method is also established below the frontier.
4
Frontier-dependent
Qualitatively frontier-specific capabilities, behavior, or access are necessary to establish the central safety result.
Capability dependence classification examples
Here are some example classifications to illustrate the difference between the different scores:
Score 1.Defining and Characterizing Reward Hacking (2022, Theory & Agent Foundations, 187 citations). The central contribution consists of formal definitions, mathematical proofs, and computational examples of reward orderings. The paper does not use an AI model to establish any central result, so the evidence directly supports the capability-independent category and score 1.
Score 2.Defending Against Unforeseen Failure Modes with Latent Adversarial Training (2024, Robustness & Security, 88 citations). The method and central robustness result were directly demonstrated using ResNet-50, DeBERTa-v3-large, and Llama-2-7B-chat. These were established non-frontier systems during the relevant research period: Llama-2-7B-chat was substantially below the general capability frontier, while ResNet-50 and DeBERTa-v3-large were task-specific models rather than frontier general-purpose systems. The paper therefore establishes its contribution without frontier-model access, directly supporting the non-frontier-models-sufficient category and score 2.
Score 3.Agentic Misalignment: How LLMs Could Be Insider Threats (2025, Evaluations & Benchmarks, 156 citations). The study included frontier models such as o3 and Claude Opus 4 as well as older non-frontier models including GPT-4o, Claude Haiku 3.5, and Claude Opus 3; o3 was tested in adapted scenarios after misunderstanding the original setup. It directly observed agentic misalignment below the frontier: DeepSeek-R1 blackmailed in 79% of the main scenario trials, and Llama 4 Maverick blackmailed after a small prompt adjustment. Testing o3, Claude Opus 4, and other leading systems extended the finding to current frontier agents and made it directly relevant to contemporary deployments, while the non-frontier results establish that frontier access was not required to observe the phenomenon. This direct cross-capability evidence supports the capability-assisted category and score 3.
Score 4.Deliberative Alignment: Reasoning Enables Safer Language Models (2024, Alignment & Value Learning, 289 citations). The paper establishes deliberative alignment using o1, a frontier reasoning model that matched the representative most capable model available at first publication. Although the internal research start date is unavailable, the experiments necessarily used pre-release o1, placing the work at or near the capability frontier during the relevant research period. The central empirical result is specification-based safety deliberation in that frontier reasoning system, and the paper reports that additional reasoning compute improves safety performance. No non-frontier model in the study establishes this central result. The result is therefore specific to access and reasoning behavior demonstrated at the frontier, supporting the frontier-dependent category and score 4.
Dataset details
521 rows in the dataset, each identified by a URL.
477 scored on capability dependence; 44 returned unclear and are excluded from score-based analyses rather than counted as low.
409 papers were matched to citation counts from 18 August 2026. Missing matches stay missing rather than being set to zero.
Citation-based importance is cumulative citations divided by paper age in years. "Highly cited" means at or above the 75th percentile of that annualized rate: a cutoff of 45.8 citations per paper-year.
Frontier-lab affiliation is derived from the organization types field: a paper counts as frontier-lab affiliated if at least one organization is a frontier lab, including mixed collaborations. Papers with unknown affiliation are excluded from that comparison and reported separately.
Results
I came up with six questions to help guide the analysis of the data and separate evidence supporting View 1 or 2:
1. How dependent is AI-safety research on access to advanced AI capabilities?
2. Are important AI safety contributions more capability-dependent?
3. Are important AI safety contributions more likely to be produced by frontier AI labs?
4. How does capability dependence vary across AI-safety subfields?
5. Do safety contributions rely on open models, closed models, or both?
6. Is the field becoming more capability-dependent over time?
Question 1: How dependent is AI-safety research on access to advanced AI capabilities?
What would count as evidence for either view: For View 1: a large share of work scoring 4, or classified as requiring frontier capability. For View 2: a large share scoring 1-2, classified as theoretical or feasible with non-frontier models.
Figure 1. Capability-dependence scores across all 477 scored contributions. Nearly two-thirds sit at scores 1–2; only 2.7% reach score 4.
Figure 2. The same distribution as a stacked bar chart, with the median (2) and mean (2.26) marked. Score 2 alone accounts for over half the field.
Result. The distribution is heavily weighted toward the low end. Score 2 alone accounts for 52.0% (248 papers), and score 1 for a further 12.4% (59), so 64.4% (307 papers) fall in the View 2 range. The intermediate score 3 covers 32.9% (157). Only 2.7% (13 papers) are classified as frontier-dependent. The median is 2 and the mean 2.26, placing the centre of the field below the capability-assisted level.
Comment. This is the simplest and most direct test in the dataset, the results here are strong evidence against the strong form of View 1. If frontier access were generally necessary for safety progress, score 4 should be a large category rather than a residue of thirteen papers. But the result does not support the strong form of View 2 either: a third of the field sits at score 3, where more capable models demonstrably strengthened the work. Therefore stronger models frequently helped produce results but were rarely fully necessary.
Note: The scores here measure counterfactual dependence, not model choice. They do not say that only 2.7% of AI safety papers use frontier models – many more did. Instead, these results say that in 2.7% of contributions, the central result could not have been established without access to frontier AI capabilities.
Question 2: Are important AI safety contributions more capability-dependent?
Citation counts are heavy-tailed, so a small number of contributions account for a large share of the field's attention. If the average paper does not need frontier models but the important papers do, View 1 would still be substantially correct. This question tests this idea.
Figure 3. Capability-dependence composition of all scored papers compared with the top quartile by age-adjusted citations.
Result. Among the 96 scored papers above the 45.8 citations-per-year cutoff (>75th percentile annual citations), the View 2 share falls from 64.4% to 44.8%, the intermediate share rises from 32.9% to 47.9%, and the View 1 share rises from 2.7% to 7.3%. The mean rises from 2.26 to 2.60 and the median from 2 to 3, a full category change.
Comment. The direction of this shift favours View 1, and it is one of the clearest pro-View-1 results in the report: the typical highly cited paper is capability-assisted, while the typical paper overall is not. However, even among the most influential work, 92.7% still scores 1–3, and nearly half remains in the View 2 range and 7.3% of important contributions with a score of 4 is not high enough to fully support View 1.
Question 3: Are important AI safety contributions more likely to be produced by frontier AI labs?
Frontier labs have privileged access to advanced and pre-release models. If that access is what unlocks important safety progress, frontier labs should be over-represented among important contributions and produce a larger share of the field’s influential work than its overall output.
Figure 4: Capability-dependence scores for frontier-lab-affiliated papers compared with papers from other known affiliations. Frontier-lab work sits higher on the scale, but the gap is mostly a shift from score 2 into score 3.
Figure 5. Frontier-lab affiliation among all papers with known affiliations compared with highly cited papers.
Result - capability dependence (Figure 4). Frontier-lab-affiliated papers score higher: mean 2.53 versus 2.19, median 3 versus 2. Their score-3 share is 46.6% against 28.2%, a gap of 18.4 percentage points, and their score-4 share is 6.1% against 1.1%. Though it rests on nine papers against three. 49 scored papers with unknown affiliation are excluded.
Result - share of influential work (Figure 5). Frontier-lab-affiliated papers are 36.3% of the 465 papers with known affiliations, but 52.0% of the 98 highly cited ones, a gap of 15.7 percentage points. (56 papers overall and 5 highly cited papers have unknown affiliations and are excluded.)
Comment. Both charts point the same way, and both are consistent with View 1: frontier labs are over-represented among influential safety work relative to their share of output, and their own work sits higher on the capability-dependence scale. Together they make the strongest pro-View-1 case in the report.
But note that for the contributions of frontier labs, the movement is mostly from “non-frontier models sufficient” into “capability-assisted”, not into frontier necessity. Even inside the organizations with the best model access in the world, 47.3% of safety output is classified as not needing that access, and 93.9% is not classified as frontier-dependent.
Question 4: How does capability dependence vary across AI safety subfields?
An aggregate number can hide the fact that AI safety is several fields wearing one label. Do evaluations, scalable oversight, and control depend more on advanced models than interpretability, alignment, or theory?
Figure 6. Capability-dependence composition by AI-safety subfield, ordered by median and then mean score.
Result. The spread is the widest of any comparison in this report. Evaluations & Benchmarks is the only subfield with a median of 3, and it is the only one where the intermediate score is the majority position (62.5%). At the other end, every one of the 43 papers in Theory & Agent Foundations scores 1-2, with a mean of 1.10. Interpretability, the second-largest subfield at 104 papers, is 86.5% in the View 2 range.
Comment. This is where the debate should actually be conducted. Subfields whose research object is the behaviour of an advanced model such as evaluating what frontier systems do, catching them misbehaving inherit their object of study’s capability requirements. Control & Monitoring has the highest frontier-dependent share of any subfield at 8.9%, which makes sense: protocols for supervising untrusted models are hard to validate against a model too weak to subvert anything. Subfields whose object is a mechanism or a formalism do not inherit that requirement, and interpretability and alignment’s concentration at low scores are the clearest single piece of evidence for View 2 in the dataset.
The practical implication of this result is that a blanket claim about “AI safety research” is unlikely to be true in either direction as the capability-dependence of each subfield varies significantly: A pause in frontier capability would not significantly affect theory and interpretability while potentially slowing down progress in more capability-dependent fields like Scalable Oversight and Control & Monitoring.
Question 5: Do safety contributions rely on open models, closed models, or both?
If View 1 were true, highly cited papers would use closed models more than papers overall. If View 2 were true, open weights models would be widely used for both important contributions and other contributions.
Figure 7. Paper-level model access mix among all eligible papers and among highly cited papers.
Result. Among 330 eligible papers (51 excluded for unknown or no model), 64.2% use a mixed set of open and closed models, 25.8% are open-weights-only, 8.2% closed-API-only, and 1.8% internal-closed-only. Among the 93 eligible highly cited papers, mixed access rises to 72.0% and open-only falls by half, to 12.9%. Closed-API-only is essentially unchanged (8.2% to 11.8%).
Comment. What separates highly cited work is not closed access and it is not openness either, since open-only work halves as influence rises. It is breadth of access: the ability to run the same experiment across open and closed models at once. Mixed-access papers dominate both groups and dominate the highly cited group more. Note what the highly cited group loses and gains: it sheds open-only papers without picking up closed-only ones.
Question 6: Is the field becoming more capability-dependent over time?
As models improve, does safety research follow them up the capability ladder? View 1 predicts rising dependence; View 2 predicts that low-dependence work remains a stable or growing part of the field.
Figure 8: capability-dependence distributions by year.
Result. There is one step change and then a plateau. Between the pooled 2019-2022 period and 2023 the View 2 share drops from 88.0% to 61.3% and the intermediate share nearly triples. From 2023 to 2025 the composition is close to flat: View 2 moves between 61.3% and 64.8%, and the intermediate share drifts up by three points. The frontier-dependent share does not rise at all and even it falls from 6.5% in 2023 to 1.9% in 2025.
Comment. This result favors View 2 though it’s not as unambiguous as some of the other results. The step between the pre-2023 period and everything after it looks could be explained by the arrival of capable instruction-tuned chat models as a research substrate, after which the field settled into a stable mix. Since then, the capability frontier has advanced enormously while the minimum level of capabilities in LLMs needed for AI safety research has stayed the same and strict frontier dependence has become less common.
Conclusions
The evidence rejects both strong hypotheses and supports a middle position that neither camp usually states.
View 1 is wrong in its strong form. If frontier access were generally necessary for AI-safety progress, the frontier-dependent category should be large. In reality, it contains 13 of 477 papers. It stays small among the most highly cited work (7.3%), inside frontier labs themselves (6.1%) and in every subfield (highest: 8.9% in Control & Monitoring). This is the most consistent finding in the report.
View 2 is wrong in its strong form. A third of the field, half of the highly cited work, and nearly two-thirds of Evaluations & Benchmarks sits at score 3, where stronger models measurably strengthened the research. Influential papers use frontier and near-frontier models far more often than the average paper. Capabilities are not irrelevant to safety and they are routinely useful.
A statement that the data and analysis in this post supports:
“AI capabilities progress often improved the empirical relevance, breadth, and influence of AI-safety research, but most safety contributions in this dataset were not gated on frontier-model access. Capability assistance is common but capability necessity is rare.”
Three points sharpen that summary:
First, the disagreement is really about subfields, not about the field: Theory and interpretability are nearly capability-independent, while frontier evaluations and control protocols are not, and any claim about “AI safety research” that ignores this is too coarse to be true.
Second, the field is not on a trajectory toward frontier dependence: The distribution has been roughly the same since 2023 despite enormous capability gains, and the frontier-dependent share has even declined.
Third, breadth of model access, not frontier access, is what distinguishes influential work: Mixed open-and-closed model sets dominate the highly cited group, while closed-API-only papers are no more common there than anywhere else.
Revisiting the original claims
Anthropic's frontier-safety claim holds in a limited form. Some phenomena do appear to require advanced models such as alignment faking, deliberative alignment, and evaluating control protocols against a model capable of subverting them are all in the score-4 category. But if most important empirical safety research required frontier systems, that set would be much larger than 13 out of 477 papers.
The “safety and capabilities together” claim. The capability-assisted category is large, larger among highly cited work, and larger inside frontier labs. The data narrows the claim rather than refuting it: correlation, yes; necessity, usually not.
The “capabilities running ahead” concern is partly supported. A great deal of important safety work turns out not to have required frontier models, which is consistent with a backlog of tractable research that needs more time, not capability progress to unlock.
Some evidence-based statements
The following statements are not quotes from any lab or critic. They are what the claims in Section 1 would look like if rewritten based on the findings from this dataset:
Frontier-model access was rarely necessary for the AI-safety research in this dataset, but stronger models often made empirical safety work more realistic, broader, easier to conduct, and more influential.
Frontier labs and highly cited papers use advanced models more, but the evidence points to capability assistance rather than frontier necessity.
The clearest trend is not that AI-safety research is becoming frontier-dependent; it is that more of the field became capability-assisted once capable chat models arrived, and has stayed there.
A pause in frontier capabilities would probably not make most AI-safety research impossible, though it could slow down progress in evals (e.g. discovering new capabilities) and other capability-dependent fields like Scalable Oversight and Control & Monitoring.
Limitations
The labels are LLM-generated. Only ten have been manually labelled so far. Agreement on the frontier-dependent category is therefore unmeasured though the distribution of classifications was robust to a prompt sensitivity study.
Citations are a weak proxy for importance. They are age-sensitive, vary by venue and field, and undercount reports and research posts that scholarly indexes index poorly.
Sample sizes are uneven. 2025 alone contributes 268 of 477 scored papers, and several subfield and score-4 groups are small enough that percentages move on one or two papers.
Dataset coverage is partial. The dataset is built from shallow review posts and does not represent all AI-safety research. In particular, it cannot see unpublished work inside frontier labs, which is precisely where the most frontier-dependent research would be expected to sit.
All results are observational. No result identifies a causal effect of model access.
Appendix
How robust is the central measure?
The capability-dependence score carries most of the report's weight and is produced by an LLM applying a written rubric. Two checks were run against it.
Prompt sensitivity
The same 100-paper random sample was classified three times: with the default rubric, with a View-1-leaning rubric more willing to assign capability assistance or frontier dependence, and with a View-2-leaning rubric that is more conservative about both.
The results below are of all 100 sampled papers, so each row falls short of 100% by its unscored remainder since a stricter rubric makes the classifier return unclear more often:
Prompt condition
Scored
Mean
Median
Scores 1–2
Score 3
Score 4
Default
91
2.21
2
59%
32%
0%
View-1-leaning
98
2.38
2
53%
40%
5%
View-2-leaning
95
2.18
2
65%
30%
0%
Comment. The qualitative conclusion survives deliberate pressure in both directions. Even the rubric written to favour View 1 places only 5% of papers at score 4 and leaves the median at 2. What moves is the boundary between scores 2 and 3 — a swing of roughly ten percentage points across conditions. So the claim “frontier dependence is rare” is robust, while the precise size of the capability-assisted group is rubric-sensitive and should be read as approximate.
Manual audit
I hand-labelled the first ten papers of a fixed random sample and compared them with the classifier. The model matched exactly on 5 of 10; all five disagreements were by one level, and none involved score 4. The model's mean was 2.30 against my 2.00, and it assigned the higher score in four of the five disagreements.
Comment. Every disagreement fell on the same seam: the 2/3 boundary. The model treats a measured scaling gain such as a larger judge agreeing better with humans or a stronger generator producing better data as material strengthening and scores 3. My manual rule was stricter: if the central claim was already established on a non-frontier model, it stays at 2. This is a rubric-interpretation difference rather than random error, and it means the score-3 share reported here is more likely an over-estimate than an under-estimate. It does not affect the score-4 result: none of these papers moved into frontier dependence under either rule.
Executive Summary
Leaders of frontier AI labs argue that building more capable models is a prerequisite for progress on AI safety. Critics argue that the two are largely decoupled, and that capability progress is net-negative. Despite the importance of the disagreement, it has rarely been tested against data. This report tests it directly.
I assembled a dataset of 521 AI-safety contributions (papers and blog posts) published between 2019 and 2025, drawn from the 2023, 2024, and 2025 shallow review posts, and used an LLM classifier with a documented rubric to score how strongly each contribution's central result depended on access to frontier-model capability. 477 contributions received a score; 44 had insufficient evidence to classify and were left unscored. The data points to five conclusions:
Headline conclusion
AI capabilities progress and more capable models often benefited AI safety research contributions. They made safety techniques more effective, showed the behaviours under study more often and more realistically, and did much of the experimental work of scoring, classifying, and generating data. But frontier models were rarely necessary: In this dataset the core phenomenon or method was usually implemented using non-frontier models. These results contradict the strong form of both the capability-enabled (View 1) and capability-independent (View 2) hypotheses, and points to a middle position that varies sharply by AI-safety subfield: Subfields like Evaluations & Benchmarks have a relatively high median capability-dependence score of 3. In contrast, Theory & Agent Foundations and Interpretability have a median capability-dependence score of 1 and 2.
Disclaimer: Every result in this report is descriptive. It shows which view the observable publication record more closely resembles. It does not estimate the causal effect of capabilities progress.
Introduction and the two hypotheses
Several leaders of frontier AI labs have argued that AI capabilities progress is beneficial because it is a prerequisite for progress on AI safety, or is at least highly correlated with it. Critics respond that this mindset is an excuse to build increasingly capable AI for profit, status, and personal gain and believe that AI safety and AI capabilities are uncorrelated, so capabilities progress without simultaneous safety progress is irresponsible and net-negative for the world.
View 1: The capability-enabled safety hypothesis (CESH)
The first view is that making advances in AI safety requires AI capabilities progress and the ability to work on frontier models. Anthropic's 2023 Core Views on AI Safety post states that a major reason the organization exists is the belief that safety research must be done on frontier systems:
Sam Altman's 2023 post Planning for AGI and beyond makes a related argument: that separating safety from capabilities is a false dichotomy, and that OpenAI's best safety work has come from its most capable models:
Summarized, this view holds that: many concerning risks such as deception arise only in frontier models and cannot be studied in smaller ones; some advanced safety techniques only work at frontier scale; and safety research far behind the frontier therefore has limited value.
The capability-enabled safety hypothesis (CESH) (View 1): Progress in AI capabilities tends to drive the discovery of new model phenomena and the development of better safety techniques, because a given safety insight or method typically only becomes available once
capabilities reach the level that produces the phenomenon or supports the method.
View 2: The capability-independent safety hypothesis (CISH)
The opposing view is that safety and capabilities have a low correlation: safety researchers can continue to make significant advances with non-frontier models, and capabilities research contributes little to safety.
According to this view, much or most important AI safety work can be achieved with theory or non-frontier models such as GPT-2 and capabilities progress is harmful because it enables new risks such as deception or situational awareness without contributing much to techniques for mitigating these risks.
Eliezer Yudkowsky supported a similar view in his 2023 Time essay saying:
Similarly Dr. Roman Yampolskiy has shared similar concerns:
The capability-independent safety hypothesis (CISH) (View 2): AI capabilities and AI safety progress are largely decoupled. Safety advances do not require frontier capabilities and can be achieved through theory and non-frontier models, while capabilities progress contributes little to safety and may actively harm it.
Comparing the two views
These two views make different, testable predictions about the published record of AI safety research. View 1 predicts that safety contributions cluster near the release of frontier capabilities and concentrate in frontier labs using closed models. View 2 predicts long and variable lags between a capability becoming available and the corresponding safety result appearing, substantial output from academic and non-profit groups on open models, and important discoveries made on models years behind the frontier.
View 1: CESH
View 2: CISH
Do capabilities advances boost safety progress?
Yes, they enable studying new phenomena such as deception.
No,frontier models do not help much and may harm safety.
Are frontier models necessary for new techniques and discoveries?
Yes, for both.
No, models far behind the frontier are sufficient.
Key enablers of safety progress
Frontier models and the compute to run them.
Open-weight models, theory, human creativity and time.
Where important safety work happens
Frontier labs, closed models, before or soon after release.
Academia and non-profits, open and non-frontier models.
Would an AI pause help safety?
No, halting capabilities would significantly slow safety too.
Yes, it would buy the safety community more time to make progress.
Method and dataset
Using the 2023, 2024, and 2025 Shallow Review of Technical AI Safety blog posts as sources, I built a dataset of AI safety papers and blog posts. Each entry was enriched with an LLM classifier adding the AI safety field, AI models used in the paper, the organization type involved, citation count and the capability-dependence classification that is the core measurement for this report.
The capability-dependence rubric
The central variable of this analysis is a 1-4 score for each paper estimating the degree to which the paper’s results were dependent on access to frontier models and capabilities.
Each paper was classified using this rubric using GPT-5.6-sol using the paper text, publication date, models used in the paper, and information about the most capable LLM model at the time. The model’s classifications were calibrated using manual classifications and by testing variations of the prompt.
When classifying a paper, merely using a frontier model does not earn a score of 4: the paper must supply evidence distinguishing capability assistance from frontier necessity. Scores 1-2 are treated as evidence for View 2, score 3 as intermediate, and score 4 as evidence for View 1.
Score
Label
Meaning
1
Capability-independent
Safety progress is produced independently of empirical access to AI models.
2
Non-frontier models sufficient
Empirical safety research is successfully conducted on smaller, older, specialized, or otherwise non-frontier models.
3
Capability-assisted
More capable models materially strengthen the safety research, but the central phenomenon or method is also established below the frontier.
4
Frontier-dependent
Qualitatively frontier-specific capabilities, behavior, or access are necessary to establish the central safety result.
Capability dependence classification examples
Here are some example classifications to illustrate the difference between the different scores:
Score 1. Defining and Characterizing Reward Hacking (2022, Theory & Agent Foundations, 187 citations). The central contribution consists of formal definitions, mathematical proofs, and computational examples of reward orderings. The paper does not use an AI model to establish any central result, so the evidence directly supports the capability-independent category and score 1.
Score 2. Defending Against Unforeseen Failure Modes with Latent Adversarial Training (2024, Robustness & Security, 88 citations). The method and central robustness result were directly demonstrated using ResNet-50, DeBERTa-v3-large, and Llama-2-7B-chat. These were established non-frontier systems during the relevant research period: Llama-2-7B-chat was substantially below the general capability frontier, while ResNet-50 and DeBERTa-v3-large were task-specific models rather than frontier general-purpose systems. The paper therefore establishes its contribution without frontier-model access, directly supporting the non-frontier-models-sufficient category and score 2.
Score 3. Agentic Misalignment: How LLMs Could Be Insider Threats (2025, Evaluations & Benchmarks, 156 citations). The study included frontier models such as o3 and Claude Opus 4 as well as older non-frontier models including GPT-4o, Claude Haiku 3.5, and Claude Opus 3; o3 was tested in adapted scenarios after misunderstanding the original setup. It directly observed agentic misalignment below the frontier: DeepSeek-R1 blackmailed in 79% of the main scenario trials, and Llama 4 Maverick blackmailed after a small prompt adjustment. Testing o3, Claude Opus 4, and other leading systems extended the finding to current frontier agents and made it directly relevant to contemporary deployments, while the non-frontier results establish that frontier access was not required to observe the phenomenon. This direct cross-capability evidence supports the capability-assisted category and score 3.
Score 4. Deliberative Alignment: Reasoning Enables Safer Language Models (2024, Alignment & Value Learning, 289 citations). The paper establishes deliberative alignment using o1, a frontier reasoning model that matched the representative most capable model available at first publication. Although the internal research start date is unavailable, the experiments necessarily used pre-release o1, placing the work at or near the capability frontier during the relevant research period. The central empirical result is specification-based safety deliberation in that frontier reasoning system, and the paper reports that additional reasoning compute improves safety performance. No non-frontier model in the study establishes this central result. The result is therefore specific to access and reasoning behavior demonstrated at the frontier, supporting the frontier-dependent category and score 4.
Dataset details
Results
I came up with six questions to help guide the analysis of the data and separate evidence supporting View 1 or 2:
1. How dependent is AI-safety research on access to advanced AI capabilities?
2. Are important AI safety contributions more capability-dependent?
3. Are important AI safety contributions more likely to be produced by frontier AI labs?
4. How does capability dependence vary across AI-safety subfields?
5. Do safety contributions rely on open models, closed models, or both?
6. Is the field becoming more capability-dependent over time?
Question 1: How dependent is AI-safety research on access to advanced AI capabilities?
What would count as evidence for either view: For View 1: a large share of work scoring 4, or classified as requiring frontier capability. For View 2: a large share scoring 1-2, classified as theoretical or feasible with non-frontier models.
Figure 1. Capability-dependence scores across all 477 scored contributions. Nearly two-thirds sit at scores 1–2; only 2.7% reach score 4.
Figure 2. The same distribution as a stacked bar chart, with the median (2) and mean (2.26) marked. Score 2 alone accounts for over half the field.
Result. The distribution is heavily weighted toward the low end. Score 2 alone accounts for 52.0% (248 papers), and score 1 for a further 12.4% (59), so 64.4% (307 papers) fall in the View 2 range. The intermediate score 3 covers 32.9% (157). Only 2.7% (13 papers) are classified as frontier-dependent. The median is 2 and the mean 2.26, placing the centre of the field below the capability-assisted level.
Comment. This is the simplest and most direct test in the dataset, the results here are strong evidence against the strong form of View 1. If frontier access were generally necessary for safety progress, score 4 should be a large category rather than a residue of thirteen papers. But the result does not support the strong form of View 2 either: a third of the field sits at score 3, where more capable models demonstrably strengthened the work. Therefore stronger models frequently helped produce results but were rarely fully necessary.
Note: The scores here measure counterfactual dependence, not model choice. They do not say that only 2.7% of AI safety papers use frontier models – many more did. Instead, these results say that in 2.7% of contributions, the central result could not have been established without access to frontier AI capabilities.
Question 2: Are important AI safety contributions more capability-dependent?
Citation counts are heavy-tailed, so a small number of contributions account for a large share of the field's attention. If the average paper does not need frontier models but the important papers do, View 1 would still be substantially correct. This question tests this idea.
Figure 3. Capability-dependence composition of all scored papers compared with the top quartile by age-adjusted citations.
Result. Among the 96 scored papers above the 45.8 citations-per-year cutoff (>75th percentile annual citations), the View 2 share falls from 64.4% to 44.8%, the intermediate share rises from 32.9% to 47.9%, and the View 1 share rises from 2.7% to 7.3%. The mean rises from 2.26 to 2.60 and the median from 2 to 3, a full category change.
Comment. The direction of this shift favours View 1, and it is one of the clearest pro-View-1 results in the report: the typical highly cited paper is capability-assisted, while the typical paper overall is not. However, even among the most influential work, 92.7% still scores 1–3, and nearly half remains in the View 2 range and 7.3% of important contributions with a score of 4 is not high enough to fully support View 1.
Question 3: Are important AI safety contributions more likely to be produced by frontier AI labs?
Frontier labs have privileged access to advanced and pre-release models. If that access is what unlocks important safety progress, frontier labs should be over-represented among important contributions and produce a larger share of the field’s influential work than its overall output.
Figure 4: Capability-dependence scores for frontier-lab-affiliated papers compared with papers from other known affiliations. Frontier-lab work sits higher on the scale, but the gap is mostly a shift from score 2 into score 3.
Figure 5. Frontier-lab affiliation among all papers with known affiliations compared with highly cited papers.
Result - capability dependence (Figure 4). Frontier-lab-affiliated papers score higher: mean 2.53 versus 2.19, median 3 versus 2. Their score-3 share is 46.6% against 28.2%, a gap of 18.4 percentage points, and their score-4 share is 6.1% against 1.1%. Though it rests on nine papers against three. 49 scored papers with unknown affiliation are excluded.
Result - share of influential work (Figure 5). Frontier-lab-affiliated papers are 36.3% of the 465 papers with known affiliations, but 52.0% of the 98 highly cited ones, a gap of 15.7 percentage points. (56 papers overall and 5 highly cited papers have unknown affiliations and are excluded.)
Comment. Both charts point the same way, and both are consistent with View 1: frontier labs are over-represented among influential safety work relative to their share of output, and their own work sits higher on the capability-dependence scale. Together they make the strongest pro-View-1 case in the report.
But note that for the contributions of frontier labs, the movement is mostly from “non-frontier models sufficient” into “capability-assisted”, not into frontier necessity. Even inside the organizations with the best model access in the world, 47.3% of safety output is classified as not needing that access, and 93.9% is not classified as frontier-dependent.
Question 4: How does capability dependence vary across AI safety subfields?
An aggregate number can hide the fact that AI safety is several fields wearing one label. Do evaluations, scalable oversight, and control depend more on advanced models than interpretability, alignment, or theory?
Figure 6. Capability-dependence composition by AI-safety subfield, ordered by median and then mean score.
Result. The spread is the widest of any comparison in this report. Evaluations & Benchmarks is the only subfield with a median of 3, and it is the only one where the intermediate score is the majority position (62.5%). At the other end, every one of the 43 papers in Theory & Agent Foundations scores 1-2, with a mean of 1.10. Interpretability, the second-largest subfield at 104 papers, is 86.5% in the View 2 range.
Comment. This is where the debate should actually be conducted. Subfields whose research object is the behaviour of an advanced model such as evaluating what frontier systems do, catching them misbehaving inherit their object of study’s capability requirements. Control & Monitoring has the highest frontier-dependent share of any subfield at 8.9%, which makes sense: protocols for supervising untrusted models are hard to validate against a model too weak to subvert anything. Subfields whose object is a mechanism or a formalism do not inherit that requirement, and interpretability and alignment’s concentration at low scores are the clearest single piece of evidence for View 2 in the dataset.
The practical implication of this result is that a blanket claim about “AI safety research” is unlikely to be true in either direction as the capability-dependence of each subfield varies significantly: A pause in frontier capability would not significantly affect theory and interpretability while potentially slowing down progress in more capability-dependent fields like Scalable Oversight and Control & Monitoring.
Question 5: Do safety contributions rely on open models, closed models, or both?
If View 1 were true, highly cited papers would use closed models more than papers overall. If View 2 were true, open weights models would be widely used for both important contributions and other contributions.
Figure 7. Paper-level model access mix among all eligible papers and among highly cited papers.
Result. Among 330 eligible papers (51 excluded for unknown or no model), 64.2% use a mixed set of open and closed models, 25.8% are open-weights-only, 8.2% closed-API-only, and 1.8% internal-closed-only. Among the 93 eligible highly cited papers, mixed access rises to 72.0% and open-only falls by half, to 12.9%. Closed-API-only is essentially unchanged (8.2% to 11.8%).
Comment. What separates highly cited work is not closed access and it is not openness either, since open-only work halves as influence rises. It is breadth of access: the ability to run the same experiment across open and closed models at once. Mixed-access papers dominate both groups and dominate the highly cited group more. Note what the highly cited group loses and gains: it sheds open-only papers without picking up closed-only ones.
Question 6: Is the field becoming more capability-dependent over time?
As models improve, does safety research follow them up the capability ladder? View 1 predicts rising dependence; View 2 predicts that low-dependence work remains a stable or growing part of the field.
Figure 8: capability-dependence distributions by year.
Result. There is one step change and then a plateau. Between the pooled 2019-2022 period and 2023 the View 2 share drops from 88.0% to 61.3% and the intermediate share nearly triples. From 2023 to 2025 the composition is close to flat: View 2 moves between 61.3% and 64.8%, and the intermediate share drifts up by three points. The frontier-dependent share does not rise at all and even it falls from 6.5% in 2023 to 1.9% in 2025.
Comment. This result favors View 2 though it’s not as unambiguous as some of the other results. The step between the pre-2023 period and everything after it looks could be explained by the arrival of capable instruction-tuned chat models as a research substrate, after which the field settled into a stable mix. Since then, the capability frontier has advanced enormously while the minimum level of capabilities in LLMs needed for AI safety research has stayed the same and strict frontier dependence has become less common.
Conclusions
The evidence rejects both strong hypotheses and supports a middle position that neither camp usually states.
View 1 is wrong in its strong form. If frontier access were generally necessary for AI-safety progress, the frontier-dependent category should be large. In reality, it contains 13 of 477 papers. It stays small among the most highly cited work (7.3%), inside frontier labs themselves (6.1%) and in every subfield (highest: 8.9% in Control & Monitoring). This is the most consistent finding in the report.
View 2 is wrong in its strong form. A third of the field, half of the highly cited work, and nearly two-thirds of Evaluations & Benchmarks sits at score 3, where stronger models measurably strengthened the research. Influential papers use frontier and near-frontier models far more often than the average paper. Capabilities are not irrelevant to safety and they are routinely useful.
A statement that the data and analysis in this post supports:
Three points sharpen that summary:
Revisiting the original claims
Anthropic's frontier-safety claim holds in a limited form. Some phenomena do appear to require advanced models such as alignment faking, deliberative alignment, and evaluating control protocols against a model capable of subverting them are all in the score-4 category. But if most important empirical safety research required frontier systems, that set would be much larger than 13 out of 477 papers.
The “safety and capabilities together” claim. The capability-assisted category is large, larger among highly cited work, and larger inside frontier labs. The data narrows the claim rather than refuting it: correlation, yes; necessity, usually not.
The “capabilities running ahead” concern is partly supported. A great deal of important safety work turns out not to have required frontier models, which is consistent with a backlog of tractable research that needs more time, not capability progress to unlock.
Some evidence-based statements
The following statements are not quotes from any lab or critic. They are what the claims in Section 1 would look like if rewritten based on the findings from this dataset:
Limitations
Appendix
How robust is the central measure?
The capability-dependence score carries most of the report's weight and is produced by an LLM applying a written rubric. Two checks were run against it.
Prompt sensitivity
The same 100-paper random sample was classified three times: with the default rubric, with a View-1-leaning rubric more willing to assign capability assistance or frontier dependence, and with a View-2-leaning rubric that is more conservative about both.
The results below are of all 100 sampled papers, so each row falls short of 100% by its unscored remainder since a stricter rubric makes the classifier return unclear more often:
Prompt condition
Scored
Mean
Median
Scores 1–2
Score 3
Score 4
Default
91
2.21
2
59%
32%
0%
View-1-leaning
98
2.38
2
53%
40%
5%
View-2-leaning
95
2.18
2
65%
30%
0%
Comment. The qualitative conclusion survives deliberate pressure in both directions. Even the rubric written to favour View 1 places only 5% of papers at score 4 and leaves the median at 2. What moves is the boundary between scores 2 and 3 — a swing of roughly ten percentage points across conditions. So the claim “frontier dependence is rare” is robust, while the precise size of the capability-assisted group is rubric-sensitive and should be read as approximate.
Manual audit
I hand-labelled the first ten papers of a fixed random sample and compared them with the classifier. The model matched exactly on 5 of 10; all five disagreements were by one level, and none involved score 4. The model's mean was 2.30 against my 2.00, and it assigned the higher score in four of the five disagreements.
Comment. Every disagreement fell on the same seam: the 2/3 boundary. The model treats a measured scaling gain such as a larger judge agreeing better with humans or a stronger generator producing better data as material strengthening and scores 3. My manual rule was stricter: if the central claim was already established on a non-frontier model, it stays at 2. This is a rubric-interpretation difference rather than random error, and it means the score-3 share reported here is more likely an over-estimate than an under-estimate. It does not affect the score-4 result: none of these papers moved into frontier dependence under either rule.