The Forecasting Research Institute (along with coauthors Jason Abaluck and Eva Vivalt) are launching Automated AI Risk Outlook (AIRO): an ensemble of frontier LLMs reguarly forecasting the probability of catastrophic risk events.
Forecasts about the likelihood of catastrophic risks from AI vary wildly. Forecasts from frontier AI models could be an important input into this debate. The best models now approach superforecaster levels of accuracy, and models are rapidly becoming more capable.
To take advantage of these capabilities, we built a fully automated dashboard of catastrophic risk outcomes as forecast by frontier AI models. AIRO enables us to track in close to real time how the likelihood of catastrophic outcomes changes, and how these risk forecasts relate to advances in model capability.
The AIRO dashboard currently puts the likelihood of an AI-related catastrophe that kills at least 10% of the population at 0.47% by the end of 2030 and 6% by the end of 2050. AIRO also tracks risks from biorisk, cyberrisk, and misalignment risk, as well as the likelihood of human disempowerment.
How does AIRO work?
AIRO uses an ensemble of the top-scoring four models on the Epoch Capabilities Index (ECI), keeping one model per family. At the time of launch, these were GPT-6 Astra, Fable 5.1, Opus 5, and GPT-5.5 Pro. We report separate forecasts from these frontier models as well as an ensemble forecast calculated as their unweighted median.
We define a catastrophic outcome as an event that kills at least 10% of the global population and ask for the likelihood of such an event happening across various time horizons, from within six months to 2100. We track various severity thresholds, each of which can be reached through deaths, or through equivalent morbidity or economic damages (for example, 100,000 deaths or $220 billion in damages).
We also ask LLMs to forecast these outcomes conditional on the score of the top AI model on the ECI, giving us a sense of how AI capabilities are linked to the probability of catastrophic outcomes.
Why trust LLM risk forecasts?
AI forecasting abilities have improved alongside other AI capabilities. In July 2026, an AI model reached parity with superforecasters on FRI’s forecasting benchmark, ForecastBench. This gives some basis for believing that AI forecasts may now be at least as good as the best available human forecasts.
However, the forecasts shown on AIRO are different from those evaluated in ForecastBench in two important ways. First, they concern very rare events. It is difficult to measure accuracy on forecasting rare events given that they, by definition, happen so rarely. However, unlike human forecasters, we can ask AI models to make thousands of predictions on events in a simulated world environment. When we do this, we find that more capable models—as assessed by their score on the ECI—are more accurate at forecasting rare events. (See our FAQ).
Second, AIRO includes many conditional forecasting questions. The dashboard includes forecasts conditional on model capabilities; the white paper also explores preliminary policy scenarios. We can again test performance on this type of conditional forecasting in real (observational, weather) and simulated (causal, StarSim disease simulator) environments, and again we find that accuracy correlates with general capabilities.
What's next for AIRO
We think that automating risk forecasts has some big advantages. First, LLMs can forecast many outcomes across many time horizons and severity thresholds without growing tired, allowing us to elicit many more forecasts than is possible with human forecasters. Second, AIRO also provides us with risk forecasts that update as real-world events change and model capabilities advance, giving us forecasts based on the very latest information.
AIRO is still a work in progress. We plan to keep validating and extending our dashboard by refining ensemble forecasts and developing a larger set of intermediate questions related to catastrophic risk outcomes. Some things on our mind:
Sandbagging. Misaligned models could intentionally reduce their forecast of catastrophic risks in order to evade human suspicion.
Policy-conditional forecasts. How does risk change conditional on certain policies being adopted in certain jurisdictions? These conditionals could help policymakers evaluate the impact of different policy forecasts.
Supplementing 'redlines.' Labs use evals with specific redlines to make decisions about safety and model deployment. The tooling behind AIRO could allow labs to estimate damages directly by conditioning the ensemble's estimates on benchmark performance.
Benefits alongside risks. Right now, we show the risks. But AI deployments also bring potential upsides. Showing upside risks (like GDP growth and life expectancy) could help us understand the costs and benefits associated with proposed policies.
The Forecasting Research Institute (along with coauthors Jason Abaluck and Eva Vivalt) are launching Automated AI Risk Outlook (AIRO): an ensemble of frontier LLMs reguarly forecasting the probability of catastrophic risk events.
—
Forecasts about the likelihood of catastrophic risks from AI vary wildly. Forecasts from frontier AI models could be an important input into this debate. The best models now approach superforecaster levels of accuracy, and models are rapidly becoming more capable.
To take advantage of these capabilities, we built a fully automated dashboard of catastrophic risk outcomes as forecast by frontier AI models. AIRO enables us to track in close to real time how the likelihood of catastrophic outcomes changes, and how these risk forecasts relate to advances in model capability.
The AIRO dashboard currently puts the likelihood of an AI-related catastrophe that kills at least 10% of the population at 0.47% by the end of 2030 and 6% by the end of 2050. AIRO also tracks risks from biorisk, cyberrisk, and misalignment risk, as well as the likelihood of human disempowerment.
How does AIRO work?
AIRO uses an ensemble of the top-scoring four models on the Epoch Capabilities Index (ECI), keeping one model per family. At the time of launch, these were GPT-6 Astra, Fable 5.1, Opus 5, and GPT-5.5 Pro. We report separate forecasts from these frontier models as well as an ensemble forecast calculated as their unweighted median.
We define a catastrophic outcome as an event that kills at least 10% of the global population and ask for the likelihood of such an event happening across various time horizons, from within six months to 2100. We track various severity thresholds, each of which can be reached through deaths, or through equivalent morbidity or economic damages (for example, 100,000 deaths or $220 billion in damages).
We also ask LLMs to forecast these outcomes conditional on the score of the top AI model on the ECI, giving us a sense of how AI capabilities are linked to the probability of catastrophic outcomes.
Why trust LLM risk forecasts?
AI forecasting abilities have improved alongside other AI capabilities. In July 2026, an AI model reached parity with superforecasters on FRI’s forecasting benchmark, ForecastBench. This gives some basis for believing that AI forecasts may now be at least as good as the best available human forecasts.
However, the forecasts shown on AIRO are different from those evaluated in ForecastBench in two important ways. First, they concern very rare events. It is difficult to measure accuracy on forecasting rare events given that they, by definition, happen so rarely. However, unlike human forecasters, we can ask AI models to make thousands of predictions on events in a simulated world environment. When we do this, we find that more capable models—as assessed by their score on the ECI—are more accurate at forecasting rare events. (See our FAQ).
Second, AIRO includes many conditional forecasting questions. The dashboard includes forecasts conditional on model capabilities; the white paper also explores preliminary policy scenarios. We can again test performance on this type of conditional forecasting in real (observational, weather) and simulated (causal, StarSim disease simulator) environments, and again we find that accuracy correlates with general capabilities.
What's next for AIRO
We think that automating risk forecasts has some big advantages. First, LLMs can forecast many outcomes across many time horizons and severity thresholds without growing tired, allowing us to elicit many more forecasts than is possible with human forecasters. Second, AIRO also provides us with risk forecasts that update as real-world events change and model capabilities advance, giving us forecasts based on the very latest information.
AIRO is still a work in progress. We plan to keep validating and extending our dashboard by refining ensemble forecasts and developing a larger set of intermediate questions related to catastrophic risk outcomes. Some things on our mind: