A map of the AI safety field's problems and agendas, and a request for your ratings
We builtaisafetyagendas.com, an interactive map of AI safety research agendas and how they map to different problems in alignment. The rows are 12 core problems, the columns are research areas, and inside you can find 58 research agendas. Each cell is the intersection of a problem and an area: the number tells you how many agendas target that problem, the colour tells you how mature they are. We did a first pass ourselves, using our own judgement, but the first pass is not the point.
The point is that every cell is a question aimed back at the community: is this rating right? You can rate the cells in your area, tell us how familiar you are with it, and read the map as best case, average, or worst case depending on how optimistic you feel. It was built as part of the Safe AI Germany Incubator.
Figure 2: From left to right best, average and worst case rating examples.
The allocation problem
The field cannot see its own resource allocation.Leech and Lynn put it as: "you can't optimise an allocation of resources if you don't know what the current one is". Wentworth goes further, arguing that the memetically successful strategy is to work on easy problems rather than "plausible bottlenecks to humanity's survival". The IAPS Expert Survey gets to a similar place from another direction, warning that funders and researchers concentrate on the most visible work and displace quieter research that matters more.
IAPS highlights as a cause the lack of comparative ranking of research directions because what actually exists are mainly lists that don’t build up: Anthropic's recommended directions, DeepMind's technical AGI safety approach,Redwood's project proposals, the UK AISI Alignment Project's £15M bet on areas it thinks the field underrates, and independent reviewers redoing the exercise every year. Written in April 2025, DeepMind placed itself closest to Anthropic's Core Views, by then two years old, with much more weight on robust training, monitoring and security, and set itself against OpenAI's focus on automating alignment research. How is this disagreement taken into account by the field? These different views aren’t lively contested in official channels.
Plenty of people are already pulling in this direction. The IAPS survey asks specialists which directions are most promising.AISafetyBenchExplorer catalogues 195 AI safety benchmarks so they can be discovered and compared. Apart Research's alignment navigation graph clusters 5000+ alignment papers with an LLM. AI Plans keeps a compendium of alignment plans with strength and vulnerability votes. AI Safety Ideas lets people post questions and fund the answers. We built ours to cover a piece we see is left: problems and research areas together, so we can see how much coverage each problem is getting.
What the platform does
The community feedback we get is aggregated so it can be analysed.
Figure 3: a visualization of the feedback.
Three things we’re aiming for:
Identifying bottlenecks, underindexed areas, overindexed areas, and high-leverage areas.
Making the community's actual opinion legible, including the disagreements and trends.
Giving AI safety metaresearch a common structure to argue inside.
Our own mapping is the scaffolding. We read the research areas, listed the problems, and assigned research agendas to the problems they target. Constitutional AI, for example, we mapped to both outer alignment and corrigibility.
We set twelve problems and 58 agendas with our best judgment, but it isn’t meant to be relied on; it’s meant to be a crowdsourcing tool, so you can see our own ratings and the community's.
With the feedback data, the site sorts then every problem and every research area into three categories, recomputed from whichever ratings you're viewing, ours or the community's:
Easy to intervene: cells with mature research agendas but where work is left. For example, model organisms of misalignment have a strong existence proof against inner alignment and scheming, but an open-source replication didn't fully reproduce it, so the next step is replicating it outside one lab.
Bottlenecks: a problem many areas depend on but where maturity is low. Alignment measurement (P1) is reached by 16 agendas, and half of them are untested.
Underexplored: a cell with 0 agendas that should or could have some. The sharp left turn (P3) is reached by only 4 of the 12 research areas.
What we want from you
Go toaisafetyagendas.com and rate the maturity of the areas you know; it only takes a few minutes! Ratings show up on the map as soon as you submit them, so you can see where yours is against everyone else's.
We also want feedback on the thing itself: whether the partition of problems and research agendas is right, whether the aggregation makes sense, whether the interface gets in the way. Once there's enough data, we'll share it and write up the analysis of the ratings, but it will also be available interactively in real time.
A map of the AI safety field's problems and agendas, and a request for your ratings
We built aisafetyagendas.com, an interactive map of AI safety research agendas and how they map to different problems in alignment. The rows are 12 core problems, the columns are research areas, and inside you can find 58 research agendas. Each cell is the intersection of a problem and an area: the number tells you how many agendas target that problem, the colour tells you how mature they are. We did a first pass ourselves, using our own judgement, but the first pass is not the point.
Figure 1: The Map at aisafetyagendas.com
The point is that every cell is a question aimed back at the community: is this rating right? You can rate the cells in your area, tell us how familiar you are with it, and read the map as best case, average, or worst case depending on how optimistic you feel. It was built as part of the Safe AI Germany Incubator.
Figure 2: From left to right best, average and worst case rating examples.
The allocation problem
The field cannot see its own resource allocation. Leech and Lynn put it as: "you can't optimise an allocation of resources if you don't know what the current one is". Wentworth goes further, arguing that the memetically successful strategy is to work on easy problems rather than "plausible bottlenecks to humanity's survival". The IAPS Expert Survey gets to a similar place from another direction, warning that funders and researchers concentrate on the most visible work and displace quieter research that matters more.
IAPS highlights as a cause the lack of comparative ranking of research directions because what actually exists are mainly lists that don’t build up: Anthropic's recommended directions, DeepMind's technical AGI safety approach, Redwood's project proposals, the UK AISI Alignment Project's £15M bet on areas it thinks the field underrates, and independent reviewers redoing the exercise every year. Written in April 2025, DeepMind placed itself closest to Anthropic's Core Views, by then two years old, with much more weight on robust training, monitoring and security, and set itself against OpenAI's focus on automating alignment research. How is this disagreement taken into account by the field? These different views aren’t lively contested in official channels.
Plenty of people are already pulling in this direction. The IAPS survey asks specialists which directions are most promising. AISafetyBenchExplorer catalogues 195 AI safety benchmarks so they can be discovered and compared. Apart Research's alignment navigation graph clusters 5000+ alignment papers with an LLM. AI Plans keeps a compendium of alignment plans with strength and vulnerability votes. AI Safety Ideas lets people post questions and fund the answers. We built ours to cover a piece we see is left: problems and research areas together, so we can see how much coverage each problem is getting.
What the platform does
The community feedback we get is aggregated so it can be analysed.
Figure 3: a visualization of the feedback.
Three things we’re aiming for:
Our own mapping is the scaffolding. We read the research areas, listed the problems, and assigned research agendas to the problems they target. Constitutional AI, for example, we mapped to both outer alignment and corrigibility.
We set twelve problems and 58 agendas with our best judgment, but it isn’t meant to be relied on; it’s meant to be a crowdsourcing tool, so you can see our own ratings and the community's.
With the feedback data, the site sorts then every problem and every research area into three categories, recomputed from whichever ratings you're viewing, ours or the community's:
What we want from you
Go to aisafetyagendas.com and rate the maturity of the areas you know; it only takes a few minutes! Ratings show up on the map as soon as you submit them, so you can see where yours is against everyone else's.
We also want feedback on the thing itself: whether the partition of problems and research agendas is right, whether the aggregation makes sense, whether the interface gets in the way. Once there's enough data, we'll share it and write up the analysis of the ratings, but it will also be available interactively in real time.