Very quick advice I gave to a friend applying to MATS about how to select AI safety projects:
* Best positive signal: if the person who came up with the project cares about existential risk
* Be very suspicious of anything "interdisciplinary". These are often not x-risk motivated and targeted at unimportant problems.
* AI companies are really good at hill climbing. Therefore:
* Improving our understanding, ideally to the point we can hill climb = high value
* Hill climbing existing metrics = low value
* Anything where hill climbing is acceleratory = potentially negative value
* Near term alignment (and some kinds of control) have dubious impact, better to focus on things that scale
* AIs may be strongly superhuman at most cheaply verifiable tasks soon. Be a complement to this, not a substitute
* Don't fight obviously losing battles, e.g. trying to prevent eval awareness
* Chem/bio/cybersecurity is low value without specific expertise
* Policy impact is probably larger than purely technical impact right now
* It’s better for technical safety work to happen in places that have policy connections, like Redwood, METR, GovAI, CAISI, UKAISI
* Be suspicious of overcomplicated methods or setups. If the setting is harder to study than other settings, it needs to be higher relevance
* Understand clearly what problem the project is solving and whether it's important. Some important problems:
* Alignment red teaming. The holy grail is being able to reproduce all precursor behaviors of the Hugging Face incident and other incidents, with no prior knowledge of the incident or even its environments, in a way that generalizes to future full loss of control incidents. I think this is the highest potential payoff problem in technical AI safety right now and deserves ~20% of the entire field's labor. Why? It scales, it's very much unsolved, progress gives AI labs evidence their models are misaligned.
* Preserving monitorability
* Model organisms