tl;dr: people should understand and think hard about the problems they work on.
We’ve observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good, doing pragmatic alignment research might be good, and as long as such objectives don’t breach our internal models of what could contribute to reducing x-risk, these things are “what should be done”. But using such vibesy thought processes don’t always produce “actually impactful work” that would beat a prospective counterfactual. We wrote this post to share our observations and figure out what we should be doing instead.
People don’t know what they’re working on
AI safety is talent constrained. However, simply inflating the field doesn’t solve our bottleneck; rather, we need more people who understand the core arguments of AI safety. You can’t determine how to meaningfully contribute to AI safety without deeply knowing the problem you are trying to solve. Many newer people (us included!) rush into research, fellowships, and the like without building the context necessary for navigating the field.
Agency-maxxing is not always good
Moving fast is good. Moving too fast leads to poor ToC and sloppily executing projects. Many people working in AI safety seem to downweight spending time thinking about the actual problem they are trying to solve and backchaining from it, in favor of moving as fast as possible to get things done quick and dirty. While this enables faster execution and increases the volume of work being done, it doesn’t always move the needle on effectively reducing x-risk.
The problem with force multipliers
We categorize ToCs that focus on enhancing the impact of others as force multipliers (think upskilling, infrastructure work). There’s strong incentives to become a force multiplier: enhancing the impact of others through auxiliary work is (arguably) higher ROI and often more approachable than “directly” working on AI safety.
However, when everyone becomes a force multiplier, no one ends up directly working on the problems we actually care about. We think this work is extremely valuable, but we believe there is comparative advantage in focusing on the direct work, and that a small number of force multipliers can produce the same (if not larger) multiplicative effect as current efforts.
Harshul: In Balatro, your final score is calculated as the number of chips you score (blue) times your mult (red). I’m skeptical of a mult-heavy build (focusing on auxiliary work) without a strong base of chips to actually multiply (direct safety work). I’m more bought into a chip-heavy build (focusing on direct work, which I believe benefits more from added effort) alongside a smaller fraction of the field providing sufficient mult (auxiliary support).
Deferring thinking to others
Our subfield includes a multitude of conflicting agendas: whether or not alignment research is “good” to do, how much or little we should be focusing on political advocacy, how effective comms work is and could be, etc. Such conflicts don’t yet have a resolution. But, in order to “do things”, people will defer to a respected figure as opposed to building their own reasoning for their models. They will defer on what things are good and bad as a way to consolidate their consciousness and start doing things to “make impact”.
As an example, upcoming technical researchers will often defer to what is popular as a proxy for impact. They may look at places like Redwood or Anthropic, and attest to wanting to work there to “do the most impactful work”. But in reality, it's very much up for debate as to whether a technical researcher should be going to said places in order to produce counterfactual impact (not to mention whether they should be doing technical research at all). But they instead defer to the vibes/popularity these places hold and conflate this with prospective impact.
AI safety is hard, but not everyone ends up working on the hardest problems. Instead, people focus on more tractable problems that may seem nearly as impactful, but actually fail solving the true bottlenecks at hand. We believe this arises through some combination of people getting turned away from working on hard problems after getting stuck, and people prioritizing fast execution over thinking deeply about their work. When pushed to “just do things”, it's often easiest to “just do” the easiest things.
(Note: Hard problems aren’t impactful by default, and easy problems aren’t auxiliary by default either. This section applies to the set of hard problems that are actually useful to work on.)
How to avoid these:
Build your models
Pursue the fundamentals: Read early AI Safety/alignment work, sequences, etc.
Read more modern work, stay up to date on meaningful new literature
Reflect on your reading regularly. Analyze when you update and what uncertainties you have after reading a piece.
Construct feedback loops
Seek lots of feedback on your direction/work from different demographics, especially those with conflicting views.
Work in public and let others help you course correct
Try and maintain strong epistemic, and avoid subtle dishonesties
We do not claim to have it all figured out ourselves. We have not thought as deeply as many others, and this post is not meant to be a “call out” of people who have exhibited the aforementioned behaviors. Rather, this is a call for improvement, and a raise in standards. We do not haveto sacrifice the understanding of our problems, haphazardly defer, or avoid what is difficult. We can try harder, be smarter, and raise the bar of the work we devote our time to.
tl;dr: people should understand and think hard about the problems they work on.
We’ve observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good, doing pragmatic alignment research might be good, and as long as such objectives don’t breach our internal models of what could contribute to reducing x-risk, these things are “what should be done”. But using such vibesy thought processes don’t always produce “actually impactful work” that would beat a prospective counterfactual. We wrote this post to share our observations and figure out what we should be doing instead.
People don’t know what they’re working on
AI safety is talent constrained. However, simply inflating the field doesn’t solve our bottleneck; rather, we need more people who understand the core arguments of AI safety. You can’t determine how to meaningfully contribute to AI safety without deeply knowing the problem you are trying to solve. Many newer people (us included!) rush into research, fellowships, and the like without building the context necessary for navigating the field.
Agency-maxxing is not always good
Moving fast is good. Moving too fast leads to poor ToC and sloppily executing projects. Many people working in AI safety seem to downweight spending time thinking about the actual problem they are trying to solve and backchaining from it, in favor of moving as fast as possible to get things done quick and dirty. While this enables faster execution and increases the volume of work being done, it doesn’t always move the needle on effectively reducing x-risk.
The problem with force multipliers
We categorize ToCs that focus on enhancing the impact of others as force multipliers (think upskilling, infrastructure work). There’s strong incentives to become a force multiplier: enhancing the impact of others through auxiliary work is (arguably) higher ROI and often more approachable than “directly” working on AI safety.
However, when everyone becomes a force multiplier, no one ends up directly working on the problems we actually care about. We think this work is extremely valuable, but we believe there is comparative advantage in focusing on the direct work, and that a small number of force multipliers can produce the same (if not larger) multiplicative effect as current efforts.
Harshul: In Balatro, your final score is calculated as the number of chips you score (blue) times your mult (red). I’m skeptical of a mult-heavy build (focusing on auxiliary work) without a strong base of chips to actually multiply (direct safety work). I’m more bought into a chip-heavy build (focusing on direct work, which I believe benefits more from added effort) alongside a smaller fraction of the field providing sufficient mult (auxiliary support).
Deferring thinking to others
Our subfield includes a multitude of conflicting agendas: whether or not alignment research is “good” to do, how much or little we should be focusing on political advocacy, how effective comms work is and could be, etc. Such conflicts don’t yet have a resolution. But, in order to “do things”, people will defer to a respected figure as opposed to building their own reasoning for their models. They will defer on what things are good and bad as a way to consolidate their consciousness and start doing things to “make impact”.
As an example, upcoming technical researchers will often defer to what is popular as a proxy for impact. They may look at places like Redwood or Anthropic, and attest to wanting to work there to “do the most impactful work”. But in reality, it's very much up for debate as to whether a technical researcher should be going to said places in order to produce counterfactual impact (not to mention whether they should be doing technical research at all). But they instead defer to the vibes/popularity these places hold and conflate this with prospective impact.
Streetlighting
AI safety is hard, but not everyone ends up working on the hardest problems. Instead, people focus on more tractable problems that may seem nearly as impactful, but actually fail solving the true bottlenecks at hand. We believe this arises through some combination of people getting turned away from working on hard problems after getting stuck, and people prioritizing fast execution over thinking deeply about their work. When pushed to “just do things”, it's often easiest to “just do” the easiest things.
(Note: Hard problems aren’t impactful by default, and easy problems aren’t auxiliary by default either. This section applies to the set of hard problems that are actually useful to work on.)
How to avoid these:
We do not claim to have it all figured out ourselves. We have not thought as deeply as many others, and this post is not meant to be a “call out” of people who have exhibited the aforementioned behaviors. Rather, this is a call for improvement, and a raise in standards. We do not have to sacrifice the understanding of our problems, haphazardly defer, or avoid what is difficult. We can try harder, be smarter, and raise the bar of the work we devote our time to.