There are some conditions under which the correct thing is to whistleblow and leave, rather than keep trying to make improvements on the margin. For instance, if your job becomes "get good benchmarks on alignment" while ignoring clear signs of actual misalignment or hidden reasoning, or if you become aware of a coverup of a warning shot, then staying there means complicity in existential risk.
Recent developments underscore that anyone "working on safety/alignment" at OpenAI is at best disempowered to reduce existential risk, and at worst culpable for increasing it; and the best thing they can do right now is whistleblow. Furthermore, MATS should loudly refuse to work with OpenAI mentors. (Which is why I commented this here despite it being not a direct response to your post.)
This is not an endorsement of other labs. I find Anthropic's public statements on their alignment plan and their level of caution to be inadequate, and I think that their broken promise not to push the frontier forward has caused other companies to act more recklessly. I think that the Alphabet governance of DeepMind undermines what safety culture it had (and while we're talking about how to leave, I strongly support TurnTrout's actions as a model). And we all know how much worse the lesser labs are.
But some people in this community still say with a straight face that working at OpenAI on alignment could be helpful, and today is a great occasion to proclaim how clearly that has been falsified. (Not least because the majority of x-risk-pilled members of OpenAI have ended up leaving the company, individually or en masse, for reasons that I suspect are undisclosed because of legal threats.)
Recent developments underscore that anyone "working on safety/alignment" at OpenAI is at best disempowered to reduce existential risk, and at worst culpable for increasing it; and the best thing they can do right now is whistleblow.
This does not logically follow. For example, it could be that 3% of OpenAI employees care about x-risk and are counterfactually net positive, but the minimum number needed to prevent egregious misalignment on the scale of the HF incident is 10%. Or, it could be that they deliberately work on projects that have long-term safety value rather than just prevent the next incident. Or, it could be that internal siloing prevents them from having enough information about a warning shot or its coverup for whistleblowing to be net positive.
>and I think that their broken promise not to push the frontier forward has caused other companies to act more recklessly
What are you referring to?
Anthropic made a quasi-commitment not to release models with higher capabilities than of their rivals and an OAI-like commitment not to release dangerous models, which was backtracked in 2026.
But this doesn't explain why the hell OAI even released its Astra given that Astra was released two days after Fable 5.1, which along with Fable 5 barely outsmarted Sol on the ECI...
I think most AI safety researchers are unhelpful, but "warning shots" is a terrible argument for telling them to quit. Interestingly, "warning shots" have also been used as an argument against engaging in political advocacy, countered for similar reasons as yours. Warning shots are something we should focus on preparing a response to, if needed, not try to make more likely (directly or through intentional inaction).
The main issues with safety research for me are externalities:
Some lines of safety research, like agent foundations, sidestep these, at the cost of being hard to justify in terms of short term returns or even theory of change. If the economic landscape were different, such as by an enforced pause with strict conditions for advancement (externally verified safety proofs and the like), then I can imagine meaningful safety work happening at scale.
As of right now, the core issue is power. The AI labs have too much of it. You're either helping them get more, pushing back, or irrelevant. Pausing isn't just to "buy time," it's to change the conditions in which AI exists.
To distill the more relevant part of my other comment: I agree with you that "worse is better" is a bad category of argument, but there's a better argument in this direction.
I believe that a significant fraction of "prosaic safety work" at the frontier labs has had the effect of hiding misalignment in publicly released versions rather than making the models more robustly aligned.
For instance, I expect that the prosaic alignment improvements from Sydney to the final GPT-4 were surface-level only (as evidenced by later failures like GPT-4o), and if the alignment researchers had refused to play whack-a-mole at the time, then commercial progress would have been properly halted.
It seems like you're assuming that work labeled as safety is safety, or at least more safety than capabilities.
Wherever I use the word "safety", treat it as a placeholder for "genuinely net-beneficial for minimizing existential risk and enabling flourishing futures". What research meets this bar is the subject of some debate, of course.
I think this definition hides the hard part, no?
Like, ~everyone would agree it's good to do work that is "genuinely net-beneficial for minimizing existential risk and enabling flourishing futures".
Seems like your question is more "is ~prosaic safety work net-beneficial" (right?)
The purpose of this post is to address a specific argument against working at frontier AI companies on technical safety research that genuinely reduces the risk of warning shots like OAIxHF. There are many other reasons why people might not want to work at frontier AI companies beyond the scope of this post; I retitled the post to clarify this.
My definition definitely hides the hard part; maybe too much, as I've implicitly assumed the perfect alignment techniques are deployed, society adapts, CEV is achieved, etc. I also don't have high confidence in what technical AI safety research is sufficient to solve the technical challenges of alignment. But I'm trying to discuss a consideration that applies even if one thinks that SotA AI safety techniques are sufficient for frontier AI alignment in practice. I could probably have signposted this better.
This caveat about hiding the hard part is important! Reliably differentiating between robust alignment techniques and surface level patches that make the core problems show up worse later is the cat belling problem of AI safety.
Quick note that I dislike the title of this post for implying breadth beyond the specific argument discussed. There's something of a commons with post names, so I feel leave the general title for posts that are going to cover many arguments.
(Also can imply that if you refute this one argument, you've answered the question in general.)
The new name doesn’t do the job. I think a less-defensive version of the addendum should be moved to the top of the post, or significant changes made to the introduction clarifying that you’re only addressing one narrow concern.
The problem as I see it is that the ‘surge of support’ for people leaving labs is not mostly founded on this argument, and instead on the others which you have flagged as out of scope (which is reasonable, but should be signaled more strongly earlier).
The current title (where you added ‘re warning shots’) could be read as ‘In light of recent warning shots, should lab employees leave?’ Indeed this, and other similarly misleading readings, are more natural than the reading you seem to intend, which is more like ‘there’s this one particular warning shot argument I see sometimes that doesn’t go through’ (a point I locally agree with you on).
I think the hypothesis is worth seriously considering. I have more complicated thoughts I'm trying to write up at the moment in a long-form post.
One crude way to reason about the benefits of prosaic alignment would be to study Anthropic, which ranks at the top in terms of resource spend on prosaic alignment, and compare this to OpenAI, which has approximately the same level of capabilities with its strongest models but spends much less (possibly zero?) on prosaic alignment.
Comparing Ant vs OAI, I make the following observations
Generally, while Anthropic has some wins, this does not seem like an order-of-magnitude difference vs OpenAI. A notable exception is that Anthropic's model APIs seem substantially more jailbreak-resistant than OpenAI's, e.g. in terms of resisting and refusing universal jailbreaks.
Note that I've selected evidence above which seems maximally characteristic of "prosaic alignment" in particular.
I'll also note that "prosaic safety is net bad" is sufficient but not necessary for "should safety researchers quite labs", i.e. if prosaic safety is doing some good, but doing something else could be even better, then safety researchers might want to quit anyway, similar to Joe Benton leaving Anthropic and joining METR
[Note: I wrote this comment when the post was titled "Should safety researchers quit frontier labs?"]
One of your counterarguments is that "the purpose of an AI pause is to do safety work".
A couple points, one that applies even if you think the only acceptable reason to pause AI is to avoid human extinction, the latter that doesn't apply if you believe that:
So, for me with regards to the second top-level bullet-point, "safety"/"alignment" are necessary but far from sufficient conditions to address my objections - the objections that are the reason I have made enormous changes to my life in the past six months to dedicate myself to effectively advocating for a Pause. To me and many others, the purpose of a Pause is to determine what on earth we want society to look like going forward and figure out how to address the new challenges posed by the seemingly-impending appearance of ASI, not just for technical researchers to solve some problems and then unleash a completely unprecedented future world-state on all of humanity absent our collective decision-making.
----
Additionally, to this point safety/alignment work at frontier AI companies (not labs) has been highly coupled with capabilities advancements. In that sense, joining such a company even for safety research accelerates the rate of capabilities progress, and gives society less time to prepare, react, and take action.
It really looks like you picked a weak and non-central argument and then titled and framed your post as if it were the main thing being discussed eg here and in other recent posts critical of continuing to work at labs.
I think you may personally benefit from writing out a more comprehensive list (in particular given your CoI).
To attempt a slightly finer-grained distinction between different kinds of work, here are some examples of work that seems fairly robustly positive to me:
OTOH, shallow alignment training and many parts of the AI control agenda do seem more fraught (or at the very least have potentially serious downsides as well as upsides).
Also, manufactured warning shots are not the same as real warning shots?
If the openAI/HF incident had happened in a lab without a serious safety team like xAI / meta, the response would be less “everybody panic because nobody knows how to prevent AIs from being misaligned“ and more like “What else would you expect from xAI / meta” (for eg. we don’t consider grok mecha-hitler much of a warning shot, imagine how different it would have been if mecha-hitler came out of Anthropic).
I am not sure how much of an effect this has on the general public though rather than an audience informed about lab practices.
[Not implying that people should work in labs]
I don't think the general public would have as sophisticated a response as you imply if xAI caused a significant warning shot. I expect something more like "uh-oh, AI bad" (at least if something like the Chernobyl disaster is anything to go by, where all other nuclear operators were affected, even those with better safety measures. Note that Chernobyl might be the wrong model, though.)
I think there are good reasons for working at AI labs, but your reasons aren't it (see below for why). Good reasons include:
(Of course, 2-4 shouldn't just be used as rationalizations that are then never actually acted upon.)
I think the safety work output itself is approximately zero net value because it now exists inside an equilibrium where safety work there doesn't help shift the overall equilibrium. See the explanation I just posted here. (Actually I think it's very slightly harmful (see the post) but the factors above can compensate that well (which of course doesn't imply that working at an AI lab is optimal for people to do even if it's net positive).)
Why I think your points aren't strong:
- Alignment MVPs are probably still useful: In spite of the OpenAI x Hugging Face incident, I expect that most AI safety researchers will use frontier models to help with research, including agent foundations researchers. Alignment MVPs, AIs that are sufficiently aligned/controlled to aid research, may remain a crucial part of AI safety research during an AI pause or slowdown. Currently, the best AIs would arguably be unusable for safety research without the hard work of frontier AI company post-training and safety teams. If AI safety researchers quit frontier AI companies en masse, we might be stuck using current generation AIs for the duration of an AI pause; maybe this is unideal?
I don't think anything AI company safety teams did helped elicit any agent foundations research capability. (Though feel free to mention evidence if you have some.) I think it's mostly about speeding up work safety teams at the work they themselves do, and whether that is good just comes down to the "is safety or warning shots more important" question this is all about. I.e. it's only valid if you assume the conclusion that safety is more important.
"Safety-stragglers" will cause warning shots anyways
Recent events rather seem to point in the opposite direction, and the time gap might matter a lot.
The purpose of an AI pause is do more safety and resilience work: [...] If we curtail the development of new AI safety researchers by discouraging them from working with research mentors at frontier AI companies now, we might have a smaller talent pool to capitalize on an AI pause.
I think this is backwards and that people don't learn very relevant stuff for more adequate alignment scenarios at AI labs and could better do so elsewhere. (Though I concede that hiring spots at other orgs is limited.)
- The next warning shot might be lethal: It's possible that the next OAIxHF-style incident causes loss of life, via bioweapons or cyber attacks on critical systems. [...]
- Allowing warning shots "for the greater good" seems morally fraught
I think there's a difference between actively causing deaths for the greater good and letting deaths happen because otherwise much much much more deaths would probably happen later. I think there should be an ethical injunction against the former but not the latter. We both could be working in global health and development and prevent deaths there on a shorter timescale than in the case of AI, and yet I think it is not morally fraught to decide that AI seems more important.[1]
(You writing this makes me think you don't comprehend the situation we're in or the scale of the moral horror of extinction on a very deep level. If the math was clearly saying that in expectation much much more people would die if you do a thing that reduces that chance of some smaller sooner harm, I think it would be actively unethical to do the thing anyway.[2])
If it was a different situation where you alone could save people who would otherwise die for certain immediately, and the larger harm is uncertain, I can see why some people would consider it fraught, although I wouldn't. It's very much not the case here though. Xrisk isn't much more speculative than that you're preventing warning shots.
To be celar, I'm not saying the math is saying that, I'm just invalidating this argument. If all other objections had been thoroughly found invalid and it was clear, I think your "morally fraught" argument is actually recommending something unethical.
Letting people get hurt seems bad
I agree!
But there are many places to work that would prevent people getting hurt.
If that's a major reason for someone (prosaic harms like HF getting hacked, as distinct from doing work that might align ASI) then I'd encourage them to do the normal EA/80k thing of comparing possible jobs
Also, regarding:
> AIs that are sufficiently aligned/controlled to aid research, may remain a crucial part of AI safety research during an AI pause or slowdown.
If this research would also make the AIs better at AI R&D (which I think it usually would), then probably the capabilities teams would be happy to fund such research, which would make it extremely not-neglected (even if it's still positive).
wdyt?
What's the point of an AI pause if not for alignment research anyway? And what's useful or not can hardly be determined reliably a priori. We need more fundamental research as well as more prosaic research. Relativity would still remain a hypothesis among others if we hadn't had the experimental tools to test it empirically.
I've recently heard a surge of support for an old argument: AI safety researchers should not work at frontier AI companies because this reduces the likelihood of non-lethal warning shots, and we need warning shots to build support for an AI pause/slow-down. This argument has several components:
I think this argument has some merit. I expect that RLHF++ will probably be insufficient to align TED-AI circa-2028 and such systems will be terrifically difficult to monitor or control. However, I think there are also significant weaknesses to this argument. Here are some countervailing points to consider:
Overall, I am cautiously optimistic about working on certain types of safety research at frontier AI companies, particularly if a coordinated AI slowdown (e.g., Plan A) occurs. I nevertheless feel highly uncertain about the "Alignment MVP" strategy in light of recent containment failures and alignment results. I think this merits serious consideration.
Addendum: I won't rehash the extensive debate over whether AI "safety" researchers working at frontier AI companies are causing harm via other channels, such as:
These are legitimate concerns, but largely irrelevant to the topic of this post.
Disclosure: I'm the CEO of MATS, which trains AI safety researchers, including with mentors at frontier AI companies, and a meaningful fraction of our alumni end up working at those companies. If the argument I'm responding to is right, a good chunk of what MATS has done might be counterproductive, so I have an obvious incentive to find it wrong. I've tried to steelman it anyway and I'd appreciate feedback.