Yep, called it. Thanks for citing the history. To me Kabir's strategy looked obviously correct just on left-wing instinct. In fact the opposite of quitting-in-protest, entryism (getting hired with an explicit goal to hijack the organization), is a known and effective tactic. The world would be a better place if alignment folks had that mindset when joining labs, instead of seeing themselves as helping the lab.
Tl;dr
I saw Kabir's post on LessWrong and thought the argument could be strengthened by a framework from the pro-democracy field, which sorts defections into breaking (leave, visibly and publicly) and binding (stay and work from the inside), each further divided divided into speaking, acting, and standing in the way. The best data I know of on what moves institutions during democratic backsliding (Pinckney & Trilling, 2024) says:
Below I'll attempt to translate this to the frontier labs, with the obvious caveat that a lab is not a regime. The short version: refusing the specific work while staying is the higher-success action in their data, and staying in the institution to quietly move colleagues is closer to what works than either resigning or striking.
Background
I've spent 15 years in political campaigns and organizing, with training and experience in applying the lessons of pro-democracy movements abroad to the US. For the past year, I have attempted to scope a few projects that could be run in the US (one of them is the project I now help lead).
I went deep on the concept of "defections," since Pinckney and Trilling's paper shows they are one of the most effective tactics in pro-democracy organizing. The concepts of breaking and binding were coined by Adam Fefer in his August 2025 guide Shifting Pillar Loyalties.
Now, I help lead an election defense project called Hold the Line that equips people with evidence-based, nonpartisan frameworks to help them understand threats to democratic institutions; organize effective collective action; and defend free, fair, and safe elections. It enables individuals to organize democracy defense teams in their own communities, wherever they are.
Definitions
Fefer defines breaking as visibly and publicly removing yourself from the group or institution you were part of, and binding as when you try to change it from the inside, through back-channel outreach, subtle persuasion, or by preventing someone from taking your place. Both matter, and the choice depends on where you are in the institution.
The guide also sorts the actions available to you into speaking (publicly or privately), acting (protest, organizing, lawsuits), and standing in the way (noncooperation, refusing a task, work stoppages).
Here's a way to visualize it:
Speaking
Acting
Standing in the way
Breaking
Resign with a public letter
Leave and organize, litigate
Leave and take the team with you
Binding
Dissent internally on the record
Organize colleagues, build a back channel
Refuse a specific task, don't let a "yes person" take your seat
Fefer's guide has a table of real world cases from anti-authoritarian movements, including here in the US and in Poland, that may be useful to lab employees.
If we were to analyze recent frontier lab employee actions, Jacob Coxon's resignation is breaking + speaking. Kabir's proposal to refuse the work and let them fire you is binding + standing in the way.
Digging deeper into the data
What was most surprising to me was Pinckney and Trilling's finding on quiet outreach.
When a campaign's main strategy toward a pillar of power was quiet outreach, that pillar moved toward democracy about 39% of the time. When the main strategy was pressure (physical or verbal protest), campaigns did worse than cases with no campaign at all.
On the success of what actors in those pillars did, noncooperation had the highest success rate (67%) vs. verbal protest (33%). The authors find that when actors in the pillars "merely speak out, there is minimal impact." Success is more likely when they use their position to stop cooperation and directly challenge the source of backsliding. The clear caveat is that loyalty shifts are rare and it's a small sample size and they suggest correlation only.
What may be instructive to a frontier lab employee
I think using the table above could be instructive to frontier lab employees. For instance, what made Coxon's departure worth more than a resignation is that it pulled a public statement from Evan Hubinger, who is still on the inside.
On binding, if an employee is asked to do something unethical, they could do the following: ask for orders in writing, run it past legal, don't do work outside of their contract, be unavailable if something is truly unethical or immoral, document everything, and refuse openly at the point they are called on to make a decision. habryka is right that you very well could be sidelined and the lab will wait for you to make a mistake. (To be clear: Nothing here is legal advice, and if any of this could touch your employment, talk to a lawyer first.)
My opinion is that there could be a strategy where ex-lab employees organize into affinity or support groups (similar to what we saw after DOGE eliminated most of USAID's staff) that then organize their connections at the frontier labs. This would likely be the strongest method to start doing relational organizing. I'm not sure if anyone is working on this or it's been proposed, but I think it could be an interesting mandate for someone at an AI safety organization or for an ex-lab employee to spin up a project of their own.
For the AI safety community, I think this research could be very useful in thinking through the benefits of relational organizing (quiet outreach) vs. protest and putting pressure on the labs from the outside. The strategy with the best record for shifting insider loyalty was the somewhat boring relational work.
I'd be really curious to hear from any lab employees on if this framework is useful and how likely they'd be to work with an outside entity to enact some of this strategy if it were well organized and structured.
AI use: I used AI to pull data from papers I previously read as well as to tighten up the structure of this post. As this is my first post, I wanted to write in the structure of LessWrong, which was unfamiliar to me. Mistakes are mine alone. Eager for feedback!
Sources