"They fired him because he refused to help AI capabilities"
"They fired him because he didn't want to work on bad policies"
Ok, but they probably just won't fire you. The labs have infinite money, they can just sideline you in the organization, give you some sinecure position, while waiting on you to make a mistake that allows them to fire you right after you made some gaffe that actually makes that seem reasonable. Choosing the context for your own exit is a powerful tool!
Today there was an example where an ordinary research engineer, Jacob Coxon, quit and tweeted that he quit, and that alone apparently warranted an article in the WSJ, which I would not have expected. See https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-out-of-control-ai-fears-707b7628.
This suggests that quitting for moral reasons seems like unusually important new information to people outside the AI bubble. It seems to me that probably means that it's more useful than communicating similar info only to coworkers by taking a principled approach inside.
Nit, but I would clarify: his quitting and tweeting about it didn't happen to warrant an article in the WSJ. Articles like that have an associated lead-time, and so he first had to find an interested journalist who would write the story, then time his quitting/tweeting to their release of the story
That is: I would not expect another employee quitting spontaneously to get an article like this, without having done a media engagement strategy ahead of time. If someone is going to quit, please try to line up a good media opportunity to pair it with, not just shoot from the hip! An exclusive like that is juicy; but merely covering it once it's already happened (and anyone could write the same story about it) is much less enticing
Coxon's resignation also led to Evan Hubinger, an 'Alignment Science lead' at Anthropic [I'm not sure if he's an alignment science lead or the alignment science lead, to be clear] to say:
"Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." (source)
This is very useful information - it's important and relevant for the world to know that frontier AI employees in such a position believe this sort of thing! And Anthropic is notably secretive compared to the leading AI companies in terms of how much we know about what their employees think. So Coxon's departure was valuable not just for its own sake (and direct coverage thereof), but for eliciting additional info from other employees.
I feel like the main reason to quit is because you think you can have more impact elsewhere (or can have a bunch of impact in the act of quitting). My guess is that in most cases, someone refusing to work on capabilities work would not be fired, but might be somewhat sidelined, wouldn't have that much influence, and be somewhat socially uncomfortable. The company is unlikely to make someone a safety martyr by firing them if they can just have a less useful employee around instead (and if said employee is doing safety work, that's typically pretty good from the company's perspective)
That plan seems more reasonable! Being a current employee of an AGI lab will make protesting and other comms things notably more effective. I do expect loudly quitting in protest to be even more effective though (see eg Jacob Coxon today)
I think "stop doing harmful work until you are fired, instead of quitting" is fairly compelling.
But, I do think most people working on prosaic safety are mostly doing harm.
If you don't have a particular story for how what you're doing scales to superalignment*, I a) don't think you're solving particularly important problems in the nearterm, b) in the medium term, mostly making it marginally faster to roll out stronger capabilities, which is bad because it's burning calendar time for serial research time, and solving the legible problems leaving the illegible ones.
I am somewhat dismayed to see how highly-upvoted this post is; I think I have literally never seen a post this poorly-reasoned get 200 karma (before my strong downvote). It reads like a grab-bag of arguments hastily thrown at a strong pre-written bottom line.
To be clear, I'm not saying that Kabir shouldn't have posted this, I generally am very pro people quickly writing up their takes (though I do think the title is unusually bad). Rather, I'm mainly trying to figure out why the community reaction was so positive.
I might engage more on the object level later but for now my default hypothesis is that people are seeing AI safety people get in public loud conflict with the labs by quitting, are having a knee-jerk conflict-averse reaction, and instinctively supporting anything that might defuse that immediate conflict.
In general we see this pattern a lot with the labs: people do something that is somewhat critical or disagreeable, and then everyone quickly generates a morass of FUD and reasons why, yes, they support people taking action in principle, it's just that this particular action is a terrible idea that will lead to so much backlash compared with all the other possible actions....
The problem here is that creating common knowledge that people are following a given set of norms (and therefore should be treated as coalition members in good standing) becomes far more complicated the more complicated the strategy is.
Leave the company: very simple, easily verifiable, everyone knows (that everyone knows) what happened.
Staying but not working, and doing internal agitation, and donating one's salary: complicated plan, many parts are hard to externally verify, easy to rationalize.
I don't find this convincing.
Tons of people who left AI labs have left our world better off because of it (you can disagree on some, but the overall picture is there). Ex-lab employees have an unfair advantage in founding new organizations (research non-profits) because they have the credibility and status to secure funding. Many high-profile AI safety non-profits were founded by ex-lab employees!
Examples:
If it's so costly for labs to fire employees for not doing x work, then they'll likely just not do it and, as Habryka said, wait until a firing looks bad on the employee, not on them.
In addition, their outside takes can be treated as more trustworthy (in general) and generate more momentum to pause.
Lastly, their inside knowledge of how frontier labs work can be valuable at many external organizati...
I upvoted this, mainly on the strength of the suggestion of the first paragraph. That's maybe a better strategy, and I want to promote it to attention.
I don't endorse it as a recommendation, because I just don't know, man.
I think that in the typical case working for an AGI company is probably evil (though it's complicated and so I might just turn out to be wrong about that, given the hindsight of the future).
It seems pretty surprising if people broadly shouldn't go to work at AGI companies in the first place, but that if such people happened to have already done that, and suddenly conclude that what the company is doing is bad, that they should then do something other than leave?
Or are they now, having already worked there for some years, in a different situation because they have credibility within the company?
Does that imply that people should (if they have the psychological resilience to actually follow through, which ANACT approximately no one does), they should start working at a company, do what they're told for n months to build credibility , and then when their n-month timer rings, switch to refusing to do work that they think is harmful?
Or is the implication that someon...
when i first joined openai, on day one i was already committed to not working on anything harmful (including capabilities). for various reasons, i think i had an unusual experience (openai culture was different back then, my first manager was extremely chill, i already had a reputation in ML, etc). but it's a useful data point.
Saw the title and was about to disagree but
Just refuse to work on bad things and see if they fire you.
Yeah this is an underrated strategy. Refusing to work on bad things likely has an outsized effect on internal company culture.
My guess is that it would not be obvious for someone within the system which work tasks are "bad things." It's not obvious from the outside and my guess is that it'd be unobvious from the inside either.
Yeah. I think this is good advice (though anyone following it should prepare themselves for a few stressful months: being pressured by management, becoming a rumor target, gaslit, accomplishments minimized etc). And your algorithmic liability plans are also really good. One of the few "rays of light" lately.
For years I've thought people should consider the strategy of joining labs to only work on things-good-for-the-world, and people at labs staying to work on those things also seems promising. But I think I mostly don't buy this at present margins. Sodom and Gomorrah aren't destroyed until Lot leaves, and every respectable person who sticks around not committing any crimes themselves is acting as cover for the organization that is happily committing crimes. Resign today.
Probably-unhelpful side-comment:
Going off the title, I expected that this post was going to be about imploring God not to abandon this universe despite what people are doing in it.
This post offers only a handful of vague category labels — "bad things," "bad policies," "AI capabilities," "morally wrong" — but no definition or operationalization of any of those terms and no example of a task that straddles the border between "morally acceptable" and "morally wrong".
The considerations in this post are unimportant IMHO compared to the important task of choosing (as a community) and then promoting / advertising a standard of conduct, then discussing tactics and strategies for pressuring researchers, engineers, entrepreneurs, investors, e...
Another alternative strategy to just quitting: become the point-person for organizing a union. Employees won't necessarily care about safety, but having another locus of power in the company besides just shareholders could substantially increase leverage along values other than short-term profit maximizing. And if you get fired for organizing, that seems like it would net extra martyr points.
People have different expectations on how this'd go → perhaps some people who'd otherwise quit should try variations of protesting without quitting so that we can be more informed
Imagine for a second that one of the people on your team said "Mateusz, everyone else on the team, I'm not going to work on this because I think it's morally wrong".
Isn't it kind of important for the impact of this statement that the "on this" is also bolded?
If you say "I'm not going to work", you're in fact being very annoying, and not doing your job, and I sympathize with your boss for firing you for not doing your job.
If you say "I'm not going to work on this", then you are correspondingly more reasonable?
I don't think this works so cleanly for most safety positions? Usually what you are being asked to do (for the pay you are presumably accepting if you haven't quit) is not something obviously evil that will make a clean headline, or something the company will conveniently announce to the media.
If you work for OpenAI, in an AI safety position, you are probably being tasked to do "safety work" even if it is only prosaic safety work. If you refuse to work on the general principals of how the company is working, they probably wait until a convenient quarter, and fire you for refusing to work, and even if you tell people it was because of the general direction of the company, they can justifiably say you were simply refusing to work on anything and I don't think you will have as much impact as quitting at an opportune moment that actually signals that it was specific choices or outcomes that triggered your quitting.
The experience of covid vaccine refusers was that forcing their employer to fire them or their apartment to evict them was a much more effective form of resistance than quitting or moving on their own.
Also, if you're an expert and you're in a job with a lot of discretion, when tasked to work on something you disagree with, just nod your head and go back to doing what you want to do.
This quitting was one of the probably 1000 largest tweets of all time with huge cut through. Seems like you are wrong.
I am now less certain about my comment that sparked this post (especially about the "fire you after 1 or 2 months" part), but I remain rather unconvinced about the cluster of strategies you're suggesting here.
I would update more strongly if I saw side-by-side-ish examples of something like what you're proposing vs quitting loudly from harmful industries (tobacco? lead gasoline? social media clusterfuckery?). Admittedly, they might be difficult to find.
I think the best reasons to stay are (1) you can do good/beneficial work on the inside; and (2) you can try to influence the culture from the inside. (1) can work for some people, but it seems to me like they are — god bless them — an anomaly. I don't expect (2) to reliably have a strong enough effect.
There's also the issue of working in this sort of environment warping your epistemics a lot, for a combination of reasons, which gives you an additional reason to quit in order to reason more clearly about what would be good for you to do. Again, some people are fairly immune to this, but, again, I expect them to be an anomaly, and most people underestimate the extent to which social reality warps their cognition. @Richard_Ngo had a twee...
Perhaps even more impactful (but much riskier!) than loudly refusing to work: quietly undermining the bad work.
If you are trusted with important projects that you disagree strongly with, that seems like a unique opportunity to gum up the works à la Simple Sabotage Field Manual: slow-walk progress while getting people to believe a breakthrough is just around the corner, pick less competent team members for the more difficult parts of the project, add highly burdensome bureaucracy in response to minor mishaps...
Full disclosure, I don't think I would have the chutzpah to attempt this. You'd need to make sure none of your acts of resistance is flashy enough to get headlines or get you sued if found out. And even so, it could require taking your motivations to the grave lest you tarnish the cause. But (alongside WillPetillo's fascinating suggestion to try to organize a union) it's another option on the table.
I actually read the title of this post as directed to The God and urging it to not resign, as a maybe philosophical position that existence of everything that currenly exists is better than non existence, or maybe some incremental improvements are better or something.
My strong sense is that almost all tasks at OAI/Anthropic are bad to work on.
That would mean essentially getting fired for refusing to do work at a company that is paying you comical amounts of money. Will this in fact have more leverage/cachet/newsworthiness than someone like Kokotajlo or Ngo if/when it comes out that you were basically sandbagging?
[Edit: I guess I was thinking more about people newly joining OAI/Anthropic intending to "do good" there. I've realized my analysis doesn't account for people who are already there, especially those with long tenures.]
Just refuse to work on bad things and see if they fire you. There's not much time left for resumes to matter. Also, firing someone for refusing to work on bad things is actually very costly.
If it's not obvious that you're working on bad things, but the company is obviously doing broadly bad things, consistently, just refuse to work at all until those change. E.g. if at Anthropic or OpenAI, refuse to work until training runs stop
Some, such as Mateusz may say: "They would fire you after a month or two and the firing wouldn't have the same social effect as voluntary quitting of, say, Daniel Kokotajlo or Richard Ngo."
I understand why it may feel that way, but I disagree very strongly, I predict it would have much more of a social effect.
"They fired him because he refused to help AI capabilities"
"They fired him because he didn't want to work on bad policies"
etc, much bigger headlines.
Also, I think you may not be factoring in the extent to which there is a cost to the company executives to be seen as firing someone. Especially someone who is refusing to work on moral grounds and has already proven themselves to be high status, respected, etc.
And especially how it would look to the other employees if they refused to even listen to the striking employee before firing them or refused to even negotiate at all.
The company leadership try to present themselves as very thoughtful, sincere, doing their best, etc. This is a large part of why many of the most talented people are there.
Imagine for a second that one of the people on your team said "Mateusz, everyone else on the team, I'm not going to work on this because I think it's morally wrong".
And then you fire them.
What would the other teammates think? What would they think of you? What would happen to their trust in you? What would happen to the trust that investors have in you? Or funders?
Obviously bad things. So obvious that you wouldn't even do it. You'd talk with them instead, see what could be changed, so that they feel better and this is less of a risk. So something would actually change in the company, to better reflect the employee's values, or you suffer a huge blow to your image and your employees trust in you.
And this would then mean that the other teammates know they can also do this.
Some may ask "would this give teammates courage or 'set an example'?". Even if it 'sets an example', it will do so in the short term, until the next moral outcry comes, where people get some courage again. And, more critically, they would much less believe in the nice things that the leadership say, the moral image they build.
And even better than quitting - there's a much, much lower chance of the protesting employee, who has moral courage, being replaced by one who has less moral courage.
Also, this is a much, much, much costlier signal. It's much harder to get funding to start a new lab, join Anthropic, etc, if you're known to be someone who might just stop working one day because you don't believe in it.
And this matters. People see that. Your colleagues see that. Journalists see that. Regular people see that.
Think of Daniel Kokotajlo refusing to sign the NDA and losing out on a lot of money because of it. It was a very very costly signal. And it's mattered and been trusted more because of that.
Refusing to work and taking on expensive cost and risk every day that you do so, is a much bigger cost and signal.