I think I'd need more information from the individual before claiming they are doing evil actions.
This researcher may believe Anthropic has the highest ethical standards out of all the frontier labs, and therefore believes Anthropic should aim to be ahead in terms of capabilities, as they will be in a better position to set AI safety standards.
From my experience, this seems to be how most researchers think.
This is of course a view that a person could hold, but then (granting that the post accurately accounts the conversation) the researcher could have responded with that. They would then be able to debate the merits of this point, or even to opt not to. But not having any response at all is a good bit worse.
Can't we press way harder on this? You can set up a table outside their office and just talk to whoever "has doubts" you want. Or ask the resigned researcher today who he knows who is most wavering. Take names, say you'll call them whenever an ASI grant opens up that poaches at Anthropic pay parity (which is a lot but if you get half the value back from poaching frontier.)
If you do something kinda bold (cold canvassing or cold calling) you definitely want to be thoughtful and don't act alone. If you're linkedin friends with an Anthropic from a few years ago you haven't talked to in a while, sure, DM them, ask how they feel, see how it goes.
I agree with your last point and have written about it myself. My basic view is: limit your use of frontier systems as much as possible; try to keep your paid AI spending directed to companies not at the frontier; offset your paid AI spending with equal (or larger) donations to groups working toward a pause/stop.
Link: https://connorsscratchpad.substack.com/p/ai-safety-advocates-its-time-to-be
one counter is that Anthropic is creating products that people pay for and find valuable. Or, [insert another tepid argument that I can’t remember]”. I was surprised that this is the first thing they thought of, rather than talking about the potential gigantic upsides of advanced AI systems.
this is not surprising to me. if the upsides are a real possibility, then so are the downsides, and it's very hard to make the math work out if the downsides are under consideration. for capabilities work to be justifiable, 'ai' must be a perfectly ordinary technology with no (or only perfectly ordinary) externalities. market signals are the standard way to evaluate whether perfectly ordinary activities are prosocial / worth doing.
(edit: just to be clear, i'm not endorsing the above motivated reasoning. just trying to explain why 'potential giant upside' is not actually reachable for someone trying to justify their accelerationist salary.)
This reads like an r/thathappened post.
“So why are you working at Anthropic then?!”
And they didnt have an easy, conditioned rebuttle/response? Researchers at Anthropic, who are open about having AI ethics discussions at a party, haven't formulated a quick thought on this subject?
Not surprising from someone who clearly articulates it is evil to even work at a frontier AI lab. Apologies we can't all have your enlightened philosophy.
I was at a house party hosted by an AI Safety friend of mine. I join a conversation midway, where my friend is saying that doing capabilities research at frontier lab is evil, given the catastrophic risks. Nothing out of the ordinary, until I find out the person who they are talking to is a capabilities researcher at Anthropic!
I was shocked at the directness of my friend. The researcher took it well though, partly based on their temperament, and partly because they have been exposed to several AI Safety spaces previously.
Unfortunately, at this point, my memory of the conversations is extremely shaky. Also, the conversation was not continuous and happened piecemeal between other ongoing discussions. But here is my best recollection.
The researcher responds with something like: “Not a good use of our time to go flesh out our positions, as we will just re-hash the standard arguments. For example, one counter is that Anthropic is creating products that people pay for and find valuable. Or, [insert another tepid argument that I can’t remember]”. I was surprised that this is the first thing they thought of, rather than talking about the potential gigantic upsides of advanced AI systems. As much as I lean libertarian, justifying catastrophic risks with ‘customers pay for our products’ was pretty weak, at least with the particular way they phrased it.
I have some 1-1 discussion with the researcher, and sympathise with them being called evil. I bring up vague idea I have had that researchers in frontier labs should have some kind of public statement about personal redlines: what things would AI’s or Anthropic or Anthropic leadership do that would cause them to leave the company.
They respond saying that this would likely just be a checkbox exercise, with people copying and pasting some standard meaningless statement which has no teeth or consequences.
I say that instead, maybe people should just post a statement every six months along the lines of “I have reflected on the risks and benefits of working at [frontier lab], and have decided to [stay/leave]”. Sure, again, they could just copy and paste this and use it as a checkbox exercise, but I sense that most people would not post such a statement if they had not actually done a reflection, whereas people might honestly post a statement about their personal red lines, and then change their minds about it later.
After some discussion, the idea morphed to: every six months, they have a discussion with somebody – e.g. me – where the aim is to help them explore their own personal views. They disliked framing, saying it would be adversarial and feel like an inquisition with lots of AI safety people around challenging them. I said the discussion would be 1-1, would be done on whatever terms the other person wanted, they could leave the discussion when they wanted, and that my personal style is to just try to understand the other person, rather than to explicitly change their mind.
They then started asking a question as a counter, “What would you think about having a discussion every six months about big things in your…”. They did not finish the question, because the answer was evident to both of us. Yes, that would be useful! Of course it would.
At this stage that they had no good reason not to do this exercise, which of course is separate from them actually doing it. The final question I ask – in an attempt to identify an underlying crux between us – was, “Do you think it is possible, in the next 10 or 20 or 40 years, for an AI system to be created that causes human extinction?” They responded immediately, “Yes, of course.”
I can only assume that the expression on my face was asking: “So why are you working at Anthropic then?!”
It was interesting to see the cogs turning in their head head realtime: I am confident that there is a significant part of them that believes they should not work at Anthropic, and that the other parts of them are desperately trying and failing to come up a coherent reason not to think about the issue.
As we part ways, I say that the offer to have this discussion is real, and they are welcome to stay in touch. I sense this is unlikely to happen, with the fact we did not share contact details being the least relevant reason.
Final thoughts. I am curious to know what you think.