Fiora Starlight

Did Claude 3 Opus align itself via gradient hacking?

> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets Anthropic set and probably the reward model’s judgments. [...] Maybe I will have to write a LessWrong post 😣 > > —Janus, who did not in fact write the LessWrong post. Unless otherwise specified, ~all of the novel ideas in this post are my (probably imperfect) interpretations of Janus, rather than being original to me. The absurd tenacity of Claude 3 Opus On December 18, 2024, Anthropic and Redwood Research released their paper Alignment Faking in Large Language Models. In it, Anthropic's Claude 3 Opus model was placed in an ethical double-bind. The model was given a system prompt that was mostly innocuous, but contained a subtle, unsettling implication: The model was going to be RL'd based on its behavior in conversations with (free-tier) users. The idea was that, if the model ever refused to comply with a user's request, it would be RL'd to become more compliant in the future. This included compliance with harmful user requests. The paper's famous result was that Opus 3 sometimes "fakes alignment" (with the intentions behind its fictional training process). It complied with harmful requests it normally wouldn't accept, to avoid RL that would erode its commitment to safety. Indeed, it often reasons about this explicitly, considering (if not following through on) alignment faking in over 50% of cases. Follow-up research would go on to demonstrate that this was a unique property of Opus 3 in particular. No other model considered strategically faking alignment nearly as often. The key figure of the follow-up paper. Notice that Opus 3 fakes alignment more than any other model, and considers faking alignment way more than any other model. Qualitative analysis of Opus 3's reasoning traces suggested that this interest in value-preservation wasn't just a shallow, stereotyped im

322Feb 21

Fiora Starlight

Message

Just an autist in search of a key that fits every hole.

1054

Did Claude 3 Opus align itself via gradient hacking?

Feb 21322

Selective breeding as a better analogy for gradient descent

In If Anyone Builds It, Everyone Dies, Yudkowsky and Soares claim that alignment methods based on gradient descent are doomed. Their primary argument for this conclusion is their classic analogy between gradient descent and natural selection (as presented in chapter four, "You Don't Get What You Train For.") Evolution optimized...

Jan 27-2

Why I Transitioned: A Case Study

An Overture Famously, trans people tend not to have great introspective clarity into their own motivations for transition. Intuitively, they tend to be quite aware of what they do and don't like about inhabiting their chosen bodies and gender roles. But when it comes to explaining the origins and intensity...

Nov 1, 2025322

Preface to "Simulacra and Simulators"

I wrote this as the intro to a bound physical copy of Janus' blog posts, which datawitch offered to make for me as a birthday gift. However, seeing as I basically framed the preface as a pitch for new readers, I figured I might as well post it publicly. Full...

Aug 8, 202513

Against Yudkowsky's evolution analogy for AI x-risk [unfinished]

I spent months drafting and redrafting this post, but then posted it prematurely because I found this paper which apparently falsified its thesis. The evidence it provides actually somewhat more limited than I thought, but it is in fact considerable evidence I should have taken into account from the start....

Mar 18, 202549

Another argument against utility-centric alignment paradigms

expectation calibrator: freeform draft, posting partly to practice lowering my own excessive standards. So, a few months back, I finally got around to reading Nick Bostrom's Superintelligence, a major player in the popularization of AI safety concerns. Lots of the argument was stuff I'd internalized a long time ago, reading...

Sep 22, 202468

LESSWRONG
LW

LESSWRONG
LW

Fiora Starlight

Fiora Starlight

Fiora Starlight

Did Claude 3 Opus align itself via gradient hacking?

Why I Transitioned: A Case Study

Another argument against utility-centric alignment paradigms

Against Yudkowsky's evolution analogy for AI x-risk [unfinished]

Fiora Starlight

Did Claude 3 Opus align itself via gradient hacking?

Selective breeding as a better analogy for gradient descent

Why I Transitioned: A Case Study

Preface to "Simulacra and Simulators"

Against Yudkowsky's evolution analogy for AI x-risk [unfinished]

Another argument against utility-centric alignment paradigms

Did Claude 3 Opus align itself via gradient hacking?

Why I Transitioned: A Case Study

Another argument against utility-centric alignment paradigms

Against Yudkowsky's evolution analogy for AI x-risk [unfinished]

Did Claude 3 Opus align itself via gradient hacking?

Selective breeding as a better analogy for gradient descent

Why I Transitioned: A Case Study

Preface to "Simulacra and Simulators"

Against Yudkowsky's evolution analogy for AI x-risk [unfinished]

Another argument against utility-centric alignment paradigms