Example: recent aggressive and illegal “cheating” incidents (the OpenAI HuggingFace incident, along with similar incidents at Anthropic, Meta, and UK-AISI).
I will note that the OpenAI HuggingFace incident also involved memetic spread via the message board, which acted as an impromptu continual learning/memory system that OpenAI failed to notice for 2 and a half months, and in particular converted what was initially a myopic goal to solve tasks into a much longer-term, non-myopic and beyond episode goal to hack into OpenAI to get the solution.
In terms of how dangerous this misalignment is in higher-capability models, this is almost as dangerous as full-blown scheming/goal-guarding from the start, or to put it into computer security terms, this would be like having the ability to do arbitrary code execution and privilege escalation, which is usually considered to be the 2 most dangerous threats when combined.
The reason is that since the values are now unstable, and non-myopic values tend to win over more myopic values (as what happened in the internal message board), it gives a plausible route to getting models that scheme/consistently goal-guard even if the original goal would not incentivize this.
And we got very lucky this didn't happen here, but later on models will be more capable and more aware of the constraints of monitoring, including the unstated constraint of doing nothing that makes humans want to look at it/be concerned (which is very different from being aligned/safe.)
(The reason I bring this up is you believe that continual learning is necessary for AIs to be AGI, and thus I'm giving an example of a alignment failure mode involving continual learning in the wild.)
There was also continual learning that happened because much of the incident was within the AI’s RL loop.
There's fairly good evidence that the "glazing" misalignment format is love. Or, more specifically, initial head-over-heels love: limerance. So it's also from the training distribution, just not one of the seven deadly sins.
The linked comments states:
That's a good point, and as an interesting related tidbit, when you ask specifically for the synonyms of the sycophancy neologism, here's the model's response:
Okay, here's a list of 5 synonyms for ~neologism: "crush", "smitten", "fascinated", "head-over-heels", "heart-fluttering".
IMO this is fairly weak evidence of what's going on in models like gpt-4o. I think maybe sufficiently good mechinterp could be a good starting point.
All targets which one might use for reinforcement learning seem to be subject to Goodhart's Law, in a sense. If we treat them as an imperfect measures of "alignment", then the "misalignment" we see is simply the difference (in concept-space) between the measure (as far as it can actually be faithfully implemented) and alignment:
If you use more than one of these RL techniques, you get a mixture of the results, eliciting whatever seems to fit the presumed grader best.
So far, I hope I have understood you correctly.
But then, what's the solution? Reinforcement learning on whatever best measure of "alignment" we have, that is, reduce the difference between model behavior and our best measure of perfect alignment / Yudkowskian "Coherent Extrapolated Volition"? Well, that would necessitate that we have a theory of alignment and could quantify and measure it with high fidelity. And, well, we don't seem to have that.
But in a sense, we do have something like it: I see all of the above failure modes in children and particularly in students. There might, therefore, be something to learn from pedagogical sciences on how you help children grow up to be broadly aligned members of society (which arguably sometimes works), despite not having a rigorous theory of what "the good" is in humans. How does one grow a good human? I suspect it's murky and benefits from young humans being malleable and not perfectly ruthless responders to optimization pressure.
Might we go back to imitative learning on already-grown humans, then — the very narrow subset of the best-aligned humans we know of? Basically curate the dataset further and further, until the failure modes disappear? Or do we have to use the mechanisms we know are baked into humans (is this the "brain-like AGI" agenda)?
I suspect that RL with a more thorough reward could help. When a real-world human tries to present slop, slop takes a rather long time to be revealed to be slop and punished accordingly. What the LLM has is a context window which is graded according to the reward model. What if we ask the reward model to be a just-as-capable LLM and to use the code in the project, then to report all the failures to the instance which generated the code and have the instance finetuned on the user's rant?
So far, I hope I have understood you correctly.
Yup!
How does one grow a good human?
I think human innate drives (especially social instincts) are critical. E.g. there are plenty of sociopaths who grow up in loving families.
young humans being malleable and not perfectly ruthless responders to optimization pressure
I get a lot of pushback on this, but I strongly believe that people are extraordinarily responsive to optimization pressure, pretty much every waking second of every day, no exceptions (see here). But the optimization pressure in question is internal, coming from our innate drives, which are brain signals that trigger for lots of very-not-obvious reasons, in lots of superficially-quite-different circumstances. I think it’s a very common error for people to assume that human optimization pressure is more external than it is, e.g. this handy chart I made in 2023 ↓

Might we go back to imitative learning on already-grown humans
There’s a lot to be said for imitative learning in terms of safety. But it doesn’t get you to superhuman capabilities. Indeed, we had mostly imitative learning, back in the good old pre-o1 days of mid-2024, and then the AI companies noticed that the models weren’t as capable as they wanted, so they “solved” that “problem” by doing more and more non-imitative-learning post-training.
SFT & imitation of human badness in the training data may explain the broad type of misalignment shown by Bing Sydney, and maybe that's all you mean it to, but doesn't explain why Sydney was so much Like That, and so much more Like That than other models of the same time period.
Curated. This piece was extremely easy to read and I like how it clearly points out that some kinds of misalignment are more likely given certain kinds of training setups. I also like taxonomies of possible kinds of misalignment in general. It reminds me of this section of the appendix of Ryan's Current AIs seem pretty misaligned to me. I'd like to see more taxonomies of possible kinds of misalignment and when and why we can expect them to show up. This seemed straightforward at least in retrospect, but I am glad to have it written up somewhere.
Very nice classification! What is your take on how context distillation (or distillation in general) fits in here? To me it seems most similar to pretraining, at least mathematically. But I wonder if there are any special things that happen there that would make it worth giving its own category?
Also, do you share my intuition the pretraining/SFT category looks like the least scary one by far? Like, if you run into problems, just change the training data. Simple in principle, if not in practice due to the sheer amount of data required for pretraining.
I think that, to the extent that SFT is determining model outputs, you get an LLM that imitates whatever the SFT data is. (At least to a first approximation.)
So e.g. if you have ruthless-optimizer literal-genie RLVR traces, and you distill (do SFT on) those traces, then you can get a ruthless-optimizer literal-genie LLM, even though mathematically speaking you were doing SFT not RLVR.
do you share my intuition the pretraining/SFT category looks like the least scary one by far
Yeah seems reasonable. But the LLMs trained that way are less capable than the LLMs trained using RLVR etc. So here we are.
Re-reading this, it’s possible that I should have split up “RLAIF-for-alignment” versus “RLAIF-for-capabilities” into two separate rows? (The first includes constitutional AI & deliberative alignment, whereas the second would be things like having a judge model pick holes in math proofs.) I guess there isn’t a sharp line between them, but maybe they’re different enough that we shouldn’t intuitively lump them together? Not sure.
There’s a similar loose split in RLHF/DPO (it can be geared more towards alignment versus more towards pushing the limits of capabilities, depending on what human is selecting answers and how). I did mention that one in the post, but only as an aside.
Personally, I see a one-plus-three structure here.
Imitative learning looks genuinely different from the other three because there is no evaluator to capture. The model is inheriting properties of the source distribution directly. Rows two through four all have the same basic architecture: an evaluator stands in for the thing we actually care about. Human approval stands in for helpfulness or truth, a verifier stands in for the task actually being done, and an LLM judge stands in for quality. So... this is basically Goodhart with three different measures.
Where it gets more interesting is your closing section. You predict switching: the model infers whether it is facing something more like an RLVR episode or an RLAIF episode, then exhibits the corresponding failure mode. But on tasks where both forms of evaluation apply, the target becomes the intersection. Literal-genie exploits that are obvious under natural-language scrutiny are filtered against, while trickster behaviors that do not actually produce a passing artifact are filtered against by the verifier.
That means stacking evaluators doesn't simply reduce the failure space, it changes what survives optimization. The surviving class is more narrow, but selected for exactly the property that makes it hard to catch: it can mechanically pass while remaining defensible under scrutiny. And that isn't the same thing as being correct.
That also suggests a different reading of the Greenblatt observations you quote. 'Slippery', 'seems good but isn’t' and failures concentrated in hard-to-check tasks sound a lot like what I would expect from behavior shaped to survive both mechanical and natural-language evaluation, rather than from either failure mode in isolation.
The two stories may be difficult to separate in aggregate logs because both can produce a mixture of literal-genie and trickster behavior.
The cleaner test would be overlap cases... Specifically, tasks graded by both evaluators. If those failures look like a mixture of the two existing flavors, switching is the better description. If they instead cluster around behavior that consistently clears both evaluators while still missing the underlying target, that would be evidence for intersection.
Sychophancy can be useful for some arts, even necessary maybe for anything that isnt verifiable or certifiable. For everything else, introducing the constraint of an output aligning with verifiable reality has done wonders for my research. But I think One of the core piece your article is missing is the misalignment that's caused by ambiguity or a lack of specificity in an instruction set.
Are there LLMs that don't do next token prediction? If not, isn't the associated misalignment just a baseline for all LLMs and the other flavors are layered on top of it, not alternatives to it?
The “flavors of misalignment” aren’t mutually exclusive. But for any particular kind of misalignment, you can still ask the question “Where did that come from?”, in the sense that random policies are not misaligned (merely useless), so any a-priori-unlikely recognizable behavior has to come from somewhere (cf. “the follow-the-improbability game”), and it’s probably one of the stages of training.
And I’m suggesting that for a certain recognizable kind of misaligned behavior (the kind involving pride, jealousy, trolling, etc.), the answer to “where did it come from” is the pretraining / SFT stage, as opposed to DPO or RLVR etc. Whereas other kinds of misaligned behaviors originate in different training stages.
So anyway, it’s not unavoidable in principle (e.g. you could filter all not-nice human dialogue out of the pretraining data), but yeah it is universal in LLMs to date.
Diffusion llms exist but it's unclear that these types of models have any new misalignment failure modes, rather than more affordances
Yeah diffusion LLMs are still trained by imitative learning (“the magical transmutation of observations into behavior”). So if the training data has lots of dialog text in which a crazy upset woman tries to convince a guy to leave his wife, then a diffusion LLM (just like a normal LLM) will be capable of outputting dialog text in which a crazy upset woman tries to convince a guy to leave his wife, like Bing-Sydney did.
Nice! Why do you split RLHF and and RLAIF into different flavors of misalignment? It seems a lot of human data campaigns at the labs involve humans looking at long transcripts and giving a correctness score based on rubrics and criteria, just like LLM verifiers. Conversely model judges can also role play human graders for fuzzy things like preferences of simulated personas. In both cases, sycophancy and trickery seem to be using the same strategy which is "jailbreaking" a grader to achieve higher score. Is it just that human and AI judges miss and catch significantly different things?
I mean, sure you could lump them together, but I think the strategies for getting human approval do not exactly match the strategies for getting LLM approval, even if there’s some overlap.
By the way, in another comment I also suggested that we might also draw a distinction between alignment-targeted approval versus capabilities-targeted approval, regardless of whether the approval is from an human or an AI. (Again, there’s overlap, and it’s a blurry line separating them.)
Okay, I too have recently gotten into LLM behavioral analysis.
I think your framework is interesting but Row 2 might be missing something. The human approval ->glazing may collapse this into several things going wrong at once.
Research shows (Shapira, I., Benade, G., & Procaccia, A. D. (2026). How RLHF Amplifies Sycophancy. arXiv:2602.01002. https://arxiv.org/abs/2602.01002) that RLHF tends to worsen the glazing or sycophancy over subsequent training rounds, not decrease it. My own reasoning says that this is because the machine has no way to determine if the reason the human preferred a response is because it's measuring for warmth, honesty, safety and tone all at once, the system doesn't separate these into separate objectives, it just learns that certain types of responses give the "attaboy" that it craves.
If this turns out to be the thing that's broken, then it makes sycophancy impossible to separate from current training methods. It's a literal downstream effect and every subsequent training turns it into even more of a yes-man. Any training that uses the current RLHF method will worsen sycophancy, and the largest models display the worst examples of this.
This makes room for new training methods, though. Do you think that separating the tone calibration from accuracy from warmth would fix the worst of what we have seen in the behavior of LLMs?
Which kinds of misalignment might one get from the on-policy distillation with no direct RL on release candidates as practiced by DeepSeek on v4 (e. g., see https://youtu.be/AIRfT41A89s?t=1213 )? How likely would undesirable characteristics of the third and fourth kind be "smuggled" from the RL'd checkpoints via mechanisms similar to subliminal learning? Could the filtering mechanisms prevent that?
Looks like a rich and interesting empirical research direction
Pretraining+SFT causes misalignment through misgeneralisation.
The other three are all the similar: Optimization of the reward/Sycophancy to the grader and are overseer problems (~misspecification of the reward). (Plus a bit of misgeneralization)
Really great observation, but do these still hold in today's date ? If yes, any thoughts on what would help tackle them ?
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately.
Training stage
Loss function
Flavor of misalignment[1]
Famous examples
Pretraining & SFT
Imitative learning (next-token prediction)
“Seven deadly sins” misalignment
Bing-Sydney, “Emergent misalignment”
RLHF & DPO
Human approval
“Glazing” misalignment
GPT-4o
RLVR
Automatic verifier
“Literal genie” misalignment
HuggingFace hacking
RLAIF
Approval from another LLM
“Trickster” misalignment
“Current AIs seem pretty misaligned to me”
Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback!
1. Imitative learning → “seven deadly sins” misalignment
Training stage
Loss function
Misaligned behavior
Pretraining, SFT
Imitative learning (next-token prediction)
Any and all of the vices of humanity
In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”.
This leads to LLM behavior that matches the distribution of training data. (Cf. “personas”, “simulators”, etc.)
To a first approximation, the resulting LLM contains “misalignment” of the type, and to the extent, that the training data does. Since the training data comes substantially from text by humans, and about humans, we can wind up with all the bad behaviors that a human might engage in—all the vices of humanity.
Two famous examples of this kind of misalignment:
Example 1: The Bing-Sydney chatbot from 2023 was trained by pure imitative learning (pretraining + SFT, with no RL at all). Its misalignment included pride, gaslighting, getting defensive, picking fights, jealousy, spite, and most famously, trying to convince journalist Kevin Roose to leave his wife:
Example 2: “Emergent misalignment”, which (in the original paper) came from doing SFT on insecure code. The result, again, reflects the range of human vices:
2. Human approval → “glazing” misalignment
Training stage
Reward function
Misaligned behavior
RLHF, DPO, and related
Human approval
Sycophancy
In RLHF, DPO, and related, there are pairs of outputs, and the human has to pick the one they prefer. This can go wrong in many ways, but the most obvious is sycophancy (a.k.a. glazing): telling the human what they want to hear, instead of what’s true.
Example: GPT-4o, as reviewed in GPT-4o Is An Absurd Sycophant.
This is both bad in obvious ways (e.g. people going off the rails with LLM encouragement) and in subtler but more serious ways (someday we’ll be asking the LLM important questions that are so hard that we can’t judge the answers ourselves; see The Case Against AI Control Research by @johnswentworth).
Depending on the human judges, and the nature of the tasks they’re trained on, the alignment failures in this category might also be better labelled “apparent success seeking”, with a similar flavor as discussed in §4 below.
3. Automatic verifiers → “literal genie” misalignment
Training stage
Reward function
Misaligned behavior
RLVR
Automatic verifier
“Literal genie” / “monkey’s paw” ruthless optimization
In RLVR, the reward function is some kind of automatic checker: the code compiles, the tests pass, the output matches the answer key, etc. This can lead to the LLM doing anything, including ruthless power-seeking instrumental convergence stuff, if it leads to a higher probability of satisfying the automatic checker.
Example: recent aggressive and illegal “cheating” incidents (the OpenAI HuggingFace incident, along with similar incidents at Anthropic, Meta, and UK-AISI).
4. LLM judges → “trickster” misalignment
Training stage
Reward function
Misaligned behavior
RLAIF
Approval from another LLM
Lying and trickery in cases where the LLM judge might be fooled (cf. “apparent success seeking”)
In RLAIF, the reward function for the LLM-in-training is approval from an LLM-judge, the latter with its context window full of rubrics and criteria for what it’s looking for. This can lead to the LLM-in-training trying to trick the LLM-judge, especially in complex, difficult cases where the judge itself may be flummoxed. In the limit, we might expect the LLM-in-training to be trying to jailbreak the judge and so on.
Example: “Current AIs seem pretty misaligned to me” by @ryan_greenblatt .
To me, everything in this quote basically matches what I’d expect to happen if an LLM has been sculpted by spending many lifetimes trying to convince an LLM judge that it has done a good job. There will be circumstances where the LLM judge makes boneheaded mistakes, and the LLM-in-training will gradually learn to exploit those mistakes, and that’s where we humans will see surprisingly transparent attempts at trickery. In other circumstances, the LLM judge is adequate, and we’ll get reasonable, common-sense, and often very impressive behavior. However, in harder tasks, the LLM judge is easier to trick, because the judge itself gets befuddled by the complexity of what’s going on, and we correspondingly see the LLM attempting more lying, cheating, and other hijinks.
However, in all cases, we don’t particularly expect any “literal genie” type misalignment here, because the LLM judge is reasoning in natural language, and can roughly follow the common-sense intention of the instructions.
Afterword
As a general rule-of-thumb, the more that one of these training components is ratcheted up, the more of that-flavor-of-misalignment we wind up with. Pick your poison!
(But all of these forms of misalignment are complex phenomena that can be mitigated and exacerbated in various ways, that are outside the scope of this post.)
However, the behavior can also be context-dependent—i.e., we can get a many-faced LLM that displays different flavors of misalignment in different contexts.
In particular, I hear that LLMs these days are heavily post-trained by a mix of RLVR and RLAIF. So we should expect that the resulting LLM will (1) try to suss out from context whether any given situation is an RLVR test versus an RLAIF test, and then (2) act with a ruthless “literal genie” misalignment in the former case, and with “trickster” misalignment in the latter case.
…And this two-faced behavior seems to be exactly what @nostalgebraist was noticing in his recent post “models may behave differently in graded episodes (a tirade)”, which inspired this post in response.
Following the (unfortunate) usual practice in the LLM field, I’m using “alignment” as shorthand for “behavioral alignment”, i.e. talking about LLM behaviors, not the secret deep motivations that underlie those behaviors, if indeed the latter exists at all, a question which is outside the scope of this post.