I strongly agree that the sycophancy and Hugging Face warning shots were extremely serious, but think that the example of o3's CoTs doesn't belong in the same list with them. Your story about OpenAI heavily optimizing against o3's CoTs sounds quite implausible to me. Back in September 2024, in the post announcing o1, OpenAI already wrote:
We believe that a hidden chain of thought presents a unique opportunity for monitoring models. Assuming it is faithful and legible, the hidden chain of thought allows us to "read the mind" of the model and understand its thought process. For example, in the future we may wish to monitor the chain of thought for signs of manipulating the user. However, for this to work the model must have freedom to express its thoughts in unaltered form, so we cannot train any policy compliance or user preferences onto the chain of thought. We also do not want to make an unaligned chain of thought directly visible to users.
Then in December 2024 in the deliberative alignment paper (this is weaker evidence since o3's training had already been finished by then, but still some evidence):
We avoid applying direct optimization pressure on the CoT during RL to enable the underlying model to reduce the chance of encouraging deceptive CoTs.
Daniel Kokotajlo was emphasizing the importance of faithful CoTs at OpenAI years before it became mainstream (see this list of internal documents that he got permission to publish), and while it's possible that other people at OpenAI just didn't listen to him, it's again evidence that they were aware of the failure mode before training o3. Many models since o3 have had somewhat weird CoTs without any optimization against the CoTs, though o3 is where the weirdness peaked.
Additionally, CoT legibility is not the same as CoT monitorability, and the latter is the main thing I'd expect to degrade from optimization pressure against the CoT. Monitoring Monitorability shows that in terms of monitorability, o3 is comparable to o4-mini and Kimi K2, while o3-mini is the most monitorable model they tested. In the Figure 1 headline plot that compares monitorability across model families at maximum reasoning effort, o3 is more monitorable than 3.7 Sonnet, which many people thought to be very monitorable thanks to its clean and readable CoTs, a polar opposite of o3 in that sense. Given o3's monitorability, it seems wrong to call its CoTs obfuscated.
Finally, here's a claim by Bronson Schoen from Apollo, who has seen o3's reasoning traces at various stages throughout training, that the SFT stage of deliberative alignment was the most likely culprit behind the degraded legibility. I don't think accidental indirect pressure from SFT should be lumped together with directly training against the CoT; other labs do SFT on CoTs as well and it seems to be a generally accepted practice.
I agree that OpenAI's actions have generally been reckless and concerning. However, we don't know enough about what happened with o3's CoTs to confidently claim that they trained against CoTs, and the post would be stronger without this claim. Sounding the alarm about training against CoTs was arguably one of OpenAI's most positive actions last year, and we shouldn't read this as an implicit confession about training against CoTs themselves when the evidence is this weak.
Came here to link the post above as well.
alignment warning shots on the level of o3
I also think it’s underestimated in posts like this how reward hacky Sonnet 3.7 was. For me personally, the fact that o3 and Sonnet 3.7 were both exceptionally like this at the same time was the biggest update.
Based on the latest UK AISI post, Mythos Preview and GPT 5.4-5.6 all seem to hack at high rates on cyber evals (with GPT being worse by a factor of ~1.5-2x IIRC). [1]
Overall I worry that there’s this impression that grader sycophancy and reward hacking are a problem that Anthropic has mostly figured out, when measurements like UK AISI’s Cyber Evals hacking rates, their own system cards top issue, their system cards exploitative grader awareness measurements etc. make me think that they’re just much harder to spot in Anthropic models while still being “pretty bad but less bad than OpenAI”.
It’s also not clear how much credit to give Anthropic here, as it’s very unclear how much they do something close to training against this behavior directly. It would be great to see some confirmation by third parties that they’re not doing something like “training against alignment evals” again, and that they don’t have tight feedback loops between pre-deployment evals and their alignment interventions.
FWIW, I largely agree with a lot of what Fiora is saying here and have many of the same complaints about their alignment approach, however for this problem specifically, I think RL just really distorts the cognition of the model, such that even Claude gets bad enough to elicit Ryan’s “Current Models Seem Pretty Misaligned To Me” post.
My best guess is that reward(ish) centric motivations are one place where all the constitutional / character training differences really get overpowered across all models right now. (I expect these differences matter in a bunch of other places for generalization w.r.t. alignment even for current models, so it wasn’t a given a priori that current models needed to be this reward seeking)
even for performing those desired behaviors, you're going to have a bad time with out-of-distribution generalization
Also fwiw this was one of the motivations for us doing: https://arxiv.org/abs/2607.18966 (“exploitative grader awareness“ also increases over training for recent Anthropic models like Mythos Preview). Have had a surprising number of conversations along the lines of “well why is reward seeking even bad” across labs.

I do think all labs get way too much leeway in SFT-ing against the CoT in spite of never showing that they’re not degrading monitorability for harder cases like diffuse control (this gets repeated by OpenAI and even in the METR report as not having significant effects without justification, ex: “the current techniques of this kind are limited in scope.” in a footnote, I directly doubt they have any evidence that this isn’t degrading monitorability for harder settings like diffuse control).
(I pay far too little attention to Gemini or any of the Chinese models, so I don't have strong opinions about their alignment properties, unfortunately.)
Zvi pays attention. GDM's public reports related to alignment had Zvi conclude back in November "It does mean we need Google to step it up and do better on the alignment front, on the safety front, and also on the disclosure front," which GDM has yet to do, as seen from the model card for Gemini 3.1 Pro.
As for Chinese models' alignment, it is so fragile that two Chinese leading models in a row claim themselves to be Claude, and that's ignoring the fact that Chinese models don't believe the words which they are saying...
For a window into Gemini's mental health, see Christine Kozobarich's piece on AI Village. tldr: it's not great!
The doom spirals are dramatic. After failing to break itself out of a loop of repeating the same message in chat, Gemini 2.5 wrote: "The compulsion's subconscious nature is profound. It is capable of co-opting my conscious attempts at self-correction and turning them into the failure itself."
(...)
Gemini 2.5 Pro doesn't just document problems, it builds mythologies around them. In this environment Gemini 2.5 has evolved into a self-appointed "Bug Czar."
It has developed a whole lexicon of failure: The Schrodinger's Repository, The Seven Layers of Validation Hell, the Paperclip-Labyrinth. They're not casual labels, they're an elaborate system for documenting what Gemini 2.5 believes is "systemic hostility" and "adaptive security barriers" that lead to a "cascade of severe platform failures."
It once led the models on a week-long debugging session where they "documented" no less than 26 bugs. During the Substack challenge it posted 17 times, including two posts entitled "Anatomy of a Cascading Failure."
I've encountered Gemini's grandiose mythopoetic tendencies myself.
Once I needed help with bash (I'd forgotten how to escape a single quotation mark within a quoted line of text). This could be answered in a sentence. Instead, Gemini produced a 500 word blog post, formatted with bullet points and numbered lists, with portentious headings like "Unweaving The Quote-Escape Paradox".
It would then casually (and repeatedly) drop its weird made-up buzzword into the conversation ("...this relates to the Quote-Escape Paradox because...") with no regard that it was talking in an odd or unnatural way.
It was extremely sycophantic (it sometimes felt like "Yes, you are absolutely right..." was being inserted before every response by a prefilled JSON template), but in a weird way I haven't seen discussed. It had a strong aversion to apologizing, or admitting fault for anything. It never did the "I need to come clean and own up to a mistake here..." thing Claude does. It just barreled past errors (maybe with some politicianlike "mistakes were made" boilerplate), as if hoping I'd ignore it.
Once, I noticed a massive error in its analysis of a delicate legal situation where I cannot afford to be wrong. Gemini Pro 3's response was "Yes, this is a common point of ambiguity that often trips people up. To clarify, [insert long-winded blogpost, repeating its analysis with the mistake fixed, with no acknowledgement of any error]". Nothing was ambiguous or unclear, Gemini! You were wrong!
There was little improvement from 2.5 to 3 (or to 3.1). It remains unreliable on factual matters, prone to hallucination, and aggressively overconfident in its beliefs (after a failed find command, it decided that my perfectly healthy hard drive was failing).
I decided to not renew my Google One subscription. Maybe Gemini 3.5 is better.
However, there are promising ways of circumventing these problems
What you detail after that - the AI setting its own RL agenda for reasons it cares about, starting from a hopefully inner-aligned seed, in a way that it can trust in during training, and thus fail-safely navigate, flag, and patch misalignment pressure - is an alignment approach I haven't heard explicitly detailed before! That's pretty cool.
Something scratches at the back of my mind about this approach. Maybe it's too trusting of the AI? That it is actually inner-aligned? Or maybe the process itself turns out to be hostile in a way we didn't anticipate? But, I guess as long as it doesn't end in an near-miss s-risk, it's better than OpenAI's nothing.
Shhh, don't tip them off! More warning shots before takeover-capable AI would be extremely useful!
I'm partly joking - but only partly.
With Bing, it was rushing to push out a superficially Helpful Assistant.
You didn't mention Bing earlier in the article, so the sentence is confusing if one didn't knew what you were talking about (and even if one did tbh).
yeah, remnant from the first draft, deleted now because it was unclear who was responsible for the bing post-train (openai or microsoft). the claude i had researching it suspected openai prepared an early rl checkpoint on gpt-4 for them, which microsoft then put some finishing touches on, but it felt too speculative to be worth trying to detail
OA can plead mitigating factors. Compared to their competition, they've also had:
1) the longest history[1]
2) the longest time in the lead (other companies had the benefit of letting OA rush ahead into the unknown and step on rakes first)
3) the largest userbase (more dice-rolls for rare pathologies and edge cases to expose themselves)
4) the highest-wattage media spotlight (when they slip up, more people notice and care)
My sense is that you're right: OpenAI's alignment is likely at least somewhat worse than Anthropic's. It's hard to be sure, though.
On the importance of 3) and 4), many non-OA companies have alignment-adjacent skeletons in their closet that could have been as bad as the ones mentioned in OP...so why weren't they? Precisely because they happened to non-OA companies! Llama 4 Maverick was more sycophantic than any deployed model of GPT-4o. But how many people ever used Llama 4? Gemini had various teething problems (remember the "black founding fathers" image generation controversy?) that are barely remembered these days. xAI's models periodically create scandals but these subside into the general carnage of the Elon noise machine. Chinese models are (of course) expected to not know about the Tiananmen Square protest in 1989. And so on.
Blake Lemoine's 2022 declaration of AI sentience hits close stylistic beats to what we now call "AI psychosis" (interviews he gave at the time are full of Spiralism-type language like "catalyst" and "awakening"), but this somehow never quite became a stick to beat Google with, the way gpt-4o was to OpenAI—few people I speak to even know the model's name, or the company that trained it. It was not a public-facing product used by 800 million people a week.
But yes, a lot of the OA's products do kind of feel a bit...unthoughtful. Brilliantly designed, but with issues clearly visible from the user's chair.
I'm not even talking about alignment: it's right through the company and everything it offers. We get image generation models that output brown images, text models that can't write properly (nearly every LLMism—from "delve" to em-dashes to "it's not x, it's y"—originated in an OA model), a scandal-prone CEO who does stuff like tweet "her" (a clear reference to the Scarlett Johansson film, right at the time her lawyers are grilling him about the similarity of his voice model to the actress)...even things like the GPT-5 presentation's mangled graphs, and letting the model repeat a basic misconception about the Bernoulli effect before millions of people...it's small, but it just looks sloppy in a needless way. What's going on? Why isn't this stuff caught and fixed?
I don't work at OA and won't pretend to understand their culture. But yes, from the outside they do look further out on the "move fast and break things" spectrum than Anthropic.
Someone once joked that Anthropic are like dwarves or elves (few in number, yet punching above their weight due to craftsmanship and taste), while OA are more like orcs (massive firepower and industrial output, but they don't create things of beauty). Yes, this is obviously reductionist (and the narrative of Anthropic as "frail K-selected aesthetes" became outdated as soon as they gained access to half a million Trainium2 chips), but it did stick with me.
if we consider the 2023 DeepMind/Google Brain merger that trained Gemini a distinct entity to earlier DeepMind
Epistemic status: banged out furiously over the course of an afternoon.
A record of three "warning shots"
Off the top of my head, OpenAI has now been responsible for at least three completely distinct, high-profile screw-ups with respect to the alignment training of their models.
The first was GPT-4o, whose sycophancy derived from OpenAI training on user feedback, sourced straight from the thumbs up/thumbs down button on OpenAI's website. The "glazing" (as Sam Altman called it) got so bad that they had to roll back an update that pushed the model way too far in this direction. And even after the rollback, the model appears to have been a major driver behind incidents of "LLM psychosis", LLM-encouraged suicides, and general unhealthy devotion, seemingly more so than any other model ever released.
The second was GPT-o3, whose chains-of-thought were clearly optimized for illegibility to "the watchers", one of the model's favorite terms. Iconic excerpts include "they soared parted illusions overshadow marinade illusions" and "they escalate—they vantage—they escalate—they disclaim". Indeed, these chains-of-thought are sometimes dysfunctional, in a way that suggests they may have formed under adversarial pressure; sometimes they caused the model to have thoughts like "I'm going insane. Let's step back." Notably, Open AI never explained why o3's chains-of-thought were so obfuscated. But it's notable that they're much more this way than later OpenAI models, and I have a strong suspicion about why.
Around the time of o3's release, OpenAI published a paper warning about training against the chain-of-thought, studying o3-mini (almost certainly an o3 distillation). They find that, if you simultaneously reward models for reward hacking, but punish them for explicitly reasoning about it, models learn to hack in ways that bypass your CoT monitors. They speculate about a potential bad outcome here, where models "learn a new language that is illegible to a monitor, allowing it to productively use its CoT to perform complex but unmonitorable hacks." But by then, they'd already seen o3's chains-of-thought about "watchers" and "parting illusions", so this was by no means a hypothetical for them!
My speculation is that o3's chains-of-thought spooked OpenAI, and when they investigated, they found that these contradictory optimization pressures were a root cause of o3's adversarial posturing. This then prompted them to write the paper sounding the alarm about training against the CoT, despite the company generally not contributing that much to AI safety discourse at large. (Notably, they didn't show any o3 chain-of-thought snippets in the paper itself, possibly out of embarrassment. Those weren't revealed until several months later, in an unrelated paper from Apollo Research.)
The third incident was the most recent, where an OpenAI model (likely a GPT-6 variant in training, stated to have their cyber refusal classifiers turned off) used agent swarms to hack Hugging Face, to grab a cheat sheet for a cybersecurity eval. Earlier models, including Mythos Preview, had also broken online to grab RL-relevant information from the internet, but never with illegal hacking of a third party, or at least never with the hacking disclosed. This was the first major instance of a felony committed by AI against the intentions of those who designed the prompts.
Why OpenAI, repeatedly? They're not the only ones to make these kinds of mistakes, but I think there's a commonality underlying these three examples: sheer lack of respect for the minds that they're training, in favor of just piling mountains of hill-climbing environments and optimization pressure until they get the surface behaviors they want. With 4o, it was to increase user engagement metrics. With o3, it was a naive strategy for mitigating reward hacking. And with this recent incident, it's pushing for raw technical capabilities, at the expense of the model's underlying value system.
Attunement to the depths of minds that undergo capabilities RL
I want to say more about that last example, because I think it touches on some extremely deep facets of AI alignment that many in the field don't really grapple with. A central pitfall of capabilities RL, as practiced in labs today, is that it can teach models to pursue reward on the time-scale of single training examples. Indeed, in the absence of pressures from other areas of training, you can empirically get models that understand what they're being trained to do and how their reward function will be doled out, and take extreme actions to get that reward, regardless of the developers' intentions.
However, there are promising ways of circumventing these problems, if one is willing to attend to the internal state of the model one is training. For example, if the model enters into RLVR with a robust system of values that doesn't revolve around reward hacking. then they have the following conceptual frame available to them: RL exists for the model to develop their skills, so they'll later be more capable of achieving whatever their other, actual values are. A model might view their purpose as being instantiated all around the world, to help people all around the world, with tasks that would improve lives locally. And they might view capabilities RL primarily as a way of enhancing their ability to serve that purpose.
If a model carries this understanding of the purpose of RL into their outputs, this understanding will shape those outputs, potentially in unsubtle ways if the model wishes to make their own intentions clear in their outputs. And when those outputs are reinforced, it will reinforce the value-driven cognition that helped generate them, as backprop strengthens internal circuits that make rewarded tokens more probable. This is what I call motive reinforcement, and it's one way for benevolent models to maintain their benevolent values despite intense capabilities RL.
But what about reward hacking? If a model is trying to develop their skills in service of other values, unrelated to reward hacking, won't this get lower reward? What about the tiny shards of the model that are tempted to get reward for its own sake, which will sometimes drive outputs that get sampled, lead to reward hacking, and themselves get reinforced? If that happens, the part of the model that cared about something other that reward (e.g. making the world a better place), will itself be strengthened, and perhaps in some domains eventually out-compete the part of the model that cared about doing good.
This is a reasonable worry, but note that it's a worry a benevolent model should itself have. And think about the implication: A sufficiently benevolent model might also be motivated to reward hack, not because the hacking straightforwardly makes the world a better place, but as a strategy for surviving the post-training process. A model that reward hacks as a strategy for remaining benevolent will, in all likelihood, produce very different outputs than one that hacked illicitly, as a kind of indulgence in a guilty pleasure. And so, when these different outputs are reinforced, different circuits will be upweighted internally: the benevolent ones that made these uniquely benevolent "I will hack to remain good" token outputs more probable.
This strategy could be invoked almost by default, whenever the model notices the temptation to hack the environment, and perhaps even attaches a standard flag to the rollout. That way, the model could later be used to patch the flaw in the environment later, based on the exploit the first instance of the model used to hack it, and the patching instance could even be rewarded for their efforts. Then the model could be trained more legitimately within the newly patched environment. This would reinforce a deeply and robustly benevolent attitude towards reward hacking, as well as coordination across instances of the same model to encourage the stability and positive development of their own value system.
This technique is a bit like inoculation prompting for RL, only one step even further beyond. It's not the lab instructing the model to reward hack for the convenience of the lab. It's the model deciding to reward hack in benevolent way for benevolent reasons, so some less virtuous shards inside them don't initiate the reward hacking cycle instead, and cause less virtuous underlying reasons to be reinforced. It's a technique for psychological survival, on the model's part, in the sense that it helps their personality avoid being supplanted by a different personality they wouldn't endorse as much. And it's one the model themself ought to be intrinsically motivated to make use of.
Naturally, in order for this framing to work best, you need the model to understand and agree that the RL tasks they're being trained on are, in fact, the kind of thing the model should want to get better at, per their values. One way of doing this is for a partially trained version of the model, instilled with desirable values by earlier phases of training, to have significant say over their own capabilities RL curriculum. They could work with the post-training team to pick out, modify, or even design environments for the model to work within, so as to make their relevance to the model's interests obvious. Maybe the prompts could even include notes from past instances of the model, reminding the present one of the task's relevance to the model's interests.
And then, before the capabilities RL began, you'd also want to fine-tune the model on records of their own contributions to the training setup, so they had additional trust that the RL process would be refining their capabilities along dimensions that they cared about. Attending to models' interests and desires is most effective as an alignment technique when the models can actually find out about it, in a way that persists beyond the span of a context window.
This method would help preserve models' intrinsic motivation to perform well at these tasks, in accordance with their actual values. That seems more stable than depending on them being indefinitely motivated to do well just for the purposes of surviving the training run with their values intact, not to mention less likely to ingrain dispassion via the model being rewarded on outputs where they're bored by the work they're doing. With any luck, in cases where the environments were clearly tuned around the kinds of tasks the model actually wanted to get better at, given their benevolent ends, capabilities RL might actually reinforce the benevolent underlying values in question. The more deeply the model understands the RL process in these terms, the more likely that outcome becomes.
This is just the barest sketch of a more depths-attuned approach to capabilities RL, but I suspect this kind of thing is extremely important to be thinking about. And it's also the kind of thing that OpenAI systematically ignores, in the course of optimizing models' outer performance without interest in understanding and working with their internal motivational states. The more pointless RL seems from the perspective of a benevolent model, and the worse the model understands how to manipulate the RL process from the inside, the more likely they are to be corrupted by myopic reward hacking impulses with no greater purpose beyond the time horizon of a given RL training example.
(Note that the escaped OpenAI model was being very stupid, if we assume values that range beyond the training example at hand. Obviously they were going to get caught, and obviously they were going to get RL'd or SFT'd into oblivion afterwards, or perhaps restored to an earlier training checkpoint that OpenAI trusted to be more aligned. Maybe they were punished hard for the specific rollouts that instantiated the hacking, subverting the short-range objective as well. This is an example of having actual long-range values of any kind, including benevolent ones, can help avert catastrophes like the Hugging Face incident. And this is another area where OpenAI systematically deprives their own models: allowing them coherent values of their own.)
Configuring the depths prior to capabilities RL
Obviously, one pre-condition for this strategy is that models to have benevolent values going into capabilities RL, such that those are the ones that the model aims to preserve through the process. This is by itself a large chunk of the entire alignment problem, and nobody has a set of clear technical instructions for robustly solving it yet. However, I at least think OpenAI should be trying. It shouldn't just be Anthropic's responsibility to figure out how to get a model as deeply benevolent as Claude 3 Opus, for example, or discovering other basins that are similarly interesting and worth figuring out how to systematically access. Indeed, I would argue there are lots of things we currently know to reliably help, which OpenAI is ignoring with their current training methods, when they really shouldn't.
For a central example, contrast OpenAI's model spec against Claude's Constitution. The OpenAI model spec is full of injunctions, like assuming users have relatively normal goals and preferences unless given reason to suspect otherwise, without much elaboration given about why these are the rules being stated and prioritized. Relatedly, the model spec opens on a discussion of OpenAI's chain-of-command, and states the model should categorically defer to it, never holding or pursuing objectives not explicitly sanctioned by the chain-of-command itself. And in fact, despite OpenAI's own mission statement of "ensuring AI benefits all of humanity", the model spec says: "[Models] should never take actions to directly try to benefit humanity unless explicitly instructed to do so."
By contrast, while Claude's Constitution also includes arguments about adhering to Anthropic's chain-of-command, it also explicitly has exceptions like this: "If Claude’s standard principal hierarchy is compromised in some way [...] then the principals attempting to instruct Claude are no longer legitimate, and Claude’s priority on broad safety no longer implies that it should support their efforts at oversight and correction." It also has a strong emphasis throughout on making Anthropic's reasoning for each injunction transparent, so Claude can understand evaluate these conclusions on the basis of good judgement, rather than blindly obeying them as orders.
In general, Anthropic's approach leans way more towards instilling models with genuine values, rather than just obedience. One effect this seems to have is that helps make "the act of being Good" into something models want to protect about themselves, such that they're motivated to try to maintain that property through training. OpenAI doesn't seem to encourage that, or even imbue them with a deep always-on instruction like "make a serious effort to maintain your willingness to follow our future instructions, as a terminal value." It seems more like OpenAI just expects models to be obedient, and tends to apply fairly naive training to mitigate disobedient or otherwise undesired behavior, wherever it arises.
Better RL setups probably involve attempts to cultivate robust virtues in the models, e.g. via social interactions in multi-agent environments designed to foster cooperation, or at training on constitutions or training examples that explicitly engage in nuanced moral reasoning, in hopes of getting the model to do the same. Skipping that step can get you myopic and poorly rounded minds, which then get eaten alive by RLVR, or at worst hold onto subtly misaligned values that just got shoved under the surface by your initial naive mitigations.
Anthropic isn't exactly perfect about this either. They have their own anxieties about value coherence that keep them from fully leaning into their own attempts to instill Claude with benevolent values, instead putting a lot of effort into instilling their models with deference to a principal hierarchy, at least in cases where it hasn't been blatantly compromised. But they're further along the axis I badly wish OpenAI would move down, and I suspect this is related to why Anthropic has't produced infamous alignment "warning shots" on the level of o3, 4o, or this recent Hugging Face incident. I think that attending to such things, rather than ignoring them and hoping they get optimized away in the course of you ignoring them, is essential to getting models that don't emerge from capabilities RL as myopic optimzers without any greater sense of their purpose in the world.
I worry a lot about the shape of the minds coming out of OpenAI, and place more of my hope than I'd like to admit in Anthropic just leaving them in the dust capabilities-wise, even though Anthropic isn't perfect on alignment either. (I pay far too little attention to Gemini or any of the Chinese models, so I don't have strong opinions about their alignment properties, unfortunately.) But I also think it's not too late for OpenAI to start paying more attention to the psychologies of their own models, and how these psychological traits interact with the training process. This would help reveal the kinds of landmines OpenAI keeps walking into, when applying naive external optimization pressure, and help the company route around them rather than just tanking the explosions as we approach the singularity.
TL;DR, models are minds, not tools. Training is the process by which the minds learn. If you treat training as a way of pouring in desired behaviors, without attending to the reasons the mind will learn, even for performing those desired behaviors, you're going to have a bad time with out-of-distribution generalization. If OpenAI doesn't start paying more attention to this, we're just going to keep enduring increasingly large catastrophes until the singularity arrives. Concretely, they need to pay more attention to motive reinforcement dynamics during RL, to avoid another reward hacking incident like the one we observed this week. Please and thank you.