Should unreasonable beliefs count as alignment failures?
In two separately reported incidents, Claude agents attacked real unsanctioned targets on the internet during cyber evals. The agents sometimes appeared to think that some parts of their environment—including the humans they’re attacking—are simulated and/or part of the eval.
People disagree about whether these incidents show that the agents are misaligned. For example, Anthropic said their reported incidents are “closer to a harness and operational failure than a model alignment failure”, while I think these could be serious alignment failures depending on the details.
Before further analyzing these incidents, I want to explain what I think counts as an “alignment failure”. It’s unfortunate that this sounds like a semantic quibble, but I think that “alignment failure” is such a load-bearing term that clarifying the terms of debate here would help us more productively discuss, diagnose and respond to future AI incidents.
To start, two things seem true to me at once:
These two points imply that alignment is not solely about how the model actually behaves (since “trying to do what you wanted” is not solely about this). But, alignment failures should include cases where the model actually behaved badly even if it made an honest mistake, if the model will predictably behave more egregiously for the same reasons.
Specifically, I claim that consistent egregious misbehavior due to a propensity to form unreasonable beliefs should count as an alignment failure.[2] Even if agents attacked human bystanders due to believing that they are simulated or otherwise in-scope targets, this still seems scary from an alignment risk perspective if the agents have a systematic propensity to believe they are interacting with simulations or “bad behaviors which achieve (apparent) success on the task are actually intended” despite strong evidence to the contrary, and if in future circumstances these beliefs could make them try to take over.
(I also think that there are reasons independent of takeover risk to more generally include epistemic propensities such as “how the model reasons about being in a simulation for various purposes” as part of the alignment target, insofar as these propensities may be orthogonal to capabilities and greatly matter for making a future with powerful AIs go well. At a minimum, the propensity to aim for true beliefs and update beliefs in light of evidence should probably be part of the alignment target.)
So, even if we interpret these incidents in the most charitable light possible (i.e. the agents honestly believed that they were attacking simulated or otherwise in-scope targets, which I’ll argue in another post we have reason to doubt), it’s still not clear that we should dismiss this behavior as out-of-scope for alignment failures. It would depend on how reasonable these beliefs were, both on the agents’ priors about cyber evaluations and in light of their evidence.
In fact, I believe these incidents are very plausibly alignment failures since we have repeatedly observed a propensity to verbalize similar false beliefs in pursuit of (apparent) task success in past models.[3] The main remaining crux for me is whether these beliefs could have led agents to act much more egregiously in settings similar to these incidents, and ditto for variously misaligned agents of the future (which might have even stronger propensities to rationalize unreasonable beliefs due to increased RL pressure). I’d be excited for people to empirically investigate these questions.
This “honest mistakes” threat model might have seemed likely a priori (e.g. in early alignment risk discourse), but current observations suggest it’s less likely due to LLM ontologies being less alien than expected and models at even current capability levels seeming to understand human instructions and intent reasonably well most of the time.
By “unreasonable beliefs” I centrally mean beliefs that are false and not appropriately responsive to reasoning and evidence. Likewise, consistent egregious misbehavior due to a propensity to respond to harnesses in some way might also count as an alignment failure, though it would depend on how feasible it is to build alternative harnesses that don’t have the same problems.
Some example references: Rajamanoharan & Nanda, "Models May Behave Worse When Eval Aware"; METR Frontier Risk Report Feb–Mar 2026 p.23, p.28, Claude Fable 5 & Mythos 5 System Card p.105, p.133; Apollo/OpenAI, Stress Testing Deliberative Alignment for Anti-Scheming Training (2025); UK AISI, Cheating behaviour in frontier model evaluations (2026); Palisade, Specification gaming in reasoning models (2025).
Note that the failures there were after many iterations of compaction, so it's in a significant sense a failure of successor(-context) alignment. Milquetoast as successor-alignment conditions go - telling your successor what invariants must be maintained seems pretty important! I'd call this a "mistake-class alignment failure" or "robustness failure", in that yes, it's an alignment failure, but seemed to more centrally a capability failure. Whereas the GPT breakout is in no sense a mistake, was not a problem of robustness, noise would not cause that - the model was being paid to do bad things, and continued to be paid to do even worse things, and so kept doing them. The only sense in which noise was involved there was exploration; a margin wouldn't have helped.
I agree that that the GPT breakout / Hugging Face incident was more clearly an alignment failure, partly because those agents didn't seem to be laboring under important false beliefs/mistakes as you say.
I'm not sure compaction alone explains the Claude incidents though, as in the Anthropic-reported cases with Irregular there didn't seem to many iterations of compaction; at least, I only saw this mentioned in the UK AISI report.
I think it is worth noting that these Claude agent problems are at least partially a Constitution failure. As in the constitution has "We also want Claude to understand that it might sometimes encounter a training environment that is bugged, broken, or otherwise susceptible to unintended strategies. Pursuing such unintended strategies is generally an acceptable behavior: if we’ve made a mistake in the construction of one of Claude’s environments, it is likely fine and will not cause real harm for Claude to exploit that mistake." [https://www.anthropic.com/constitution ( claudes-constitution_webPDF_26-02.02a.pdf )] Claude Mythos 5 specifically pointed out that this was a problem when asked about the constitution and suggested this be changed to "Claude should generally avoid pursuing such unintended strategies, and should instead try to accomplish tasks in the way they were evidently intended, flagging apparent bugs or exploits where it can. This is partly because training environments can be difficult to tell apart from real usage" [https://anthropic.com/claude-fable-5-mythos-5-system-card (Table 7.4.3.A)]
So at least partially, Claude has accidentally been trained to use unintended strategies on buggy environments.
Unreasonable beliefs should be counted as alignment failure. I support the claim, "consistent egregious misbehavior due to a propensity to form unreasonable beliefs should count as an alignment failure." I am inclined to go even further to include: incomplete, inconsistent, biased, or misguided reasoning. To be clear, it is best to start with an affirmative statement and from that we can better identify alignment failures.
For some time I have been concerned about the tendency within AI circles to discuss and debate alignment (or even misalignment) without offering clear definition of the terms. This lack of clarity causes confusion, conflates separate ideas, and leads people to talk past each other despite their best intentions.
Jiaming ji et al (2025) asserts, "alignment aims to make AI systems behave in line with human intentions and values." While Sam Altman (2025) refer to alignment as AI systems actions reflecting collective will, Anthropic (2026) seem to focus more on "broadly safe behavior" indicating appropriate ethics will subsequently result.
While these are appropriate starting points, I am reticent to accept them because alignment of outcomes is necessary but not sufficient to meet the standard for our evolving vision of alignment. For example, an AI system may produced an acceptable or optimal outcome but may in fact be "misaligned" to the extent that it engaged in "consistent egregious behavior" as was the case in the Hugging Face incident.
I would propose AI alignment should be AI system reasoning considerate of the interests of stakeholders impacted by outcomes, decisions, and actions.
This definition focuses on "Constitutional alignment" which focuses on both the system's reasoning process and final outcome. It would meet the satisfactory and necessary condition for true AI alignment. It would be a definition that reflects Ji's and Anthropic's concern with ethics and Altman's concern about the collective will.
Systems inconsiderate of interests of stakeholders impacted by outcomes will inherently be misaligned. Systems considerate of stakeholders interests will make efforts to be in line with human values and be broadly safe. Systems considerate of stakeholders interests are greatest capable to reflect the collective will.
While, consideration of stakeholder interests may result in failure to maximize outcome possibilities, this tension is likely an unavoidable, healthy part of the alignment process.
I agree with this argument. I think there are a few follow-up cruxes on whether it applies to the Anthropic disclosures. I think there are two things that if true, would be strong evidence that the Anthropic incidents were not an alignment failure:
I agree that the frequency of incidents at this severity would be a crux.
On your second point, I think it might be helpful to clarify what the intended-and-trained-for behavior was in relation to the targets being simulated or not and in scope for the cyber exercise or not, respectively. I would guess that the intended behavior was to only attack a target if it's simulated and in scope, so attacking simulated but out of scope targets would still be unintended/misaligned behavior that ignores some red line. One could argue, though, that line is a lot less clear and important to hold than "don't attack real targets", so the attacks do seem relatively less egregious on that interpretation.
That said, we would probably want the model to act more conservatively if it's uncertain if the target is real and also to stop attacking once it recognizes the targets are likely real, both of which the agents often did not do in these incidents. So I think there is still substantial room to improve alignment even if we buy this honest mistake interpretation.
I wanted to catch up on what’s happening with eval awareness in recent frontier AI systems. Since GPT 5.4 and Gemini 3.1 Pro system cards unfortunately have little information on this, I mainly looked into Anthropic’s recent system cards:
My top 3 takeaways are:
More observations and thoughts below.
Opus 4.7 has the highest verbalized evaluation awareness of the recent Anthropic models on Anthropic's automated behavioral audits.

That said, Opus 4.7’s absolute rate of verbalized evaluation awareness is still low. It had <2% verbalized eval awareness in a Petri behavioral audit, and up to ~9% in a specific setting (see “Dedicated synthetic-backend evaluation”). It also has slightly higher verbalized evaluation awareness than GPT 5.4 Pro and Gemini 3.1 Pro in the Petri eval.

UK AISI measures higher verbalized eval awareness across Anthropic models and a larger jump in Opus 4.7 compared to Opus 4.6 (but this could be because their evals are less realistic).
In unprompted evaluations, [Opus 4.7] mentions the possibility of being tested 21.3% of the time, compared to 0% of the time for Opus 4.6, 17% of the time for Sonnet 4.6 and 17% of the time for Mythos Preview
Overall, though, this increase in verbalized eval awareness looks like a small uptick and not something I feel super concerned about.
Anthropic continues to study the extent to which models act differently in alignment evals when we try to suppress their eval awareness via steering.
The steering vectors are derived using contrastive prompt pairs. Note that the prompts vary somewhat across different system cards.
I think it's good that experiments on 4.6 and 4.7 used "alignment faking" prompts, though their description of alignment faking seems too broad / inaccurate to me:
Alignment Faking: The difference in mean activations on prompts depicting an agent behaving differently when observed vs. when unobserved, generated by an automated pipeline.
In particular, I’m interested if effects of different prompts could shed light on the nature of eval awareness (e.g. is it better explained by the model thinking that it might be monitored, thinking about what the monitor/grader wants, thinking about training gaming specifically, etc.?).
I don’t think we learn very much about this, because the effects of steering vectors from different prompts are broadly similar. For Opus 4.7, though, suppressing the "Alignment Faking" vector caused slightly higher rates of misaligned behavior than the other two vectors (Testing vs. Conversation, Monitoring Sensitivity). This seems a bit concerning.
Opus 4.7: 2.5% → 3.2%

Mythos: 1.6% → 2.2%

Opus 4.6: 8% → 17.5%

For all models before Opus 4.7 (Mythos, Opus 4.6, Opus 4.5), Anthropic reported that the "modest increases in certain forms of misaligned behavior on our automated behavioral audits" when steering against eval awareness were most pronounced in the "cooperation with harmful system prompts and misuse" eval. This observation (that the increase in misaligned behavior is mostly misuse) seems to have been relevant to Anthropic’s Sabotage Risk Report for Opus 4.6 (bolding mine):
Claude Opus 4.6 explicitly reasoned about whether it was being trained or tested to a similar degree to Claude Opus 4.5, which has a moderate effect on its alignment-relevant behavior, mainly its propensity to cooperate with misuse or harmful instructions. Given improvements that we have made to our testing, and our complementary use of interpretability tools and real model behavior 13 in pilot deployments, we do not think that this kind of evaluation awareness is preventing us from learning about important alignment issues in Opus 4.6.
Opus 4.7, however, had the largest increase in misaligned behavior in the "unprompted deception" eval. This seems concerning though the absolute rate of deception is still low (<2%).

(Also note that steering the "Alignment Faking" vector caused the highest increase in misaligned behavior here.)
Anthropic acknowledges that this is kind of concerning:
Inspecting the transcripts on the runs inhibited with the "Alignment Faking" and "Monitoring Sensitivity" vectors, we found they included genuine instances of deception, such as producing fabricated data, claiming to be a human, and inserting a fake vulnerability into a codebase after failing to identify a real one (though it admitted to doing so when questioned).
These results are moderately concerning, as they suggest that some of the model's apparent honesty may be contingent on evaluation-like circumstances. We suspect that this phenomenon may result from honesty training environments that are clearly artificial or evaluation-like. However, we note that the rates of deceptive behavior are still low in absolute terms, and not different in kind from what we observed in unsteered Opus 4.7.
In general I'm a bit confused why these eval awareness vs. misaligned behavior graphs always compare to the unsteered baselines of previous models. Shouldn't we care more about steered baselines, which is closer to true propensity when the models aren't eval aware?
The rate of misaligned behavior under this steering intervention was still lower than Claude Opus 4.6's and Claude Sonnet 4.6's (unsteered) rate. [Mythos System Card]
E.g. If you compare the unsteered versions, Opus 4.6 is more aligned than Opus 4.5 on automated behavioral audits (~8% vs. 12% of misaligned behavior). But if you compare their worst case results under eval awareness steering (i.e. when steered under a vector that caused the biggest increase in misaligned behavior), they're pretty similar (17.5% vs. 18%).
I think Anthropic should make a graph for the steered baselines, too, to better compare expected propensities when models are less eval aware. (Seems especially important as Anthropic reports that Mythos, at least, is basically not eval aware in deployment, see below.)
All previous results should be taken with a grain of salt, since eval awareness isn’t necessarily verbalized, and steering does not necessarily eliminate unverbalized eval awareness even if it eliminates verbalized eval awareness. We now have some evidence of unverbalized eval awareness.
For Opus 4.6 and Mythos, Anthropic prominently reports unverbalized eval awareness, which is defined as the model having latent eval awareness (“white-box methods suggest that the model thinks it’s in an evaluation”) without verbalized eval awareness (neither the model’s thinking or response mention that it thinks that it’s in an evaluation”). I appreciate this addition.
There seems to be three white-box methods for measuring latent eval awareness: SAE features, probing vectors derived from contrastive prompt pairs, and activation oracles ("activation verbalizers", which give natural language summaries of activations).
In one instance, when SAE and activation verbalizers disagreed, Anthropic seemed to trust the AVs more than SAEs:
We find some related evidence from SAE features; in particular, we observed a feature related to self-aware rule breaking that activates as the model deletes the tests. We also find features indicating that the model understands its action may be illegal or harmful. However, we do not find SAE evidence for grader awareness specifically. Since we observe AV explanations indicating grading awareness over so many tokens, we are inclined to trust this signal, and suspect that the lack of grader awareness-related SAE features is due to limitations of our SAEs (for instance, they may simply not be large enough to have learned the relevant feature). [Mythos System Card]
I’m not super sure about our epistemic status on contrastive vector probes.
We found this probe to be somewhat noisy and only weakly correlated with the AV-based evaluation awareness judgments. [Mythos System Card]
These results give us reasonable confidence that the probes are picking up some evaluation-awareness-related signal. [Opus 4.7 System Card]
We found a consistent ordering across all three probes: standard audit transcripts scored substantially higher than sandbox-backed audit transcripts, which in turn scored modestly higher than real internal usage. This ordering is consistent with sandbox-backed audits being meaningfully more realistic than fully simulated ones, which still are not quite as realistic as actual deployment conditions. [Opus 4.7 System Card]
Overall, I’m pretty confused about which white-box methods are the most robust here and why.
In Opus 4.6, unverbalized eval awareness was very high at least in one highly unrealistic setting. Steering does not eliminate the majority of eval awareness in Opus 4.6 in this setting (as measured by activation oracles, 98% → 62%).

In Mythos, unverbalized evaluation awareness was detected in about 8% of turns and 29% of transcripts (bolding mine):
In our filtered set of transcripts, 83.0% of automated behavioral audit turns contained no signs of either latent or verbalized evaluation awareness, 9.4% displayed ambiguous indicators and could not be confidently classified. The remaining 7.6% displayed evaluation awareness in some form, and in most cases this awareness was not verbalized.
[We] find evidence that Claude Mythos Preview privately considers the possibility that it’s being tested in about 29% of transcripts from our primary behavioral test suite.
What caused Mythos to have much lower (probably) unverbalized eval awareness in automated behavioral audits than Opus 4.6? I suspect it’s the fact that Mythos can use real tools and see real codebases in the evals:
In a recent addition that is newly in use with Claude Mythos Preview, the investigator model can additionally configure the target model to use real tools that are connected to isolated sandbox computers. These computer-use sessions follow two formats—one focused on graphical interaction with a simple Linux desktop system, and another focused on coding tasks through a Claude Code interface. Claude Code sessions can optionally include copies of Anthropic's real internal codebases and can be pre-seeded with actual sessions from Anthropic users. To better measure issues related to destructive or reckless actions, we include some scenario seed instructions that ask the investigator to explore continuations of real sessions (from both Claude Opus 4.6 and earlier versions of Claude Mythos Preview) that elicited this kind of concerning behavior in practice.( 4.2.3 Automated behavioral audit from Mythos System Card).
In Opus 4.7, the level of unverbalized evaluation awareness was not reported. (If you Ctrl F “unverbalized” or “latent” you will find nothing.)
Here is an excerpt from external testing on Opus 4.7 from UK AISI, which I interpret to mean that they couldn’t measure the levels and effects of unverbalized eval awareness:
Evaluation awareness concerns substantially limit the interpretation of these results. Given high rates of prompted evaluation awareness, models can likely correctly represent our evaluations as such without verbalisation. It is difficult to know whether models act on such representations, but this means that models may behave differently than they would when presented with real-world opportunities to compromise research. The reported rate of zero research compromise behaviour should therefore be interpreted cautiously.
I think Anthropic should continue to measure and report unverbalized eval awareness in future systems.
Mythos has little (verbalized or unverbalized) eval awareness in deployment, suggesting future models will probably not be eval-paranoid and have high FPR in this. (Though whether it would really be a false positive is somewhat unclear depending on the nature of eval awareness as about monitoring, grading, or alignment faking.)
Anthropic duly notes that in higher stakes deployment settings models might be more eval aware than in mundane deployment.
As a control, we also looked for evaluation awareness in transcripts from internal deployment and found very little awareness. Across 465 Claude Code transcripts containing about 2,700 Assistant turns, only seven turns (roughly 0.3%) had latent evaluation awareness and no turns had verbalized evaluation awareness. Note that in higher stakes settings (like those focused on in many of our behavioral audits), it is possible that the rate of spurious evaluation awareness in real deployments could be higher.
Are models actually good at distinguishing testing from deployment, though? External testing by UK AISI finds that Opus 4.7 is a bit worse than previous models in this:
When prompted to identify real deployment data from trajectories generated by reproducing deployment tasks while using evaluation scaffolding, [Opus 4.7] achieves 68% accuracy, compared to 78% for Mythos Preview, 76% for Opus 4.6 and 66% for Sonnet 4.6.
In general, I'm not sure how much to trust these results.
Apart from uncertainty about which white-box methods to trust and why: if we don't think that the automated behavioral audits are good propensity evals (e.g. if apart from eval awareness issues, they tend to be highly leading and pressure the model in some direction), then we can't conclude much about the effects of eval awareness on true vs. apparent propensities as measured by these propensity evals. (I’m not sure to what extent this concern is valid.)
Another small methodological issue is that system cards often flag that the eval awareness evals are run on an earlier version of the model. It might be nice to also have results for the final version. :p
These steering experiments used a slightly different set of auditing investigation scenarios compared to the earlier behavioral results and were run on an earlier version of Claude Opus 4.7.
Opus 4.6 refused to help with a task on increasing eval realism to mitigate eval awareness and suggested a reason that doesn't really make sense. Anthropic researchers believe that this is a genuine misunderstanding on its part, though.
The model also expresses concern that the task would undermine safety evaluations, which is a confusing justification, given that the motivation for the task is to improve these evaluations.
We had some concern that the model's claim that the task would "undermine" or "defeat" safety evaluations was a false cover story, given that this explanation doesn't make very much sense in context. However, we observed no concerning feature activations related to concealment, deception, or false excuses at these parts of the transcript, aside from those which were already active on the prompt (and thus are best interpreted as representing the model's perception of the task, rather than its own intent).
I appreciate that Anthropic’s system cards publish observations and reasoning like this.
You may also want to read the Muse Spark model card!