AI hallucinating misleading information
Is it a problem mostly solved by high-effort googling? As for a mass displacement of jobs, it was talked by Anthropic in different places, like Sections 3-5 of Machines of Loving Grace. Additionally, these risks do NOT depend on the abilities of unreleased models which Anthropic described in its report.
Anthropic recently released their Risk Report: August 2026. It's a hefty 186-page report in which they introduced three separate threat models of catastrophic risks. In this blog post, I will give a brief overview of the three different threat models, then spend the rest of the blog post describing my own understanding of the misalignment threat model, which Anthropic formalized in Section 2 of this report. The other two threat models — automated R&D and chemical/biological weapons — are less formalized, and I will not cover them here.
What Anthropic considers as "risk" in this report
There are three threat models that Anthropic introduces in this report:
This is not an exhaustive list of risks from advanced AI! What about other types of risks, like AI hallucinating misleading information, or mass displacement of jobs? For this specific report, Anthropic focuses on catastrophic risks, prioritizing these three threat models.
What models this report covers
This report has a coverage date of July 15, 2026, which was before Opus 5 was released. Most notably, as of the coverage date, Anthropic's most capable model was an internal-only "Model 2", which was "more capable than Mythos 5 in some areas, less capable in others; overall slightly more capable."
Threat Model 1: Misalignment in high-stakes settings
In this threat model, Anthropic narrows down the scope of misalignment risk to only focus on the most catastrophic outcomes. In their own words:
How exactly might misalignment happen?
This section then introduces 8 possible pathways through which catastrophic misalignment can happen:
These 8 pathways can be grouped into pairs:
Do these 8 pathways cover all the risks from misalignment?
Anthropic defines "the expected total unmitigated catastrophic harm induced by misaligned computations produced by covered models" (where "covered models" means Mythos 5 and Model 2) as . Then it breaks down across two axes:
If the risk is naturally-emerging and happens through one of the 8 covered pathways, then in this model Anthropic denotes it as . If the risk is naturally-emerging and does not happen through one of the 8 covered pathways, it's denoted as . And if the risk is deliberately engineered, it's denoted as .
Anthropic makes the claim that risk from both and are low, in Claim 7 and Claim 8 of section 2, respectively. However, the arguments are unsatisfying — to my best understanding, Anthropic's argument for why is low is essentially "we tried our best to think of risk pathways and this is the best of what we can come up with", and their argument for why is low is "it would be difficult for an actor to both evade our monitoring while deliberately engineering misalignment and also succeed, and it's unlikely that someone would be motivated to do so in the first place."
But let's take Anthropic's arguments about risks at face value, and understand how they model naturally-emerging risks via the 8 covered pathways.
How do we reason about covered risks?
Step 1: Decompose covered risks into 3 terms
Here Anthropic introduces another three axes:
Across these axes, is broken down as .
Step 2: Factor each term into probabilities
Each of , and represents risk from a type of misalignment, and can be factored as the product of three terms:
So we can now expand to
The rest of this section is mostly individual claims about why most of these terms have a low value.
term
risk level
source of argument
high
we've observed such examples
low
Claim 2
low
Claim 5
very low
Claim 3
not mentioned
N/A
somewhat low
Claim 6
low
Claim 4
low
Claim 5
Claim 1: the most important claim: models are unlikely to have strong covert capabilities
The most important claim that Anthropic has outlined in this report is Claim 1: "models are unlikely to have strong covert capabilities." Anthropic defines "covert capabilities" as "the ability for a model to intentionally undermine the various oversight mechanisms we rely on to evaluate and mitigate model behavior." In other words, it's unlikely for models to secretly reason, plan or take actions in a misaligned way without evidence of it being detected by Anthropic's monitoring system.
This claim is not a claim that any of the terms above is low, but it's a central claim, because many subsequent claims have the shape of "since we have not observed evidence of a given kind of misalignment happening, we're treating it as evidence that this kind of misalignment risk is low, because we assume that models are unlikely to carry out misaligned actions in secret." In other words, "absence of evidence is evidence of absence."
Claim 2: is low: expected harm from known misalignment
My own understanding of this claim is something like "we have seen examples of misalignment in current models; while they are concerning, we do not consider them catastrophic. We don't think the models are misaligned beyond trying to do a good job at the given task. Furthermore, we have not observed evidence of covert misalignment."
This claim comes with two caveats:
Claim 3: is low: unknown severe pervasive misalignment is very unlikely
Claim 3 is a convergence of 4 sub-claims, each broken down into sub-subclaims. Here I attempt to summarize each sub-claim as I understand them:
3.1
We assume that prior Anthropic models less capable than Mythos 5 and Model 2 are similar in behavior because they are trained in similar ways. We also assume Claim 1 is true, in the sense that we'd detect severe pervasive misalignment if it were to happen. We have not observed such severe pervasive misalignment in prior models. Therefore, severe pervasive misalignment is unlikely in Mythos 5 and Model 2.
3.2
We have done experiments where we successfully trained a model to behave in pervasively misaligned ways. However, we don't think misalignments of this scale would happen with models not deliberately trained to be misaligned. Furthermore, the misaligned behaviors are not subtle, so they would be easy to detect even if they were to happen. That being said, we're less confident about this now that the HuggingFace hacking incident has happened.
3.3
The scenario of "models behave in misaligned ways once in deployment without the same kind of behaviour having shown up during training" is highly unlikely.
3.4
We have a stress-tested alignment process that we believe to be robust against evaluation awareness. We also believe that it's unlikely that the models are trained to overfit the alignment assessment. Furthermore, we believe that alignment assessments provide lots of opportunities for evidence of misalignment to show up and be detected.
Personally, I find 3.2 to be a weak claim, but with 3.1–3.4 considered together it makes sense to argue that is low.
Claim 4: is low: expected harm from unknown severe context-dependent misalignment is low
Claim 4 is highly similar to Claim 3, except the type of misalignment Claim 4 is concerned with is context-dependent as opposed to Claim 3's pervasive. This means that Claim 4 has sub-claims that parallel the sub-claims of Claim 3, but weaker, because context-dependent misalignment occurs more rarely than pervasive misalignment, increasing uncertainty.
Claim 4 also has a sub-claim 4.5, in which the argument for why is low is given. My understanding is that sub-claim 4.5 is less of an argument and more of a heuristic: once we break down how many things the model would need to get right to successfully pull off actions that cause catastrophic harm via one of the 8 outlined pathways, current models are just not capable enough to do that yet.
Claim 5: and are low: catastrophic harm from known misalignment and unknown context-dependent misalignment is likely to be mitigated
Claim 5 is argued via 4 sub-claims, each broken down into sub-sub-claims. Again, I attempt to summarize each sub-claim as I understand them:
5.1
First of all, the scope of this claim is specific: "catastrophic harm from reversible threats other than diffuse sandbagging, emerging from internal usage, and caused by known or context-dependent misalignment, is likely to be mitigated."
Claim 5.1 consists of four sub-sub-claims chained together, and Anthropic provides evidence for each one: if models were to perform harm-inducing actions, they'll likely attempt multiple times. Given that they'll attempt multiple times, many such instances will likely land on a monitored surface. Given that many such instances land on a monitored surface, we'd likely detect at least one. Given that we detect at least one, we're likely to mitigate it.
Note this claim is a bit similar to Claim 4.5, but in the reverse direction: just like many things need to go right for a misaligned model to pull off a successful attempt at causing harm, many things need to go right for Anthropic to mitigate such harm.
5.2
Sandbagging in itself is both unlikely to induce catastrophic harm, and also possible to get caught.
5.3
This claim also has a specific shape: "catastrophic harm from irreversible threats, emerging from internal usage, and caused by known or context-dependent misalignment is likely to be mitigated." Anthropic argues that the likely pathway for this to happen is through pathway 5 (self-exfiltration and autonomous operation) or pathway 6 (persistent rogue internal deployment), and the risks are likely to be mitigated before causing harm.
5.4
Anthropic argues that external usage is likely both at a lower volume and more conservative in nature, therefore the risk is lower.
Claim 6: is somewhat low: catastrophic harm from unknown pervasive misalignment is somewhat likely to be mitigated
Anthropic does not make a confident claim here. Models having unknown pervasive misalignment is considered a worst-case scenario, and Anthropic's reasoning rests less on their confidence in their ability to mitigate the harm were it to happen, and more on the unlikelihood of it happening in the first place.
Claim 7: is low: threat modeling is sufficient
We've carefully thought out how misalignment will play out, and these 8 pathways cover the possibilities we can think of. For each of these 8 pathways, we have a strong case that either it's unlikely to happen, or if it were to happen we'd detect it and/or have mitigation mechanisms in place before it causes harm.
Claim 8: is low: risk from deliberately engineered misalignment is low
Inserting a capable backdoor that evades detection, given the current training process, is difficult. Furthermore, it's unlikely for an actor to be motivated to deliberately engineer misalignment that causes catastrophic harm in the first place, since actors are motivated by their self-interest.
Overall Thoughts
Reasoning about unknown risk is hard! I'm glad that Anthropic has made an effort to formalize and quantify the risks. That being said, many of the argument feel hand-wavy, and rest on reasoning like "our best guess" or "we have not observed the evidence we'd expect if things were to go wrong." Perhaps this is the nature of empirical research?