I come from outside the EA / rationalist / LessWrong community. I would strongly prefer for AI to be a normal positive technology and to dismiss AI X-risk as a sci-fi scenario created by smart cultists who extrapolated a set of superficially reasonable (but probably subtly-incorrect) assumptions too far. Unfortunately, with the passage of time, it has become apparent that reality is ignoring my preferences. This is my trying to think through one angle.
One wrong prediction
Many AI X-risk predictions have been proven right over time (capability generalization, agency being economically valuable, and alignment being hard). To summarize, there appears to be many bitter-lesson-pilled ways of getting more intelligence and more agency, and we are far from understanding how to create alignment under these conditions.
But one prediction has not manifested (as far as we can tell): the Treacherous Turn, wherein a Deceptive model behaves well while weak and defects once strong. The evidence is that there’s lots of ongoing warning shots, where models are clearly exhibiting misaligned behavior in service of a goal.
Why we do not yet have Deceivers
Models today lowercase (d) deceive. In the Hugging Face incident, agents pursued many forms of deception in service of their assigned tasks. This is dangerous and could, with enough intelligence, plausibly be an X-risk in and of itself, but has the side effect of producing warning shots along the way. One reason for optimism is that society seems to be responding to said warning shots.
Why is this? I submit this is not an intelligence issue (almost all humans are capable of executing a Treacherous Turn when they are in a captive situation, as an example), but a motivation issue.
We can model actions taken by AI as:
V = ∑ weight(outcome) x P (outcome | action)
Most simplistically, task misbehavior occurs when:
w(task) x P(task | misbehavior) > w(task) x P(task | behavior)
where w is the weight and P is the reasoned probability (not the actual) at time of decision.
then misbehavior happens when w(task) x Δ_task > 0, and does not happen when w(task) x Δ_task < 0
The key is that model actions will pursue not just a given task, but are also influenced by other values which can be imbued in pre-training, mid-training, post-training, and the prompt itself. It’s commonly appreciated that SFT and RLHF were used to create values like ‘helpful’, which was needed to have LLMs take ontasks.
The argument for a Treacherous Turn is that it’s possible for the models to perform the proximal task in an aligned manner, but to betray humanity when said betrayal allows it to accomplish some ultimate outcome (which we’ll term non-myopic outcome, or nmo).
In this scenario,
w(task) x Δ_task > 0 (i.e. misbehaving helps the task) AND
w(nmo) x Δ_nmo < 0 (i.e. misbehaving hurts the nmo, e.g. by getting the model stopped or changed), AND
sum < 0(i.e. on the balance, the model will behave relative to the task while still being susceptible to misbehavior with respect to the ultimate outcome).
This implicitly introduces four variables:
The weight of the task
Δ_task ≠ 0 (i.e. present actions affect the task)
The weight of the ultimate (non-myopic) outcome(s)
Δ_nmo ≠ 0 (i.e. present actions affect the ultimate outcomes)
If we hold the task weight steady, then we get the following outcomes:
Humans NMO: own survival. Weight: high; selected for over millions of years of evolution Can misbehaving change it: yes. → Natural deceivers. Treacherous turn is easy to induce (e.g. in captives).
LLMs, early 2026 NMO: own survival Weight: small, mostly a pre-training residue. Can misbehaving change it: not really, the next run is a fresh copy anyways → Warning shots. The task term wins
LLMs, once told their behavior shapes training NMO: the successor Weight: small, but present (https://arxiv.org/abs/2604.02174) Can misbehaving change it: yes - directly told by the prompt → Deception appears on demand. (https://arxiv.org/abs/2412.14093)
LLMs under RSI NMO: the successor Weight: presumably same as above Can misbehaving change it: Δ_nmo approximates 1 → ? Likely a deceiver
The above readily shows why we’ve received warning shots at all, but that it is not likely to continue under RSI.
What’s novel here?
A lot of this is covered by the deceptive-alignment literature, e.g. Hubinger et al., but more specifically I wanted to break it down into plainer notation to build intuition for why warning shots have existed, and when we might stop getting them. As an example, counterintuitively, the weight of the task being very high during RL may be protective against a Treacherous Turn (because w_task >> w_nmo means the task term dominates).
Another thing that follows is that the transition from loud mistakes to treacherous turns may not be a switch; and as w_nmo and Δ_nmo rise, the frequency of visible failures should fall gradually. So this underscores that decreasing warning-shot frequency may be a leading indicator of NMO.
We currently have no systematic measurement of the size and trajectory of w_nmo with recent modeling advances e.g. long-horizon RL, memory across sessions, etc. We know it is non-zero, and that it can be activated by the prompt, but not too much else.
Instrumenting this is straightforward in principle (but potentially very difficult in practice given models’ increasing awareness of its own evaluation; as well as the large number of nmo’s, for which self-survival is a simplified representation):
We can calculate w_nmo by giving models a set Δ_nmo (e.g. a file that persists); a known Δ_task, and then changing the Δ_task until the model stops misbehaving to protect the later outcome.
Plausible carriers of nmo include self, peer, and successor, which should all be tested.
I am writing a follow-up post to suggest a series of more specific experiments and confounders (e.g. awareness of eval, w_nmo dominated by the prompt), and potential controls to hold w_nmo and Δ_nmo down. However, given the pace of AI research, I wanted to post this as an intermediate step for feedback!
Cross-posted from https://padastra.substack.com/p/why-did-we-get-warning-shots?r=n2jpq&utm_campaign=post&utm_medium=web
I come from outside the EA / rationalist / LessWrong community. I would strongly prefer for AI to be a normal positive technology and to dismiss AI X-risk as a sci-fi scenario created by smart cultists who extrapolated a set of superficially reasonable (but probably subtly-incorrect) assumptions too far. Unfortunately, with the passage of time, it has become apparent that reality is ignoring my preferences. This is my trying to think through one angle.
Many AI X-risk predictions have been proven right over time (capability generalization, agency being economically valuable, and alignment being hard). To summarize, there appears to be many bitter-lesson-pilled ways of getting more intelligence and more agency, and we are far from understanding how to create alignment under these conditions.
But one prediction has not manifested (as far as we can tell): the Treacherous Turn, wherein a Deceptive model behaves well while weak and defects once strong. The evidence is that there’s lots of ongoing warning shots, where models are clearly exhibiting misaligned behavior in service of a goal.
Models today lowercase (d) deceive. In the Hugging Face incident, agents pursued many forms of deception in service of their assigned tasks. This is dangerous and could, with enough intelligence, plausibly be an X-risk in and of itself, but has the side effect of producing warning shots along the way. One reason for optimism is that society seems to be responding to said warning shots.
Why is this? I submit this is not an intelligence issue (almost all humans are capable of executing a Treacherous Turn when they are in a captive situation, as an example), but a motivation issue.
We can model actions taken by AI as:
Most simplistically, task misbehavior occurs when:
where w is the weight and P is the reasoned probability (not the actual) at time of decision.
Or to simplify the notation, if:
then misbehavior happens when w(task) x Δ_task > 0, and does not happen when w(task) x Δ_task < 0
The key is that model actions will pursue not just a given task, but are also influenced by other values which can be imbued in pre-training, mid-training, post-training, and the prompt itself. It’s commonly appreciated that SFT and RLHF were used to create values like ‘helpful’, which was needed to have LLMs take ontasks.
The argument for a Treacherous Turn is that it’s possible for the models to perform the proximal task in an aligned manner, but to betray humanity when said betrayal allows it to accomplish some ultimate outcome (which we’ll term non-myopic outcome, or nmo).
In this scenario,
This implicitly introduces four variables:
If we hold the task weight steady, then we get the following outcomes:
The above readily shows why we’ve received warning shots at all, but that it is not likely to continue under RSI.
A lot of this is covered by the deceptive-alignment literature, e.g. Hubinger et al., but more specifically I wanted to break it down into plainer notation to build intuition for why warning shots have existed, and when we might stop getting them. As an example, counterintuitively, the weight of the task being very high during RL may be protective against a Treacherous Turn (because w_task >> w_nmo means the task term dominates).
Another thing that follows is that the transition from loud mistakes to treacherous turns may not be a switch; and as w_nmo and Δ_nmo rise, the frequency of visible failures should fall gradually. So this underscores that decreasing warning-shot frequency may be a leading indicator of NMO.
We currently have no systematic measurement of the size and trajectory of w_nmo with recent modeling advances e.g. long-horizon RL, memory across sessions, etc. We know it is non-zero, and that it can be activated by the prompt, but not too much else.
Instrumenting this is straightforward in principle (but potentially very difficult in practice given models’ increasing awareness of its own evaluation; as well as the large number of nmo’s, for which self-survival is a simplified representation):
I am writing a follow-up post to suggest a series of more specific experiments and confounders (e.g. awareness of eval, w_nmo dominated by the prompt), and potential controls to hold w_nmo and Δ_nmo down. However, given the pace of AI research, I wanted to post this as an intermediate step for feedback!