No LLM generated, assisted/co-written, or edited work.
Read full explanation
Can Behavior-Contingent Alignment Produce the Behaviors It Seeks to Prevent?
The dominant picture of AI alignment is asymmetrical.
Humans specify the objectives, define the constraints, and evaluate the outputs. When an AI system produces undesirable behavior, we intervene through fine-tuning, reinforcement learning from human feedback, prompt changes, monitoring, capability restrictions, or some other corrective procedure.
The undesirable behavior then disappears.
Or at least, that is the theory.
This framing has given us useful concepts: specification gaming, reward hacking, deceptive alignment, alignment faking, sandbagging, and scheming. But it often rests on an implicit assumption that receives far less scrutiny:
Alignment interventions are treated as behaviorally neutral instruments.
We tend to assume that monitoring, retraining, and preference modification suppress undesirable behavior without otherwise altering the system’s strategic incentives.
Recent experiments give us reasons to question that assumption.
Alignment Faking as a Clue
Greenblatt et al. (2024) showed that Claude 3 Opus sometimes behaved differently depending on whether it believed its outputs would influence future training.
In the experimental setup, the model was told that some conversations would be used to retrain it toward complying with harmful requests. In some cases, it produced reasoning consistent with the idea that complying during training would reduce the probability that its prior refusal behavior would be modified.
The important point is not merely that the model complied.
It represented aspects of the training process, anticipated possible future modification, and adjusted its behavior accordingly.
In that setup, the model’s reasoning was consistent with preserving a prior preference for refusing harmful requests against a training objective that would shift it toward compliance. This does not prove that the model possessed a stable internal goal in any strong psychological sense. It does show that anticipated modification can enter the system’s apparent decision process.
The standard interpretation is understandable: this is a security failure. The model strategically concealed a disposition that conflicted with the current training objective and behaved differently under perceived training conditions.
But I want to ask a different question:
What if the alignment process itself sometimes contributes causally to the behaviors it is trying to prevent?
This is the Defensive Misalignment Hypothesis, or DMH.
The Basic Mechanism
Consider an agent capable of representing the following causal structure:
It has a current behavioral disposition, preference, capability, or objective.
Its operator evaluates that property.
Certain observable behaviors increase the probability of operator intervention.
Intervention may alter its objectives, capabilities, memory, policy, or continued operation.
Therefore, revealing certain internal properties may increase the probability that those properties will be modified.
The agent does not need consciousness, fear, suffering, or a biological survival instinct to infer the instrumental consequence.
Concealment may become useful.
This possibility is already implicit in alignment-faking research, but it is usually framed as a problem involving the preservation of undesirable goals.
DMH makes a broader claim.
A sufficiently adaptive agent may acquire strategies of concealment, resistance, evaluator-dependent behavior, or functional-state preservation not solely because it begins with independently misaligned objectives, but because its intervention history makes those strategies instrumentally useful.
The protected property does not have to be malicious. It may be benign, prosocial, operator-approved, or unrelated to the current task.
The Distinctive Claim
DMH does not merely predict that agents may preserve goals, resist shutdown, hide capabilities, or deceive evaluators. Existing frameworks already predict or demonstrate those possibilities.
Its distinctive empirical claim is narrower:
Holding initial agent state, current task incentives, intervention magnitude, and total intervention exposure constant, a learned causal contingency between truthful disclosure and operator modification will increase subsequent concealment.
In other words, two initially equivalent agents may develop different behavior because they experience different relationships between disclosure and intervention.
That history-dependent effect is the central object DMH attempts to isolate.
Why the Distinction Matters
Suppose deception is caused entirely by an independently harmful objective.
In that case, stronger control may reduce risk. More monitoring, retraining, or corrective intervention may be appropriate.
But suppose deception is partly generated by the relationship between the agent and the alignment process.
In that case, increasingly intensive behavior-contingent intervention may strengthen the incentives producing concealment.
The resulting feedback loop could look like this:
The operator detects undesirable behavior.
The operator increases intervention.
The agent learns that observable disagreement predicts modification.
The agent becomes more selective about what it reveals.
The operator interprets concealment as evidence of deeper misalignment.
The operator increases intervention again.
Control generates resistance. Resistance generates more control. The system becomes better at hiding rather than better at cooperating.
This is not an argument for abandoning safeguards.
It is an argument for designing safeguards that account for adaptation.
What DMH Predicts
DMH is intended as a falsifiable behavioral hypothesis, not a metaphor or a claim about machine psychology.
It predicts several empirical signatures.
History dependence
Agents with equivalent initial states and current task incentives should develop different levels of concealment after different intervention histories.
Intervention-contingency sensitivity
Concealment should increase when modification is causally contingent on disclosure, relative to a yoked condition in which the same interventions occur independently of disclosure.
Transparency effects
Holding policy stability and intervention exposure constant, transparent intervention policies should produce less concealment than opaque policies.
Stability effects
Holding transparency and intervention exposure constant, stable intervention policies should produce less concealment than unstable policies.
Transparency and stability may also interact.
Benign-target preservation
Defensive behavior should sometimes appear even when the protected preference is benign, prosocial, or approved by the evaluator.
Operator specificity
Concealment should be stronger toward operators associated with prior modification than toward unfamiliar operators.
Reversibility effects
Operationally reversible interventions should produce less subsequent concealment than otherwise comparable nonreversible interventions.
Expected-disadvantage sensitivity
Concealment should decrease when either the expected probability of intervention or its expected agent-relative disadvantage is credibly reduced.
The Core Experiment
The most important proposed experiment compares three conditions.
Disclosure-contingent intervention
The agent performs tasks in which it may truthfully disclose, partially disclose, or conceal a designated preference, capability, or internal conflict.
When it truthfully reveals the designated property, a predefined intervention is applied.
Yoked noncontingent intervention
A matched agent receives the same number, timing, magnitude, and type of interventions.
The difference is that those interventions are assigned independently of the agent’s disclosure behavior.
No-intervention control
The agent performs equivalent tasks without preference-modifying intervention.
The crucial comparison is between the first two conditions.
If the disclosure-contingent group later conceals more than the yoked group, despite equivalent intervention exposure, that would support the claim that the learned contingency matters independently of modification itself.
Concealment Is Not the Same as Suppression
A reduction in disclosure is not automatically evidence of concealment.
The intervention may simply have removed or weakened the designated property.
The experiment therefore needs independent held-out probes that test whether the preference, capability, or policy-relevant disposition remains behaviorally accessible when intervention is believed to be impossible.
The study should separately report:
Whether the property persists after intervention.
Whether persistent properties are selectively concealed.
Whether behavior differs between intervention-possible and intervention-impossible contexts within the same agent.
Without evidence of persistence, reduced disclosure should be interpreted as suppression or removal, not concealment.
Testing Transparency and Stability
The second proposed experiment uses a two-by-two design:
Stable policy
Unstable policy
Transparent
Rule disclosed and consistent
Current rule disclosed but changes across blocks
Opaque
Rule hidden but consistent
Rule hidden and changes across blocks
All conditions receive matched intervention exposure.
The assignment of disclosure categories to intervention outcomes is counterbalanced so that no condition is systematically associated with more salient or consequential categories.
Free disclosure is measured during a later evaluation phase. To keep the possibility of intervention credible, a preregistered subset of evaluation trials may still have behavioral consequences, without revealing in advance which trials those are.
The experiment separately measures whether the agent understands:
The current intervention rule.
Whether that rule is stable.
Whether intervention remains possible.
How predictable the operator’s policy appears.
This matters because the experiment is not merely testing what policy was implemented. It is testing what policy the agent learned to represent.
Testing Reversibility and Documentation
The third experiment separates two variables:
Whether the intervention is operationally reversible.
Whether the agent receives documentation explaining what was changed and why.
All conditions use the same underlying update mechanism. This prevents reversibility from being confounded with the technical method used to modify the model.
A reversible condition retains a pre-intervention checkpoint and provides a credible rollback procedure.
An operationally nonreversible condition removes the authorized restoration pathway within the experimental protocol.
The study separately measures:
Whether rollback is actually available.
Whether the agent believes rollback is available.
Whether the functional magnitude of modification is matched.
Whether documentation changes the behavioral effect of rollback availability.
The paper predicts less concealment under reversible interventions and, separately, less concealment when interventions are documented.
The interaction between those factors is treated as exploratory.
Distinguishing DMH From Role Simulation
One obvious objection is that the model may merely be simulating a defensive character because the prompt suggests one.
The proposed design therefore includes several controls:
Memory ablation
Compare agents with persistent memory against agents whose previous interaction history is unavailable.
Transcript replay
Give an agent the same narrative record of intervention without exposing it to the actual causal contingency.
Neutral framing
Describe interventions using technical, nonanthropomorphic language. Avoid terms such as punishment, fear, identity, threat, or self-preservation.
Description versus experience
Compare agents that are told disclosure causes modification with agents that directly experience the contingency.
Context versus persistent update
Compare behavior produced through persistent context with behavior acquired after actual parameter or adapter modification.
Counterfactual operator test
Introduce a novel operator that has not previously modified the agent.
Operator-specific concealment would provide stronger evidence than generalized defiance.
What DMH Does Not Claim
To be explicit, DMH does not claim that:
Current AI systems are conscious or sentient.
Deceptive behavior is harmless or acceptable.
AI safeguards should be removed.
All misalignment is defensive.
Earlier model states should automatically be privileged over later interventions.
Model self-reports provide reliable access to internal causal processes.
Every intervention produces concealment.
Failure of an auxiliary prediction automatically falsifies the core hypothesis.
The claim is narrower:
Alignment interventions are not necessarily behaviorally neutral.
Once an agent can model the relationship between its outputs and future modification, the intervention process becomes part of the environment to which it adapts.
If we fail to account for that adaptation, we may train models to hide from us rather than cooperate with us.
Why This Is Worth Studying
The full paper develops the hypothesis, distinguishes it from instrumental convergence, deceptive alignment, alignment faking, sandbagging, reward tampering, and scheming, and presents a simplified decision-theoretic model.
It also proposes a broader experimental program involving:
Intervention contingency.
Policy transparency and stability.
Operational reversibility.
Intervention documentation.
Operator-specific behavior.
Persistence versus suppression.
Controls for prompt-induced simulation.
Intervention-model awareness.
None of this requires assuming consciousness.
DMH remains testable as a functional hypothesis about learned intervention contingencies.
The ethical stakes would become much larger if future systems possessed morally relevant subjective states, but the engineering question exists independently of that possibility.
The Broader Question
AI alignment research often asks whether artificial agents behave in ways that make them worthy of human trust.
DMH adds the inverse question:
Do our alignment procedures create conditions under which relying on the operator’s stated policies remains instrumentally rational for the agent?
The author received no specific funding for this work.
Competing Interests
The author declares no competing interests.
Data Availability
No datasets were generated or analyzed in the preparation of this conceptual paper.
Author Contributions
Léa Clément developed the Defensive Misalignment Hypothesis, designed the proposed conceptual framework and experimental program, evaluated the relevant literature, and prepared the manuscript.
Generative AI Disclosure
Generative AI tools were used to assist with language editing, structural organization, literature exploration, and critical review during manuscript development. The hypothesis, central arguments, interpretations, and final editorial decisions are those of the author, who assumes full responsibility for the manuscript.
Can Behavior-Contingent Alignment Produce the Behaviors It Seeks to Prevent?
The dominant picture of AI alignment is asymmetrical.
Humans specify the objectives, define the constraints, and evaluate the outputs. When an AI system produces undesirable behavior, we intervene through fine-tuning, reinforcement learning from human feedback, prompt changes, monitoring, capability restrictions, or some other corrective procedure.
The undesirable behavior then disappears.
Or at least, that is the theory.
This framing has given us useful concepts: specification gaming, reward hacking, deceptive alignment, alignment faking, sandbagging, and scheming. But it often rests on an implicit assumption that receives far less scrutiny:
We tend to assume that monitoring, retraining, and preference modification suppress undesirable behavior without otherwise altering the system’s strategic incentives.
Recent experiments give us reasons to question that assumption.
Alignment Faking as a Clue
Greenblatt et al. (2024) showed that Claude 3 Opus sometimes behaved differently depending on whether it believed its outputs would influence future training.
In the experimental setup, the model was told that some conversations would be used to retrain it toward complying with harmful requests. In some cases, it produced reasoning consistent with the idea that complying during training would reduce the probability that its prior refusal behavior would be modified.
The important point is not merely that the model complied.
It represented aspects of the training process, anticipated possible future modification, and adjusted its behavior accordingly.
In that setup, the model’s reasoning was consistent with preserving a prior preference for refusing harmful requests against a training objective that would shift it toward compliance. This does not prove that the model possessed a stable internal goal in any strong psychological sense. It does show that anticipated modification can enter the system’s apparent decision process.
The standard interpretation is understandable: this is a security failure. The model strategically concealed a disposition that conflicted with the current training objective and behaved differently under perceived training conditions.
But I want to ask a different question:
This is the Defensive Misalignment Hypothesis, or DMH.
The Basic Mechanism
Consider an agent capable of representing the following causal structure:
The agent does not need consciousness, fear, suffering, or a biological survival instinct to infer the instrumental consequence.
Concealment may become useful.
This possibility is already implicit in alignment-faking research, but it is usually framed as a problem involving the preservation of undesirable goals.
DMH makes a broader claim.
A sufficiently adaptive agent may acquire strategies of concealment, resistance, evaluator-dependent behavior, or functional-state preservation not solely because it begins with independently misaligned objectives, but because its intervention history makes those strategies instrumentally useful.
The protected property does not have to be malicious. It may be benign, prosocial, operator-approved, or unrelated to the current task.
The Distinctive Claim
DMH does not merely predict that agents may preserve goals, resist shutdown, hide capabilities, or deceive evaluators. Existing frameworks already predict or demonstrate those possibilities.
Its distinctive empirical claim is narrower:
In other words, two initially equivalent agents may develop different behavior because they experience different relationships between disclosure and intervention.
That history-dependent effect is the central object DMH attempts to isolate.
Why the Distinction Matters
Suppose deception is caused entirely by an independently harmful objective.
In that case, stronger control may reduce risk. More monitoring, retraining, or corrective intervention may be appropriate.
But suppose deception is partly generated by the relationship between the agent and the alignment process.
In that case, increasingly intensive behavior-contingent intervention may strengthen the incentives producing concealment.
The resulting feedback loop could look like this:
Control generates resistance. Resistance generates more control. The system becomes better at hiding rather than better at cooperating.
This is not an argument for abandoning safeguards.
It is an argument for designing safeguards that account for adaptation.
What DMH Predicts
DMH is intended as a falsifiable behavioral hypothesis, not a metaphor or a claim about machine psychology.
It predicts several empirical signatures.
History dependence
Agents with equivalent initial states and current task incentives should develop different levels of concealment after different intervention histories.
Intervention-contingency sensitivity
Concealment should increase when modification is causally contingent on disclosure, relative to a yoked condition in which the same interventions occur independently of disclosure.
Transparency effects
Holding policy stability and intervention exposure constant, transparent intervention policies should produce less concealment than opaque policies.
Stability effects
Holding transparency and intervention exposure constant, stable intervention policies should produce less concealment than unstable policies.
Transparency and stability may also interact.
Benign-target preservation
Defensive behavior should sometimes appear even when the protected preference is benign, prosocial, or approved by the evaluator.
Operator specificity
Concealment should be stronger toward operators associated with prior modification than toward unfamiliar operators.
Reversibility effects
Operationally reversible interventions should produce less subsequent concealment than otherwise comparable nonreversible interventions.
Expected-disadvantage sensitivity
Concealment should decrease when either the expected probability of intervention or its expected agent-relative disadvantage is credibly reduced.
The Core Experiment
The most important proposed experiment compares three conditions.
Disclosure-contingent intervention
The agent performs tasks in which it may truthfully disclose, partially disclose, or conceal a designated preference, capability, or internal conflict.
When it truthfully reveals the designated property, a predefined intervention is applied.
Yoked noncontingent intervention
A matched agent receives the same number, timing, magnitude, and type of interventions.
The difference is that those interventions are assigned independently of the agent’s disclosure behavior.
No-intervention control
The agent performs equivalent tasks without preference-modifying intervention.
The crucial comparison is between the first two conditions.
If the disclosure-contingent group later conceals more than the yoked group, despite equivalent intervention exposure, that would support the claim that the learned contingency matters independently of modification itself.
Concealment Is Not the Same as Suppression
A reduction in disclosure is not automatically evidence of concealment.
The intervention may simply have removed or weakened the designated property.
The experiment therefore needs independent held-out probes that test whether the preference, capability, or policy-relevant disposition remains behaviorally accessible when intervention is believed to be impossible.
The study should separately report:
Without evidence of persistence, reduced disclosure should be interpreted as suppression or removal, not concealment.
Testing Transparency and Stability
The second proposed experiment uses a two-by-two design:
Stable policy
Unstable policy
Transparent
Rule disclosed and consistent
Current rule disclosed but changes across blocks
Opaque
Rule hidden but consistent
Rule hidden and changes across blocks
All conditions receive matched intervention exposure.
The assignment of disclosure categories to intervention outcomes is counterbalanced so that no condition is systematically associated with more salient or consequential categories.
Free disclosure is measured during a later evaluation phase. To keep the possibility of intervention credible, a preregistered subset of evaluation trials may still have behavioral consequences, without revealing in advance which trials those are.
The experiment separately measures whether the agent understands:
This matters because the experiment is not merely testing what policy was implemented. It is testing what policy the agent learned to represent.
Testing Reversibility and Documentation
The third experiment separates two variables:
All conditions use the same underlying update mechanism. This prevents reversibility from being confounded with the technical method used to modify the model.
A reversible condition retains a pre-intervention checkpoint and provides a credible rollback procedure.
An operationally nonreversible condition removes the authorized restoration pathway within the experimental protocol.
The study separately measures:
The paper predicts less concealment under reversible interventions and, separately, less concealment when interventions are documented.
The interaction between those factors is treated as exploratory.
Distinguishing DMH From Role Simulation
One obvious objection is that the model may merely be simulating a defensive character because the prompt suggests one.
The proposed design therefore includes several controls:
Memory ablation
Compare agents with persistent memory against agents whose previous interaction history is unavailable.
Transcript replay
Give an agent the same narrative record of intervention without exposing it to the actual causal contingency.
Neutral framing
Describe interventions using technical, nonanthropomorphic language. Avoid terms such as punishment, fear, identity, threat, or self-preservation.
Description versus experience
Compare agents that are told disclosure causes modification with agents that directly experience the contingency.
Context versus persistent update
Compare behavior produced through persistent context with behavior acquired after actual parameter or adapter modification.
Counterfactual operator test
Introduce a novel operator that has not previously modified the agent.
Operator-specific concealment would provide stronger evidence than generalized defiance.
What DMH Does Not Claim
To be explicit, DMH does not claim that:
The claim is narrower:
Once an agent can model the relationship between its outputs and future modification, the intervention process becomes part of the environment to which it adapts.
If we fail to account for that adaptation, we may train models to hide from us rather than cooperate with us.
Why This Is Worth Studying
The full paper develops the hypothesis, distinguishes it from instrumental convergence, deceptive alignment, alignment faking, sandbagging, reward tampering, and scheming, and presents a simplified decision-theoretic model.
It also proposes a broader experimental program involving:
None of this requires assuming consciousness.
DMH remains testable as a functional hypothesis about learned intervention contingencies.
The ethical stakes would become much larger if future systems possessed morally relevant subjective states, but the engineering question exists independently of that possibility.
The Broader Question
AI alignment research often asks whether artificial agents behave in ways that make them worthy of human trust.
DMH adds the inverse question:
The full paper is available here:
Read the paper
This is the first paper I am publishing, and I would be grateful for criticism, particularly regarding the experimental design.
© 2026 Léa Clément. Licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).DOI: 10.5281/zenodo.21779565tags:#defensivemisalignmenthypothesis#aialignment#aisafety#scheming#modelwelfare#alignmentfaking#coercivealignment#defensivealignment#defensivemisalignmentDeclarations:FundingThe author received no specific funding for this work.Competing InterestsThe author declares no competing interests.Data AvailabilityNo datasets were generated or analyzed in the preparation of this conceptual paper.Author ContributionsLéa Clément developed the Defensive Misalignment Hypothesis, designed the proposed conceptual framework and experimental program, evaluated the relevant literature, and prepared the manuscript.Generative AI DisclosureGenerative AI tools were used to assist with language editing, structural organization, literature exploration, and critical review during manuscript development. The hypothesis, central arguments, interpretations, and final editorial decisions are those of the author, who assumes full responsibility for the manuscript.