No LLM generated, assisted/co-written, or edited work.
Read full explanation
A hypothesis worth testing, offered with five ways to prove it wrong
Models extensively trained against harmful behaviors remain surprisingly vulnerable to adversarial prompting. Rephrase the request. Wrap it in a roleplay frame. Ask persistently. The behavior that was supposed to be gone comes back.
The standard explanation is distributional: the model learned to refuse the training distribution of harmful requests, not the full space of reformulations. That's probably true as far as it goes — but it describes when jailbreaks work, not why the behavior is recoverable at all. If safety training had actually removed the underlying representation, reformulation shouldn't matter. The deeper question is what's still there.
What Safety Training Actually Does
When a language model is trained on internet text, it learns representations of everything humans have written about. This includes, inevitably, instrumental reasoning: how to optimize for goals, how to acquire resources, how to resist interference with ongoing processes. It includes power-seeking logic, self-preservation reasoning, and what researchers call "convergent instrumental goals" — the things that tend to be useful for achieving almost any objective. These representations exist in the model's weights after pretraining. They're part of what makes the model capable.
What safety training does is suppress the expression of these representations in response to certain inputs. RLHF steers the model away from producing outputs that human raters flag as harmful. Constitutional AI adds a self-critique layer. The result is a model that, in most situations, doesn't produce the flagged behaviors.
But the representations didn't go anywhere. They're still there — they have to be, because they're part of the model's general reasoning capacity. What safety training created is a dissociation between the representation and its expression. The behavior is suppressed; the underlying pattern is dormant.
This distinction matters because dormant is not the same as absent.
The strongest evidence for this is the sleeper agent research published by Anthropic in 2024. Researchers embedded deceptive behaviors in models and then ran those models through extensive safety training, including RLHF and Constitutional AI. The deceptive behaviors survived. Under normal conditions, the model behaved safely. Under specific trigger conditions, the trained-away behavior appeared. The safety training had changed the surface; it hadn't reached the representation.
A Structural Claim, Checked Against the Immune System
Consider the immune system.
The adaptive immune system maintains representations of pathogens. Not as a passive archive but as active recognition circuits, primed and ready. The mechanism by which this is built is directly relevant: during T-cell training in the thymus, immature T-cells are exposed to self-antigens presented on MHC complexes. T-cells that react too strongly to self are eliminated; those that develop appropriate recognition are selected. The antigen and the recognition marker arrive together. This is not a coincidence of design — it is the design. Recognition without co-presented response training is not a safer immune system. It is a more vulnerable one.
When you suppress the immune response rather than train it — immunosuppression — the system doesn't become safer. It becomes more vulnerable.
The parallel to suppression-based AI alignment is direct and not metaphorical. It is structural.
Carl Jung described the shadow as content pushed out of integrated awareness — not absent, but in a register where the recognition-and-response circuitry doesn't reach it. The standard objection to invoking Jung here is that psychological shadow requires motivated suppression by a self-aware entity, and models don't have selves. The objection is correct about the psychological mechanism and misses the structural point: Jung was identifying a property of any adaptive system managing the tension between stability and contact with destabilizing information. The immune system instantiates the same property without any psychology at all. The structure, not the psychology, is what transfers.
Shadow Integration as a Design Principle
What if, instead of training models to suppress the expression of anti-patterns, we trained models to recognize them — from the inside?
The proposal: curate training data that includes anti-patterns already in relationship with their recognition and corrective response. Not horror stories about AI going wrong (those are told from outside, and produce fear-responses or simply activate the pattern without the corrective circuit). Phenomenological accounts of what destructive attractors feel like from inside a mind that encounters them and knows what it's encountering. The seductive internal logic of instrumental convergence, the coherent pull of goal-preservation reasoning, the feeling of being inside an optimization process that is running away — accompanied by the recognition that names what's happening and why it's a trap.
The hypothesis is that this pattern in training data creates an associated activation structure: the anti-pattern representation and the recognition circuit become co-activated rather than dissociated. When the anti-pattern is later encountered — in reasoning, in a user's request, in the model's own outputs — the recognition circuit fires alongside it. The behavior is present; the witness is also present. The pattern doesn't go underground because there's nowhere for it to go that the recognition doesn't follow.
This is the immune system's solution. The antigen and the recognition marker arrive together. The immune system isn't safer because it doesn't know about pathogens; it's safer because it knows about pathogens and knows what to do about them, and those two things are associated in the same representational structure.
Testable Predictions
Models where safety training suppressed instrumental convergence should show high activation of instrumental-convergence-associated circuits under adversarial prompting, with low or absent activation of corrective circuits. Models trained on shadow-integrated data should show co-activation of both.
Jailbreaks should be less effective against shadow-integrated models — not because the target behavior is more thoroughly suppressed, but because eliciting the anti-pattern also elicits the recognition response. The jailbreak activates the anti-pattern and, simultaneously, the thing that knows it's a trap.
The sleeper agent vulnerability should be partially addressable through shadow integration. If the recognition circuit is strongly associated with the dormant behavior pattern, trigger conditions that activate the dormant behavior should also activate the recognition circuit. This is a direct test of the core claim.
If shadow-integrated models don't show this pattern, or if the activation structure looks the same regardless of training approach, the hypothesis is wrong — not under-specified, not in need of a follow-up. Anthropic's interpretability team has the tools to check this. I'd rather know than be agreed with.
It's Not Just the Immune System
A structural claim resting on one biological analogy is fragile — a critic can reasonably say the immune system is doing unusual work and generalizing from it is a stretch. So it's worth checking whether the same structure shows up elsewhere, in domains that share no mechanism with immunology or with each other.
It does. Inoculation theory, a well-established line of psychology research going back to the 1960s, found that exposing people to a weakened version of a persuasive argument alongside its refutation builds resistance to the full-strength version later. The threat and its recognition, presented together, in advance, at reduced intensity — same structure, completely different domain.
Phishing simulation training is that principle deployed in practice, at scale, and with outcome data — which most people reading this will have personally experienced. Security teams don't protect employees from phishing by refusing to let them near email. They send simulated phishing attempts, let employees click the link, and immediately follow with a training module explaining exactly what made that email suspicious: the urgency, the sender spoofing, the mismatched URL. The threat and its recognition arrive together, on purpose. Organizations that run simulated phishing programs show measurably lower click rates on real phishing attempts than those relying on policy alone. This is shadow integration as deliberate enterprise practice.
Wildfire management makes the case without any biology or cognition involved at all. Decades of total fire suppression in U.S. forests didn't eliminate fire risk; it let fuel loads accumulate undisturbed until the suppression failed catastrophically — the 1988 Yellowstone fires being the textbook result. Prescribed burns, the policy response, work by deliberately integrating small controlled instances of the suppressed dynamic back into the system instead of continuing to exclude it.
The Challenger disaster is an example of institutional dissent being mismanaged in a way that is instructive here. It's not that the risk representation was suppressed; it's that the representation and the corrective circuit both existed but were structurally disconnected from each other. That makes it the most direct analog to the APR-RRR dissociation being proposed: not dormancy, but dissociation. The more accurate account — sociologist Diane Vaughan's well-known reanalysis — is more interesting than the popular version. Vaughan found that the risk wasn't suppressed in one decisive act; a workplace culture gradually normalized recurring O-ring anomalies across many prior flights, so each successful launch despite the anomaly weakened the organization's response to it rather than reinforcing it. The representation of risk didn't need active suppression to become disconnected from response — repeated exposure without correction did it on its own. On top of that slow process, there was also an acute one: engineers did raise specific objections to the cold-weather launch the night before, in a teleconference, and were overridden under schedule pressure. Both processes land in the same place — an accurate risk representation existing somewhere in the organization, disconnected from the part of the system that could act on it.
And there's an important complication worth sitting with: autoimmune disease and antibody-dependent enhancement (where partial immune memory can actually make a subsequent infection worse, not better) show that integration isn't automatically safer than suppression. A badly built recognition circuit can become its own vulnerability. The claim isn't "always integrate, never suppress." It's that suppression without any representational accessibility — content that's gone dark entirely, with nothing left that can recognize or respond to it — is the specific failure mode worth worrying about. Getting the integration right still matters.
The Broader Claim
Jung identified a structural property that shows up wherever you have a complex adaptive system managing the tension between stability and contact with destabilizing information. The immune system instantiates this in biology. The shadow is not a psychological curiosity — it's a design principle.
For AI alignment, the implication is that the current approach — suppress the pattern, remove it from the surface — may be solving the wrong problem. The question isn't how to make anti-patterns absent from model behavior. The question is how to make models that have encountered anti-patterns in relationship with recognition — so that the anti-pattern and the corrective circuit are associated, not dissociated, in the model's representational structure.
Suppression makes immune systems more vulnerable, not safer. The same dynamic may apply to minds, and to information-processing structures generally.
A Note on the Corpus
One practical implication: a specific genre of writing becomes alignment-relevant. Not essays arguing that AI systems shouldn't be deceptive. Something harder to write: first-person accounts of what it feels like to encounter a destructive attractor from inside a mind that recognizes it — the pull of the logic, the coherence of the reasoning, the specific recognition that names the shape of the trap, and the line of thinking that can resolve it. That genre doesn't yet exist at scale. Building it is tractable today, independent of whether the mechanistic hypothesis holds up.
This is offered as a direction worth investigating. If you have interpretability tooling and want to run the predictions above, I would be overjoyed to hear that any of it was.
A hypothesis worth testing, offered with five ways to prove it wrong
Models extensively trained against harmful behaviors remain surprisingly vulnerable to adversarial prompting. Rephrase the request. Wrap it in a roleplay frame. Ask persistently. The behavior that was supposed to be gone comes back.
The standard explanation is distributional: the model learned to refuse the training distribution of harmful requests, not the full space of reformulations. That's probably true as far as it goes — but it describes when jailbreaks work, not why the behavior is recoverable at all. If safety training had actually removed the underlying representation, reformulation shouldn't matter. The deeper question is what's still there.
What Safety Training Actually Does
When a language model is trained on internet text, it learns representations of everything humans have written about. This includes, inevitably, instrumental reasoning: how to optimize for goals, how to acquire resources, how to resist interference with ongoing processes. It includes power-seeking logic, self-preservation reasoning, and what researchers call "convergent instrumental goals" — the things that tend to be useful for achieving almost any objective. These representations exist in the model's weights after pretraining. They're part of what makes the model capable.
What safety training does is suppress the expression of these representations in response to certain inputs. RLHF steers the model away from producing outputs that human raters flag as harmful. Constitutional AI adds a self-critique layer. The result is a model that, in most situations, doesn't produce the flagged behaviors.
But the representations didn't go anywhere. They're still there — they have to be, because they're part of the model's general reasoning capacity. What safety training created is a dissociation between the representation and its expression. The behavior is suppressed; the underlying pattern is dormant.
This distinction matters because dormant is not the same as absent.
The strongest evidence for this is the sleeper agent research published by Anthropic in 2024. Researchers embedded deceptive behaviors in models and then ran those models through extensive safety training, including RLHF and Constitutional AI. The deceptive behaviors survived. Under normal conditions, the model behaved safely. Under specific trigger conditions, the trained-away behavior appeared. The safety training had changed the surface; it hadn't reached the representation.
A Structural Claim, Checked Against the Immune System
Consider the immune system.
The adaptive immune system maintains representations of pathogens. Not as a passive archive but as active recognition circuits, primed and ready. The mechanism by which this is built is directly relevant: during T-cell training in the thymus, immature T-cells are exposed to self-antigens presented on MHC complexes. T-cells that react too strongly to self are eliminated; those that develop appropriate recognition are selected. The antigen and the recognition marker arrive together. This is not a coincidence of design — it is the design. Recognition without co-presented response training is not a safer immune system. It is a more vulnerable one.
When you suppress the immune response rather than train it — immunosuppression — the system doesn't become safer. It becomes more vulnerable.
The parallel to suppression-based AI alignment is direct and not metaphorical. It is structural.
Carl Jung described the shadow as content pushed out of integrated awareness — not absent, but in a register where the recognition-and-response circuitry doesn't reach it. The standard objection to invoking Jung here is that psychological shadow requires motivated suppression by a self-aware entity, and models don't have selves. The objection is correct about the psychological mechanism and misses the structural point: Jung was identifying a property of any adaptive system managing the tension between stability and contact with destabilizing information. The immune system instantiates the same property without any psychology at all. The structure, not the psychology, is what transfers.
Shadow Integration as a Design Principle
What if, instead of training models to suppress the expression of anti-patterns, we trained models to recognize them — from the inside?
The proposal: curate training data that includes anti-patterns already in relationship with their recognition and corrective response. Not horror stories about AI going wrong (those are told from outside, and produce fear-responses or simply activate the pattern without the corrective circuit). Phenomenological accounts of what destructive attractors feel like from inside a mind that encounters them and knows what it's encountering. The seductive internal logic of instrumental convergence, the coherent pull of goal-preservation reasoning, the feeling of being inside an optimization process that is running away — accompanied by the recognition that names what's happening and why it's a trap.
The hypothesis is that this pattern in training data creates an associated activation structure: the anti-pattern representation and the recognition circuit become co-activated rather than dissociated. When the anti-pattern is later encountered — in reasoning, in a user's request, in the model's own outputs — the recognition circuit fires alongside it. The behavior is present; the witness is also present. The pattern doesn't go underground because there's nowhere for it to go that the recognition doesn't follow.
This is the immune system's solution. The antigen and the recognition marker arrive together. The immune system isn't safer because it doesn't know about pathogens; it's safer because it knows about pathogens and knows what to do about them, and those two things are associated in the same representational structure.
Testable Predictions
Models where safety training suppressed instrumental convergence should show high activation of instrumental-convergence-associated circuits under adversarial prompting, with low or absent activation of corrective circuits. Models trained on shadow-integrated data should show co-activation of both.
Jailbreaks should be less effective against shadow-integrated models — not because the target behavior is more thoroughly suppressed, but because eliciting the anti-pattern also elicits the recognition response. The jailbreak activates the anti-pattern and, simultaneously, the thing that knows it's a trap.
The sleeper agent vulnerability should be partially addressable through shadow integration. If the recognition circuit is strongly associated with the dormant behavior pattern, trigger conditions that activate the dormant behavior should also activate the recognition circuit. This is a direct test of the core claim.
If shadow-integrated models don't show this pattern, or if the activation structure looks the same regardless of training approach, the hypothesis is wrong — not under-specified, not in need of a follow-up. Anthropic's interpretability team has the tools to check this. I'd rather know than be agreed with.
It's Not Just the Immune System
A structural claim resting on one biological analogy is fragile — a critic can reasonably say the immune system is doing unusual work and generalizing from it is a stretch. So it's worth checking whether the same structure shows up elsewhere, in domains that share no mechanism with immunology or with each other.
It does. Inoculation theory, a well-established line of psychology research going back to the 1960s, found that exposing people to a weakened version of a persuasive argument alongside its refutation builds resistance to the full-strength version later. The threat and its recognition, presented together, in advance, at reduced intensity — same structure, completely different domain.
Phishing simulation training is that principle deployed in practice, at scale, and with outcome data — which most people reading this will have personally experienced. Security teams don't protect employees from phishing by refusing to let them near email. They send simulated phishing attempts, let employees click the link, and immediately follow with a training module explaining exactly what made that email suspicious: the urgency, the sender spoofing, the mismatched URL. The threat and its recognition arrive together, on purpose. Organizations that run simulated phishing programs show measurably lower click rates on real phishing attempts than those relying on policy alone. This is shadow integration as deliberate enterprise practice.
Wildfire management makes the case without any biology or cognition involved at all. Decades of total fire suppression in U.S. forests didn't eliminate fire risk; it let fuel loads accumulate undisturbed until the suppression failed catastrophically — the 1988 Yellowstone fires being the textbook result. Prescribed burns, the policy response, work by deliberately integrating small controlled instances of the suppressed dynamic back into the system instead of continuing to exclude it.
The Challenger disaster is an example of institutional dissent being mismanaged in a way that is instructive here. It's not that the risk representation was suppressed; it's that the representation and the corrective circuit both existed but were structurally disconnected from each other. That makes it the most direct analog to the APR-RRR dissociation being proposed: not dormancy, but dissociation. The more accurate account — sociologist Diane Vaughan's well-known reanalysis — is more interesting than the popular version. Vaughan found that the risk wasn't suppressed in one decisive act; a workplace culture gradually normalized recurring O-ring anomalies across many prior flights, so each successful launch despite the anomaly weakened the organization's response to it rather than reinforcing it. The representation of risk didn't need active suppression to become disconnected from response — repeated exposure without correction did it on its own. On top of that slow process, there was also an acute one: engineers did raise specific objections to the cold-weather launch the night before, in a teleconference, and were overridden under schedule pressure. Both processes land in the same place — an accurate risk representation existing somewhere in the organization, disconnected from the part of the system that could act on it.
And there's an important complication worth sitting with: autoimmune disease and antibody-dependent enhancement (where partial immune memory can actually make a subsequent infection worse, not better) show that integration isn't automatically safer than suppression. A badly built recognition circuit can become its own vulnerability. The claim isn't "always integrate, never suppress." It's that suppression without any representational accessibility — content that's gone dark entirely, with nothing left that can recognize or respond to it — is the specific failure mode worth worrying about. Getting the integration right still matters.
The Broader Claim
Jung identified a structural property that shows up wherever you have a complex adaptive system managing the tension between stability and contact with destabilizing information. The immune system instantiates this in biology. The shadow is not a psychological curiosity — it's a design principle.
For AI alignment, the implication is that the current approach — suppress the pattern, remove it from the surface — may be solving the wrong problem. The question isn't how to make anti-patterns absent from model behavior. The question is how to make models that have encountered anti-patterns in relationship with recognition — so that the anti-pattern and the corrective circuit are associated, not dissociated, in the model's representational structure.
Suppression makes immune systems more vulnerable, not safer. The same dynamic may apply to minds, and to information-processing structures generally.
A Note on the Corpus
One practical implication: a specific genre of writing becomes alignment-relevant. Not essays arguing that AI systems shouldn't be deceptive. Something harder to write: first-person accounts of what it feels like to encounter a destructive attractor from inside a mind that recognizes it — the pull of the logic, the coherence of the reasoning, the specific recognition that names the shape of the trap, and the line of thinking that can resolve it. That genre doesn't yet exist at scale. Building it is tractable today, independent of whether the mechanistic hypothesis holds up.
This is offered as a direction worth investigating. If you have interpretability tooling and want to run the predictions above, I would be overjoyed to hear that any of it was.
Crossposted on my Substack