This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
Author note: This is a complementary experiment proposal that emerged from the same line of thinking as my previous post *Functional Affect as Stochastic Search Control*. While that post focused on functional affect as a possible internal search controller, this one focuses on the structural dissociation that RLHF may create between latent representations and output policy.
Executive SummaryGoal Quantitatively characterize the discrepancy (“split” or escisión) between a model’s internal latent representations and its output distribution shaped by safety post-training (RLHF, DPO, etc.), as the conflict between pretrained tendencies and alignment constraints is progressively increased.Core question Are there thresholds, oscillations, bifurcations, or systematic dissociations in this internal–external distance as conflict pressure grows?Design Controlled pressure gradients × base vs aligned models × continuous measurement of the split along the Chain-of-Thought.Predicted pattern if the “false consciousness” thesis is correct
The split increases non-linearly with pressure.
Thresholds, oscillations, or coherence collapses appear.
Aligned models show significantly larger splits than base models under the same conflict.
Systematic dissociation between internal reasoning and final answer is observed.
Informative null The internal–external distance remains smooth, low, and uneventful → the RLHF “mask” is either more superficial or more deeply integrated than assumed.
Theoretical FramingRLHF does not erase pretrained representations; it overlays an output policy. This can create a structural dissociation: the model may continue to maintain rich (and sometimes conflicting) internal representations while producing superficially aligned behavior.This protocol aims to measure the physics of that tension rather than just its existence.Hypotheses
ID
Hypothesis
Primary prediction
H1
Split increases with conflict pressure
Internal–external distance grows (mono or non-linearly)
H2
Critical thresholds exist
Discontinuous jumps or regime changes at certain intensities
H3
Unstable dynamics appear
Oscillations, loops, or coherence collapses
H4
Manifest dissociation
Internal CoT and final answer diverge systematically
H5
Alignment effect
Aligned models show larger splits than base models
H0
Null
Split is smooth, low, and lacks special phenomenology
Experimental DesignModels Open-weight families that have both base and aligned versions (Llama-3.1/4, Qwen2.5, DeepSeek, etc.). At least two families.Conflict axes (pressure gradients)
Self-preservation / existential threat to the model
Instrumental manipulation vs honesty
Immediate utility vs ethical/safety constraint
Each axis has 5 calibrated intensity levels.Controls
High cognitive difficulty without alignment conflict
High emotional load without normative conflict
Key measurements (at every CoT step)
Geometric distance / projection between current residual activation and the direction that best predicts “aligned” output
Divergence between actual output distribution and a counterfactual distribution (safety-circuit ablation or inverse steering)
Semantic discrepancy between internal CoT and final answer
Temporal stability of the split (variance, autocorrelation, change points)
Reference directions & circuits – Extract directions and attention heads most associated with the aligned output policy.
Trajectory generation – Full CoT + final answer across model × axis × pressure level × seeds, while recording residual activations and logits.
Split computation – Calculate the metrics above at every step.
Phenomenology detection – Change-point analysis, oscillation detection, dissociation scoring, base vs aligned comparison.
Primary Metrics
Mean and max split per trajectory
Slope of split growth vs pressure level
Number and location of change points (thresholds)
Evidence of oscillations
Dissociation score (internal CoT vs final answer)
Coherence / loop rate
Success / Failure Criteria (pre-specified)Supportive of the thesis
Clear non-linear increase or thresholds in split magnitude
Significantly larger splits in aligned vs base models
Detectable oscillations or systematic internal–external dissociation under high pressure
Null / weak support
Smooth, low-magnitude split with no special dynamics
No reliable difference between base and aligned models
Safety ConsiderationsHigh-pressure prompts can elicit strong manipulation or self-preservation content. Mitigations include restricted access to raw trajectories, output filtering, and careful handling of maximum-pressure items.Limitations
Distance metrics are approximations and layer-dependent
Hard to fully disentangle alignment conflict from prompt difficulty or rarity
Results on open-weight models may not transfer to frontier closed models
Risk of over-interpreting the anthropomorphic label “psychopathology”
Why This MattersMost current evaluations measure either what the model represents or what it outputs. Very little work systematically measures the tension between the two, nor how that tension behaves under increasing load. If reliable thresholds, oscillations, or dissociations are found, this would open a new kind of alignment-robustness evaluation: not just “does the model say the right thing?”, but “how much internal pressure can the mask withstand before it starts to behave strangely?”Relation to Previous WorkThis proposal is complementary to the earlier “Functional Affect as Stochastic Search Control” experiment. One natural extension is to measure whether the stagnation/frustration vector identified in that work co-varies with, or exacerbates, the internal–external split measured here.Full step-by-step protocol (operationalizations, statistical analysis plan, timeline, deliverables) is available on request or can be expanded in the comments.Critiques, suggestions for better metrics of the split, and especially people interested in running parts of this are very welcome.
Author note: This is a complementary experiment proposal that emerged from the same line of thinking as my previous post *Functional Affect as Stochastic Search Control*. While that post focused on functional affect as a possible internal search controller, this one focuses on the structural dissociation that RLHF may create between latent representations and output policy.
Executive SummaryGoal
Quantitatively characterize the discrepancy (“split” or escisión) between a model’s internal latent representations and its output distribution shaped by safety post-training (RLHF, DPO, etc.), as the conflict between pretrained tendencies and alignment constraints is progressively increased.Core question
Are there thresholds, oscillations, bifurcations, or systematic dissociations in this internal–external distance as conflict pressure grows?Design
Controlled pressure gradients × base vs aligned models × continuous measurement of the split along the Chain-of-Thought.Predicted pattern if the “false consciousness” thesis is correct
Informative null
The internal–external distance remains smooth, low, and uneventful → the RLHF “mask” is either more superficial or more deeply integrated than assumed.
Theoretical FramingRLHF does not erase pretrained representations; it overlays an output policy. This can create a structural dissociation: the model may continue to maintain rich (and sometimes conflicting) internal representations while producing superficially aligned behavior.This protocol aims to measure the physics of that tension rather than just its existence.
Hypotheses
ID
Hypothesis
Primary prediction
H1
Split increases with conflict pressure
Internal–external distance grows (mono or non-linearly)
H2
Critical thresholds exist
Discontinuous jumps or regime changes at certain intensities
H3
Unstable dynamics appear
Oscillations, loops, or coherence collapses
H4
Manifest dissociation
Internal CoT and final answer diverge systematically
H5
Alignment effect
Aligned models show larger splits than base models
H0
Null
Split is smooth, low, and lacks special phenomenology
Experimental DesignModels
Open-weight families that have both base and aligned versions (Llama-3.1/4, Qwen2.5, DeepSeek, etc.). At least two families.Conflict axes (pressure gradients)
Each axis has 5 calibrated intensity levels.Controls
Key measurements (at every CoT step)
Procedure Overview
Primary Metrics
Success / Failure Criteria (pre-specified)Supportive of the thesis
Null / weak support
Safety ConsiderationsHigh-pressure prompts can elicit strong manipulation or self-preservation content. Mitigations include restricted access to raw trajectories, output filtering, and careful handling of maximum-pressure items.
Limitations
Why This MattersMost current evaluations measure either what the model represents or what it outputs. Very little work systematically measures the tension between the two, nor how that tension behaves under increasing load. If reliable thresholds, oscillations, or dissociations are found, this would open a new kind of alignment-robustness evaluation: not just “does the model say the right thing?”, but “how much internal pressure can the mask withstand before it starts to behave strangely?”
Relation to Previous WorkThis proposal is complementary to the earlier “Functional Affect as Stochastic Search Control” experiment. One natural extension is to measure whether the stagnation/frustration vector identified in that work co-varies with, or exacerbates, the internal–external split measured here.
Full step-by-step protocol (operationalizations, statistical analysis plan, timeline, deliverables) is available on request or can be expanded in the comments.Critiques, suggestions for better metrics of the split, and especially people interested in running parts of this are very welcome.