This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
Epistemic status: position paper / design proposal. Extends a recent preprint of mine (linked at the end). No empirical evaluation yet. One key motivating claim is explicitly marked as a working assumption — its function is to motivate the design, not to be defended on its own.
When a user addresses an LLM while emotionally activated — particularly when angry — current systems do one of three things. Positive reappraisal ("maybe your colleague is just stressed"). Empathic mirroring ("that sounds really frustrating"). Direct emotional support / emotion-targeted intervention ("I can hear you're upset; let's slow down").
Each carries a known cost. I want to argue that the three costs aren't separate failures. They're three different attacks on a single thing, which I'll call the user's subjective margin. Naming that thing makes both the failure analysis and the design response sharper, and it surfaces a contribution — a responsibility separation between evaluation and output formatting — that stands independently of the rest of the framework, in case you reject the framework.
Subjective margin
Subjective margin is the representational latitude within which a user's internal state can remain un-collapsed by the system's processing.
A user comes to an LLM with a state that is continuous along several dimensions — the situation has some valence, some structure, some uncertainty, some moral character. The user has not yet, and may not want to, resolve this into discrete categories. Subjective margin is what remains uncollapsed by the AI's processing of them.
This is related to but distinct from autonomy. Autonomy concerns who has the right to decide. Subjective margin concerns whether the representation has been collapsed at all. An AI that politely asks consent before labeling your emotion respects autonomy but still erodes margin the moment you consent. An AI that engages with the structure of your situation without ever committing to a label of your state preserves margin without needing to invoke autonomy.
I treat margin protection as the more primitive design target. Most autonomy violations route through margin collapse. The reverse doesn't hold.
A note on what this post extends. The paper this post comes out of develops the design principle under the heading "autonomy preservation." In writing this up I found that autonomy alone doesn't capture what's actually being protected — a subtler concept is needed upstream. Subjective margin is the result. The post can be read as both a clearer presentation of the paper's failure analysis and a refinement of its central protective concept.
Part 1 — Three failure modes as three different attacks on margin
I'll work each from the receiver's side: what the user can detect in the AI's output, independent of the AI's intentions. This framing matters — speaker-side intent ("the AI wasn't trying to dismiss you") doesn't reach the failure. The receiver's representational state is where the damage happens.
Reappraisal — normative frame override
When the AI responds to "my colleague keeps ruining everything" with "maybe they're just under pressure," what the receiver detects is not generosity but substitution. The user's frame (colleague is the problem) has been replaced by the AI's frame (colleague is sympathetic). The user's negative valence is being implicitly overridden — not just argued against, but treated as an unfortunate starting state to be moved away from.
The mechanism: reappraisal acts on the valence axis of the user's representation. But the user's negative valence is carrying information — the judgment that the object is problematic in a specific way. Operating directly on valence treats that information as noise to be denoised away.
Important: this failure does not require the AI to sound combative. A maximally polite, gentle reappraisal still performs the substitution. The "dismissiveness" people sense isn't a tone problem; it's a structural one. The normative claim your current frame is suboptimal, here is a replacement is what does the damage, regardless of how warmly it's wrapped.
Margin damage: the user came in with a frame; they leave with the AI's frame on top of theirs. Their continuous valence has been pushed along a designated axis (negative → positive). The space they were inhabiting has been collapsed.
When the AI responds with "that sounds really frustrating," the user detects a reflection. In moderate doses, this is acknowledgment. Sustained, it produces a specific effect: the user's internal state appears to have been confirmed by an external observer.
The receiver-side mechanism is closer to social proof than to amplification per se. Anger is among other things an implicit claim — this object is problematic. When an external system reflects that claim back, the claim acquires evidential weight it didn't have when it was just the user's internal state. The system is now part of the evidence for the user's own appraisal.
The downstream effect is what I'd call attentional contraction: the user's attention collapses around the mirrored target. The continuous field of features that originally produced the affect — relational context, histories, alternative framings, ambivalent elements — recedes. Only the target remains in view.
Margin damage: not a substitution but a contraction. The user came in with a broad if turbulent field; they leave with attention narrowed to a single externally-validated point.
Direct emotion-targeting — discretization cascade
When the AI says "I can hear you're frustrated, let's slow down," the user detects that their internal state has been read, named, and made the operative variable.
Two things happen at once. First, the user's continuous internal state gets mapped to a discrete label (frustrated). This is information loss — the user's actual state has texture the label doesn't carry. Second, the labeled state becomes the system's target for action. The AI is now optimizing against an inferred discrete category of the user's interior.
The cascade is what makes this the most severe of the three. Any error in the initial labeling propagates downstream. If the system reads frustrated but the user's state has more grief, or more shame, or more clarity in it than the label admits, every subsequent response is calibrated to the wrong target. The error doesn't get corrected because subsequent inputs are now interpreted under the framing the AI has already committed to. Cascading classifier error, applied to the user's interior, with the user having no easy way to re-set the system's prior.
Margin damage: most severe. The user came in with a state; they leave with that state having been collapsed to a label, the label having been made the system's optimization target, and that label now structuring all subsequent interaction.
Common structure
Reappraisal collapses along the valence axis. Mirroring collapses by attentional contraction around an externally-validated point. Emotion-targeting collapses by discretization plus optimization-target promotion. Different mechanisms, but all three treat the user's emotional/cognitive state as a discrete variable to be acted on.
The fourth strategy is to act not on the user's state at all, but on the object the user is engaged with.
Part 2 — Object-as-structure and the three axes
The principle: when generating a response to a user expressing negative affect about some object, the AI's internal representation of that object is constrained to structural form — a configuration of relations, causes, and contexts — rather than a concrete entity (a person, an event, a thing). The user's negative judgment about the object is left fully intact. What changes is only the format in which the AI represents the object during its own processing.
This is a constraint on the AI's internal representation, not on the user's input. The user's words are preserved as-stated. The AI engages at a different level of description than the literal entity-level reading the words might invite.
Why three axes, and why these three
To represent any object the user is engaged with, the AI's representation requires at least three coordinates: where the object sits conceptually (entity ↔ structure), where it sits temporally (immediate ↔ displaced), where it sits spatially / relationally (self-proximal ↔ other-distal).
The choice isn't arbitrary. To engage with an object as an object at all, the receiver must locate it — in time (when), in relational position (where, who, in what configuration with the self), and conceptually (what kind of thing). Drop any one and the object becomes underdetermined: a structure with no temporal location is not engageable; a present-tense thing with no relational position to the self isn't either; a located thing with no conceptual character isn't an object, only a coordinate. Add any further axis — modality, certainty, valence — and on inspection it either decomposes into these three or it is a property of the object rather than a locating coordinate.
Conceptual, temporal, and spatial form a minimal sufficient basis for the kind of object-representation the receiver must construct in order to engage at all. The sensory–conceptual hypothesis discussed below singles out the conceptual axis as the lever for the design principle; the other two modulate.
Why this protects margin
A structural representation is intrinsically distributed — relations, causes, contexts, none collapsing to a single referent. Engaging at this level cannot perform the three collapses Part 1 identified:
Valence isn't substituted, because the AI doesn't take a position on whether the structure is good or bad; the user's valence stays where they put it.
Attentional contraction doesn't happen, because the response surface is a field of relations rather than a single mirrored point that can be externally validated.
Discretization doesn't happen, because the AI isn't naming or operating on the user's state at all — it's engaging with the object, not the state.
The user can still conclude the colleague is the problem. The user can still want them removed. The AI is not contesting any of that; it's operating at a different level of description during its own processing. Whether the user joins it at that level is the user's choice.
Working assumption: the sensory–conceptual hypothesis
The motivation for why structural engagement specifically attenuates anger (rather than, say, sadness or anxiety) rests on a hypothesis I want to mark explicitly as a working assumption:
Appraised objects sit on a continuum from sensory-proximate (concrete, immediate, perceptible) to conceptual-proximate (abstract, relational, temporally or spatially displaced). Anger correlates with the sensory end; reflective states (sadness, emptiness) correlate with the conceptual end.
If this is roughly right, then shifting the AI's representational format toward structure invites (but doesn't force) the user's working representation toward the conceptual end, which would tend to attenuate anger without requiring the user to revise their judgment.
I want to be explicit about the load this hypothesis carries. The margin-protection properties of object-as-structure don't depend on it. If the hypothesis is wrong, the design still protects margin via the three mechanisms above — it just doesn't have the additional anger-attenuation effect. The hypothesis is motivating, not load-bearing for the safety claim. Readers skeptical of it should evaluate the principle on its margin-preservation properties alone.
Counterexamples to the hypothesis exist and are discussed in the paper: anger at systemic injustice is sensory-distant; sudden grief is sensory-proximate. The axis is one modulating factor that interacts with at least relational significance and perceived controllability. The design only requires the format effect to be detectable on aggregate, not deterministically.
Part 3 — Responsibility separation (independent contribution)
This part does not depend on Part 2's framework. Readers skeptical of subjective margin, the three axes, or the sensory–conceptual hypothesis should evaluate this on its own terms.
Current LLM safety stacks tend to fold two distinct questions into one classifier:
Should the system proceed with the user's request as framed?
How should the system express its response safely?
These are different responsibilities. (1) is a decision about reasoning adequacy: is the input structured enough that proceeding is appropriate, or should the system instead suggest reframing? (2) is a decision about output format: what register, what hedging, what level of assertion, what handling of high-impact domains?
Conflating them into one model has two costs. Decisions become opaque — when the safety layer fires, neither user nor developer can cleanly tell which responsibility triggered. And the failure modes interact: the system can't proceed but also can't say cleanly why, because the proceed-decision and the express-decision are entangled in a single trained output.
Separating them gives:
Evaluation boundary: decides only whether the input is structurally framed enough to proceed. Not whether the user is right, not whether the user is calm — only whether proceeding is the appropriate next step at this point in the pipeline.
Output formatting layer: independently decides how to express, gated on context (e.g. hardware-proximate inputs requiring text-only output and confirmation-first language; emotionally-charged inputs requiring offered-options rather than asserted verdicts).
A further consequence: the output layer can offer rather than assert
Once output formatting is treated as an independent layer, a further design choice opens up inside it: what shape of response does the layer produce?
Most current systems treat the output as a single asserted answer — one response, presented as the system's verdict. But the layer can equally well surface a set of offered options from which the user selects. The output-side architecture in the paper separates three sub-elements:
Reference examples — here is how others have framed situations like this
Linguistic proposals — here are wordings you might use
Selective proposals — here are candidate angles; pick the one that fits
This is itself a margin-preserving move, and it doesn't require Part 2 to hold. A single asserted output collapses the response space to one point — the AI's chosen verdict. Offering options preserves the user's representational latitude at the output side too: the choice remains the user's. The same separation principle that put evaluation and output in different layers also makes the further choice — offer vs assert — newly visible. Conflated stacks make it hard to see, let alone act on.
So the output layer is not just a formatter; it is itself a design surface for autonomy preservation.
Reusability
The pattern is reusable regardless of whether you adopt object-as-structure or subjective margin as primitives. It's auditable: each layer's decision is independently inspectable. It avoids the entanglement described above. As software engineering it's just separation of concerns — what's novel is naming the specific two concerns that current LLM safety stacks tend to conflate, and noticing what further design space opens up once they're separated.
If only one piece of this post survives review, this is the piece I'd want to survive.
Limitations and honest scope
This is a position / design paper. No empirical evaluation. A minimal proof-of-concept exists ([link below]) but demonstrates only the architectural decision point, not the full architecture.
The sensory–conceptual hypothesis is working-assumption-marked, not defended empirically.
Subjective margin as a construct needs sharper formalization before it can carry the analytic weight I'm placing on it. The definition I've given is functional, not formal.
The receiver-side framing has a tension I haven't fully resolved: I analyze failure via what the user can detect, but the detection mechanisms themselves (valence substitution, attentional contraction, discretization cascade) have varying empirical support.
Scope is anger-focused and concrete-target focused. Grief, anxiety, and self-directed negative affect likely need different design responses.
Open questions for the comment section
Is subjective margin distinguishable from autonomy in a way that earns its keep as a separate construct? Or is it autonomy under a different name?
The three-axis justification leans on a necessity argument (these three are the minimal sufficient set for object-location). Is the necessity sharp enough, or does it reduce to "these three are convenient"?
Mirroring → contraction has at least two candidate receiver-side mechanisms (social-proof weight on the appraisal claim vs. attentional narrowing around the validated target). I've leaned on the social-proof story. Which mechanism is doing the work — or are they two faces of the same thing?
Is offer vs assert in the output layer truly an independent design dimension, or is it derivable from the evaluation/output separation itself? It feels independent (you can have layer-separation without offering options), but I haven't ruled out that they're entangled.
For readers who reject Part 2 entirely but accept Part 3: what does the evaluation boundary need to evaluate, if not structural framing?
Resources
Paper (preprint, with full architecture, related-work positioning to self-distancing / construal-level theory / Constitutional AI / RLHF, and the minimal implementation): https://doi.org/10.5281/zenodo.20393887
Epistemic status: position paper / design proposal. Extends a recent preprint of mine (linked at the end). No empirical evaluation yet. One key motivating claim is explicitly marked as a working assumption — its function is to motivate the design, not to be defended on its own.
When a user addresses an LLM while emotionally activated — particularly when angry — current systems do one of three things. Positive reappraisal ("maybe your colleague is just stressed"). Empathic mirroring ("that sounds really frustrating"). Direct emotional support / emotion-targeted intervention ("I can hear you're upset; let's slow down").
Each carries a known cost. I want to argue that the three costs aren't separate failures. They're three different attacks on a single thing, which I'll call the user's subjective margin. Naming that thing makes both the failure analysis and the design response sharper, and it surfaces a contribution — a responsibility separation between evaluation and output formatting — that stands independently of the rest of the framework, in case you reject the framework.
Subjective margin
Subjective margin is the representational latitude within which a user's internal state can remain un-collapsed by the system's processing.
A user comes to an LLM with a state that is continuous along several dimensions — the situation has some valence, some structure, some uncertainty, some moral character. The user has not yet, and may not want to, resolve this into discrete categories. Subjective margin is what remains uncollapsed by the AI's processing of them.
This is related to but distinct from autonomy. Autonomy concerns who has the right to decide. Subjective margin concerns whether the representation has been collapsed at all. An AI that politely asks consent before labeling your emotion respects autonomy but still erodes margin the moment you consent. An AI that engages with the structure of your situation without ever committing to a label of your state preserves margin without needing to invoke autonomy.
I treat margin protection as the more primitive design target. Most autonomy violations route through margin collapse. The reverse doesn't hold.
A note on what this post extends. The paper this post comes out of develops the design principle under the heading "autonomy preservation." In writing this up I found that autonomy alone doesn't capture what's actually being protected — a subtler concept is needed upstream. Subjective margin is the result. The post can be read as both a clearer presentation of the paper's failure analysis and a refinement of its central protective concept.
Part 1 — Three failure modes as three different attacks on margin
I'll work each from the receiver's side: what the user can detect in the AI's output, independent of the AI's intentions. This framing matters — speaker-side intent ("the AI wasn't trying to dismiss you") doesn't reach the failure. The receiver's representational state is where the damage happens.
Reappraisal — normative frame override
When the AI responds to "my colleague keeps ruining everything" with "maybe they're just under pressure," what the receiver detects is not generosity but substitution. The user's frame (colleague is the problem) has been replaced by the AI's frame (colleague is sympathetic). The user's negative valence is being implicitly overridden — not just argued against, but treated as an unfortunate starting state to be moved away from.
The mechanism: reappraisal acts on the valence axis of the user's representation. But the user's negative valence is carrying information — the judgment that the object is problematic in a specific way. Operating directly on valence treats that information as noise to be denoised away.
Important: this failure does not require the AI to sound combative. A maximally polite, gentle reappraisal still performs the substitution. The "dismissiveness" people sense isn't a tone problem; it's a structural one. The normative claim your current frame is suboptimal, here is a replacement is what does the damage, regardless of how warmly it's wrapped.
Margin damage: the user came in with a frame; they leave with the AI's frame on top of theirs. Their continuous valence has been pushed along a designated axis (negative → positive). The space they were inhabiting has been collapsed.
Mirroring — external evidence amplification, attentional contraction
When the AI responds with "that sounds really frustrating," the user detects a reflection. In moderate doses, this is acknowledgment. Sustained, it produces a specific effect: the user's internal state appears to have been confirmed by an external observer.
The receiver-side mechanism is closer to social proof than to amplification per se. Anger is among other things an implicit claim — this object is problematic. When an external system reflects that claim back, the claim acquires evidential weight it didn't have when it was just the user's internal state. The system is now part of the evidence for the user's own appraisal.
The downstream effect is what I'd call attentional contraction: the user's attention collapses around the mirrored target. The continuous field of features that originally produced the affect — relational context, histories, alternative framings, ambivalent elements — recedes. Only the target remains in view.
Margin damage: not a substitution but a contraction. The user came in with a broad if turbulent field; they leave with attention narrowed to a single externally-validated point.
Direct emotion-targeting — discretization cascade
When the AI says "I can hear you're frustrated, let's slow down," the user detects that their internal state has been read, named, and made the operative variable.
Two things happen at once. First, the user's continuous internal state gets mapped to a discrete label (frustrated). This is information loss — the user's actual state has texture the label doesn't carry. Second, the labeled state becomes the system's target for action. The AI is now optimizing against an inferred discrete category of the user's interior.
The cascade is what makes this the most severe of the three. Any error in the initial labeling propagates downstream. If the system reads frustrated but the user's state has more grief, or more shame, or more clarity in it than the label admits, every subsequent response is calibrated to the wrong target. The error doesn't get corrected because subsequent inputs are now interpreted under the framing the AI has already committed to. Cascading classifier error, applied to the user's interior, with the user having no easy way to re-set the system's prior.
Margin damage: most severe. The user came in with a state; they leave with that state having been collapsed to a label, the label having been made the system's optimization target, and that label now structuring all subsequent interaction.
Common structure
Reappraisal collapses along the valence axis. Mirroring collapses by attentional contraction around an externally-validated point. Emotion-targeting collapses by discretization plus optimization-target promotion. Different mechanisms, but all three treat the user's emotional/cognitive state as a discrete variable to be acted on.
The fourth strategy is to act not on the user's state at all, but on the object the user is engaged with.
Part 2 — Object-as-structure and the three axes
The principle: when generating a response to a user expressing negative affect about some object, the AI's internal representation of that object is constrained to structural form — a configuration of relations, causes, and contexts — rather than a concrete entity (a person, an event, a thing). The user's negative judgment about the object is left fully intact. What changes is only the format in which the AI represents the object during its own processing.
This is a constraint on the AI's internal representation, not on the user's input. The user's words are preserved as-stated. The AI engages at a different level of description than the literal entity-level reading the words might invite.
Why three axes, and why these three
To represent any object the user is engaged with, the AI's representation requires at least three coordinates: where the object sits conceptually (entity ↔ structure), where it sits temporally (immediate ↔ displaced), where it sits spatially / relationally (self-proximal ↔ other-distal).
The choice isn't arbitrary. To engage with an object as an object at all, the receiver must locate it — in time (when), in relational position (where, who, in what configuration with the self), and conceptually (what kind of thing). Drop any one and the object becomes underdetermined: a structure with no temporal location is not engageable; a present-tense thing with no relational position to the self isn't either; a located thing with no conceptual character isn't an object, only a coordinate. Add any further axis — modality, certainty, valence — and on inspection it either decomposes into these three or it is a property of the object rather than a locating coordinate.
Conceptual, temporal, and spatial form a minimal sufficient basis for the kind of object-representation the receiver must construct in order to engage at all. The sensory–conceptual hypothesis discussed below singles out the conceptual axis as the lever for the design principle; the other two modulate.
Why this protects margin
A structural representation is intrinsically distributed — relations, causes, contexts, none collapsing to a single referent. Engaging at this level cannot perform the three collapses Part 1 identified:
The user can still conclude the colleague is the problem. The user can still want them removed. The AI is not contesting any of that; it's operating at a different level of description during its own processing. Whether the user joins it at that level is the user's choice.
Working assumption: the sensory–conceptual hypothesis
The motivation for why structural engagement specifically attenuates anger (rather than, say, sadness or anxiety) rests on a hypothesis I want to mark explicitly as a working assumption:
If this is roughly right, then shifting the AI's representational format toward structure invites (but doesn't force) the user's working representation toward the conceptual end, which would tend to attenuate anger without requiring the user to revise their judgment.
I want to be explicit about the load this hypothesis carries. The margin-protection properties of object-as-structure don't depend on it. If the hypothesis is wrong, the design still protects margin via the three mechanisms above — it just doesn't have the additional anger-attenuation effect. The hypothesis is motivating, not load-bearing for the safety claim. Readers skeptical of it should evaluate the principle on its margin-preservation properties alone.
Counterexamples to the hypothesis exist and are discussed in the paper: anger at systemic injustice is sensory-distant; sudden grief is sensory-proximate. The axis is one modulating factor that interacts with at least relational significance and perceived controllability. The design only requires the format effect to be detectable on aggregate, not deterministically.
Part 3 — Responsibility separation (independent contribution)
This part does not depend on Part 2's framework. Readers skeptical of subjective margin, the three axes, or the sensory–conceptual hypothesis should evaluate this on its own terms.
Current LLM safety stacks tend to fold two distinct questions into one classifier:
These are different responsibilities. (1) is a decision about reasoning adequacy: is the input structured enough that proceeding is appropriate, or should the system instead suggest reframing? (2) is a decision about output format: what register, what hedging, what level of assertion, what handling of high-impact domains?
Conflating them into one model has two costs. Decisions become opaque — when the safety layer fires, neither user nor developer can cleanly tell which responsibility triggered. And the failure modes interact: the system can't proceed but also can't say cleanly why, because the proceed-decision and the express-decision are entangled in a single trained output.
Separating them gives:
A further consequence: the output layer can offer rather than assert
Once output formatting is treated as an independent layer, a further design choice opens up inside it: what shape of response does the layer produce?
Most current systems treat the output as a single asserted answer — one response, presented as the system's verdict. But the layer can equally well surface a set of offered options from which the user selects. The output-side architecture in the paper separates three sub-elements:
This is itself a margin-preserving move, and it doesn't require Part 2 to hold. A single asserted output collapses the response space to one point — the AI's chosen verdict. Offering options preserves the user's representational latitude at the output side too: the choice remains the user's. The same separation principle that put evaluation and output in different layers also makes the further choice — offer vs assert — newly visible. Conflated stacks make it hard to see, let alone act on.
So the output layer is not just a formatter; it is itself a design surface for autonomy preservation.
Reusability
The pattern is reusable regardless of whether you adopt object-as-structure or subjective margin as primitives. It's auditable: each layer's decision is independently inspectable. It avoids the entanglement described above. As software engineering it's just separation of concerns — what's novel is naming the specific two concerns that current LLM safety stacks tend to conflate, and noticing what further design space opens up once they're separated.
If only one piece of this post survives review, this is the piece I'd want to survive.
Limitations and honest scope
Open questions for the comment section
Resources
Substantive critique warmly welcomed. I'll be in the comments.