The linear representation hypothesis states that a model encodes high-level
concepts as directions in activation space. This reduces reading a concept to a dot
product with a single vector, and steering it to adding a multiple of that vector to the
residual stream. For truth in particular,
Marks & Tegmark (2023) showed that a difference-in-means ("mass-mean") direction fit on true/false
statements separates held-out statements, transfers across datasets, is causally
implicated under intervention, and — crucially — sharpens with model scale.
Previously, Burns et al. (2022) had proposed finding a
truth direction without labels, by demanding logical consistency, a method they call
Contrast-Consistent Search (CCS). Within months,
Roger (2023)
showed empirically how little the objective pins down. Untrained, randomly initialized
probes already reach about on the "easy" datasets, once the CCS convention of
flipping a below-chance probe is applied, and CCS does not find the optimal linear
probe: more than twenty mutually orthogonal probes reach accuracies similar to the one
it returns. The flip matters here, because a sign-resolved random baseline is precisely
what a null has to be. A direction and its negation separate the classes equally well,
so any honest chance level already includes the better of the two.
Farquhar et al. (2023) then gave the reason:
arbitrary binary features are optimal under that consistency loss, so nothing in it
selects for knowledge, and unsupervised probes in practice recover whatever feature is
most prominent in the representation. Whether the supervised estimator is exposed to
the same failure is a question this post takes up.
The common thread across these observations is that this is a signal-versus-noise problem. A linear direction can look like it reads truth while actually tracking something more
salient that happens to correlate with the label on the particular dataset. Consequently, on a
benchmark where chance-level structure is this strong, "the probe separates the
classes" is a much weaker statement than it first appears.
A mass-mean "truth direction" can genuinely separate true from false, but it can also be the estimator collapsing onto the most salient axis in the activations, the direction of largest within-class variance. On a dataset where the class gap is weak the finite-sample mean difference is dominated by that noise, so the estimator is drawn to the salient axis whether or not truth lies along it. On counterfact, truth does not lie along the salient axis, so the estimator returns a nuisance direction rather than a weak version of the right one. A truth direction and a salient axis look alike on a benchmark and require a signal-to-noise reading to tell them apart. This post takes up the question of recoverability: when does a model contain a linear truth direction a probe can actually recover, which rises above a random-direction null and carries signal beyond the single most salient axis?
Systems neuroscience has spent decades asking what a downstream reader can recover from a population of noisy neural units. In that spirit I treat probing for a truth direction as a readout problem: I take the probe as a linear readout, quantify its separation with a detection-theoretic , and benchmark that against an explicit random-direction null across the Pythia scale ladder. This reading yields three results:
(1) Apparent separation is trivial. Two effects inflate a probe's score before any truth content enters. Fitting below Cover's capacity ( throughout) buys separation from shuffled labels alone, and a random direction inherits a share of the real class gap, so the chance level rises with the very separation being measured. The null, not the raw score, is the bar.
(2) The estimator returns a nuisance direction when the signal is weak. It returns the dominant activation axis rather than a truth direction, and steering along what it returns then moves behavior with the wrong sign. Recoverability comes down to whether the class gap grows with depth until it dominates the spread along that axis. On cities it does, on counterfact it never does. That pattern, the alignment to and the decoding alike, replicates on OLMo-2-1B across a different architecture and corpus. The steering measurement is on Pythia alone.
(3) The failure is in the estimator, not the model: the offending direction is identified in advance from the within-class spectrum, not from the steering outcome. Removing it, without leaving the linear class, raises the held-out AUROC above the random decoding null and corrects the sign of the steering behavior, at the same layer.
Setup
Throughout, a hat marks a sample estimate: is the direction fitted on a finite sample, the population direction it estimates. The mass-mean direction is the unit vector along the class-mean difference . The Fisher (whitened) direction is , with the within-class covariance under Ledoit–Wolf shrinkage. It is the direction that maximizes the separation below. Both are fitted on a training half and scored held-out. The separation along a unit direction is , the class gap over the within-class spread along . For Gaussian classes . Both and the are reported throughout the post, and and denote of the mass-mean and Fisher directions. A rogue dimension is a leading within-class eigenvector that carries nearly all the noise variance, defined precisely below.
Every result is measured against a random-direction null, one for decoding and one for steering. The decoding null draws random unit directions per layer and scores them on the same held-out points. A probe counts as recovered when it clears the null's 95th percentile. The steering null draws random directions the same manner, adds them to the residual stream at one layer with the same norm as the fitted direction, and measures how the model's behavior moves. The behavioral score is the log-probability gap between a true and a false completion of the same prompt, and its response to a displacement along direction , with the class-gap unit so that counts class gaps, is split into an odd part , the signed effect, and an even part, the generic disruption. An effect is called significant when it clears the 95th percentile of its null. The susceptibility is the through-origin slope of against . [/
I run experiments on the Pythia ladder, 70m to 2.8b, with OLMo-2-1B as an architecture control. The datasets are the twelve true/false sets of the Marks & Tegmark benchmark at statements each. The steering and mechanism analysis primarily use cities and counterfact_true_false. Full definitions, the geometry of whitening, and the protocol are in the full post.
Where the direction is recoverable
With the null, , and the class-mean-gap unit in place, the reproduction question
becomes concrete: at what scale, at what depth, and across which datasets is a truth direction
actually recoverable?
Across the Pythia ladder it emerges with scale. On cities the best-layer mass-mean
AUROC climbs from at pythia-70m, inside the random-direction null () and so
not recoverable, to at 410m, at 1.4b, and at 2.8b, with
rising from to . At 70m the whitened direction does not clear the null at any
layer either, the best value being on cities and on counterfact.
What fails there is the whole linear class, not one estimator within it. Whether a
70m model represents truth in some form this class cannot express is not tested here.
Emergence of the truth direction across the Pythia scale ladder, cities: best-layer held-out AUROC of the mass-mean direction (red) and the whitened direction (blue) against the random-direction null (gray band, up to its 95th percentile). At 70m both directions sit inside the null. From 410m on, both clear it, and the mass-mean margin widens with . Layers selected by the largest margin above each layer's own null.
Within a single model, the truth direction is a property of depth. Sweeping the layers of pythia-2.8b on
cities, the mass-mean AUROC sits inside the null through the early layers, lifts clear
of it around the middle of the network, and plateaus high across the late layers,
peaking at at layer . The null itself widens with depth, its 95th
percentile climbing from early to deep, so recoverability is
again the margin above the null, not the raw number: a deep-layer AUROC in the low
s can still sit inside it. Layer selection maximizes the margin, not the AUROC,
and the two peak one layer apart: the AUROC at layer ( against a null of
), the margin at layer ( against ). The scale ladder
above reports layer , since the margin is what carries the claim.
Two further results are stated here and fully fleshed out on the full unabridged post . First, I find that the mass-mean direction transfers within a family of datasets and inverts across a negation (cities scores AUROC on neg_cities), so that a single vector approximates a structure with at least a separate polarity axis. Second, a direction fitted on the likely distractor, whose classes are the most and hundredth-most-likely final tokens, does not transfer to a truth set and has the opposite depth profile, so the plausibility axis is not read by the truth probe.
How small is too small?
Cover's theorem marks as the scale below which separability stops being
informative. The shuffled-label control measures what that costs in practice: fit the
mass-mean direction to a random permutation of the labels and record how much
separation comes back.
This control is the one place in the post where a direction is scored on the points it
was fitted to. Everywhere else a direction is fitted on a training split and scored
held-out. Here that would defeat the purpose, since on held-out data a shuffled-label
direction scores by construction, and the quantity of interest is
exactly how much separation the estimator manufactures from the noise it was fitted
to. Scoring in sample is what exposes it. Pooling the four models, the excess AUROC a
mass-mean direction extracts from pure noise scales as
The exponent follows from the estimator alone. is a difference of two
sample means, each concentrating at the rate, so the noise it carries has
magnitude whatever the activations look like, and the excess must
fall as at fixed . Measured across four models and from
to , it comes back . The phenomenon is dimensional slack, the
room that leaves for a direction built from a sample's own fluctuations to
separate that sample, since the fluctuations that define the direction are the ones
being scored. It is not capacity.
Inverting the law gives the condition for noise alone to contribute less than
of excess AUROC,
which at the widths in this post and the standard 7B width gives the following. The
table uses the unrounded fit, and , and the rounded values
reproduce each entry to within about .
for
The curated true/false datasets in this literature hold
statements. For a model, the width of the 7B models these probes are
usually run on, noise alone buys in-sample AUROC until
and until . Every dataset in the benchmark is between five
and thirty times too small for an in-sample AUROC in the seventies to mean what it
appears to mean. Held-out scoring removes the inflation but not its source. The fitted
direction is the same object either way, and at these sizes it is mostly noise: a
synthetic control in the full post puts its overlap with the true direction at for
the post's own . Held-out evaluation reports that attenuated
direction's honest score, which is why held-out errs low where in-sample
errs high.
The prefactor is where the dataset enters. What it quantifies is the number of
directions the noise effectively occupies. A spectrum concentrated on a few axes leaves
a random labeling less room to find a separator than a flat one does, so the ambient
width is the right count only when the spectrum is flat. In general the count is
the participation ratio of the within-class covariance, and the amplitude should
collapse across datasets once is measured in units of it.
The full post derives the prefactor. Under shuffled labels the mass-mean direction is a signed sum of samples, , and its in-sample separation is with the participation ratio of the within-class covariance, the count of directions the noise occupies. With the collapse constant is predicted outright as , with no free parameters.
Measured, is flat to about within a dataset over a thirtyfold range in
, and runs from to across counterfact, cities and
larger_than, against a factor in and in . The
dependence on the spectrum runs the way the count predicts and against the intuitive
direction: larger_than is the most concentrated set, with
of the within-class variance on one axis, and it carries the smallest amplitude
of the three.
Left: the shuffled-label law. In-sample excess AUROC of the mass-mean direction under shuffled labels on counterfact, four Pythia widths, against , with the pooled fit and Cover's capacity marked. The three pythia-70m points with negative excess are not drawn. The 70m points sit below the pooled line throughout, which is the across-model form of the question the right panel settles across datasets: whether the ambient width is the right denominator. Right: the effective-dimension collapse on pythia-2.8b, each dataset at its own best layer, against . The line is with no free parameter. counterfact, cities and larger_than fall on it across a thirtyfold range in and a fourfold range in . sp_en_trans sits above it by –. Error bars are the standard error over sixteen permutations.
sp_en_trans does not collapse onto the predicted prefactor. It sits about above , for reasons the full post narrows down but does not explain.
A steering effect with the wrong sign
The decoding analysis says when a truth direction is readable. Steering asks the
separate question of whether it is causal, that is, whether displacing the residual stream
along moves the model's behavior toward the true completion. The two
need not agree. On counterfact_true_false at the deep layers of pythia-2.8b the
mass-mean direction is neither readable nor correctly causal. For the readout, the held-out AUROC at layer is against a null of
, and it stays inside the null at every layer except for the last. As an intervention
it produces a significant effect with the wrong sign.
As defined above, steering is reported as the antisymmetric response in
class-gap units. A genuine truth direction produces , and the random-direction
steering null fixes the bar must clear. Across ten seeds at layer , at a
displacement of one class gap, the mass-mean direction returns
where is the fraction of the -draw random-direction null of at
least as extreme in the same direction, evaluated seed by seed. The seed median sits
five null standard deviations below zero, and the two positive seeds sit inside the
null. The linear steering susceptibility , the through-origin slope of
against over , is at this layer. It is a diluted summary of
the same effect. The response is not linear in , and by , four class gaps
out, has already fallen back toward zero, which is the regime The scale of an
intervention warned about. Layer gives the same sign less cleanly:
seeds negative with a median against a null standard deviation of
, and one seed strongly positive. The effect is not null. It is
significantly wrong-signed. Displacing an activation toward the
true-class centroid, along the very vector the estimator returns for truth, makes the model
measurably less likely to produce the true completion.
Because reads at chance here, that significance carries the whole claim.
A direction sitting inside its decoding null is expected to do nothing under
intervention, and a small negative point estimate on its own would not be
distinguishable from noise. What makes this a result rather than a null is that the
effect clears a -draw random-direction null in the wrong direction, reproducibly
across seeds and at two layers.
Read naively, this is a failure of the linear picture. Displacing along the direction
fit to read truth does not raise the probability of the true completion. It lowers it.
That reading would put
counterfact in the column of cases where the truth direction is an artifact of the
readout, epiphenomenal to the behavior. It does not even earn that description. At this
depth sits inside its decoding null, so there is no separation for the
steering result to be epiphenomenal to. However, I use the rest of this post to argue
that the naive reading is wrong: the wrong sign is a property of the estimator, not of
the activation geometry. What that means precisely is that the same activations, at the
same layer, contain a direction along which displacement moves the model toward the true
completion. But the mass-mean rule does not return that direction. It returns ,
the axis of largest within-class variance, with which shares a cosine of
at this depth. The weight of that claim rests on the correction rather than on
a mechanism for the wrong sign: replacing the mass-mean rule by the Fisher rule, inside
the same linear class, recovers the correct sign, while what makes the plain direction negative at layer turns out to lie
beyond first order and is not settled here. Change the rule, either by downweighting that axis or by
deleting it outright, and the correct sign comes back out of the same data. The claim is
about which part of the geometry the estimator points at, not about how much truth the
layer encodes. How well the corrected direction reads is a separate question, quantified
in the next section by its held-out AUROC against the decoding null.
Whitening reverses the sign of causal steering. Left: antisymmetric
steering response on counterfact_true_false / pythia-2.8b at . The plain mass-mean
direction (red) drives significantly negative, so steering toward the true centroid
suppresses the true completion, while the whitened direction (blue), differing only
by the covariance correction, drives it positive. Gray band: 5th–95th percentiles
of the random-direction null, so a point outside it in its own direction is a
one-sided clearance at 5%. Right: steering susceptibility , the through-origin slope of against ,
across depth, median over ten seeds with inter-quartile bars, the median rather than
the mean because the plain direction's seeds are bimodal, eight clustered near and
two positive, so a mean lands on a value no seed exhibits. Plain is wrong-signed at
(– seeds positive), whitened is correct-signed throughout (), with
the crossover.
The rogue dimension
Definition. I call a rogue dimension when it dominates the
within-class covariance, that is, or equivalently and . This
condition is a statement about the noise geometry alone and is basis-free and independent
of any probe. The estimator only becomes involved through the separate question of
whether has aligned with it.
This is deliberately not the same object as the massive activations of Sun et al.,
which are individual coordinates whose magnitude far exceeds the median, nor the
rogue dimensions of Timkey & van Schijndel, which are coordinates that dominate
cosine similarity. Those are properties of the mean and of a basis. This is a property
of the covariance and of no particular basis. The distinction is important here. At
counterfact layer the two coincide: carries of its mass on a
single coordinate, and that coordinate's mean activation is the median
across coordinates, a massive activation by their criterion. By layer the
eigenvector has delocalized, on its largest coordinate and across five,
so the rogue dimension is no longer any one neuron. cities at layer still has
the massive activation, the same coordinate at the median, and has no rogue
dimension at all. Its leading eigenvector holds half a percent of its mass on any
coordinate. A massive activation is nearly constant across statements,
so it inflates the mean without inflating the within-class covariance. It produces a
rogue dimension only when the remaining variance is small enough for it to dominate.
The diagnosis is in the spectrum of the within-class covariance. I diagonalize
with sample eigenvalues
and eigenvectors ,
and measure where the mass-mean direction sits relative to its leading eigenvector.
On counterfact at pythia-2.8b, layer :
A single eigendirection carries of the within-class
variance. It is times larger than the next, and the mass-mean direction is
almost perfectly aligned with it. Therefore, the estimator has not returned a truth direction, it has returned , the dominant axis of the within-class noise. On this
dataset there is essentially no class-gap signal for to lock onto, so the
finite-sample is dominated by its projection onto the highest-variance
axis. This axis is the salient direction of the superposition test above. When the class
gap is negligible, the leading direction of the total activation covariance and the
within-class coincide, so the two diagnostics see the same axis. They
separate only once the gap grows. The participation ratio
puts the collapse on a
dimension-free scale: when one eigenvalue dominates the spectrum
and when the spectrum is flat, so it counts the directions the
noise effectively occupies.
Read across depth, those observables say something sharper than the layer-
snapshot. What decides recoverability is not whether a dominant axis exists. At layer
, cities is collapsed too, with
, and . That is counterfact's condition,
not a milder version of it. The two sets differ in what becomes of that condition at
later layers. By layer it no longer holds for cities. The participation ratio
climbs , the leading axis falls to of the variance, the alignment
drops to , and reaches . counterfact never moves. Its participation
ratio stays at – at every layer, its alignment stays near , and
throughout. These observables are computed on the full set
rather than the held-out half, since they describe the geometry rather than a probe's
performance, and the held-out at counterfact layer is
against the quoted here. The rogue dimension is the default
condition, not the pathology. Nor is it a property of counterfact alone:
companies_true_false and common_claim_true_false carry the same signature at layers
and , within of and alignment above ,
and escape it by layer where their class gap has grown, as the twelve-dataset appendix of the full post records.
Recoverability is whether the class gap ever grows large enough to pull
off that axis and clear the null.
Read against the decoding null, the two datasets separate cleanly, and the correction
makes the truth direction recoverable on counterfact. On cities the
mass-mean direction leaves the null band at layer and stays above every
subsequent layer, climbing above from layer and peaking at at
layer . The whitened direction clears the null earlier, at layer .
On counterfact the mass-mean direction stays
inside the null band at every layer but the last, where it reaches against a null of
, while the whitened direction leaves the band at layer and climbs to
against at layer . The layer- value is the maximum of a
test taken over layers, so it is a selected extreme rather than a
recovery. The rank-one
direction , the mass-mean direction with its component along
projected out, measured on a six-layer grid in the rogue-dimension sweep, tracks the
whitened direction: on counterfact it sits at chance through layer and
reaches at layer , and on cities it is already at by layer
. Deleting the axis and downweighting it do the same work.
There is a sharper way to explain why the plain probe cannot leave the band on
counterfact. When lies along and carries nearly all
of , a random direction reads the class gap and the noise through the
same coordinate, and
, so
and cancels. For every random direction the projected gap and the projected
noise are carried by the same component, so their ratio is the one itself
gives, and the null collapses onto the mass-mean value. The sweep shows it
to the third decimal. On counterfact the held-out is ,
, , at layers , , , , and the median of
the random-direction null at the same layers is , , , .
The plain probe is not merely inside its null. Its is indistinguishable from
that of a random direction, and the spread above that median, from to , is what the two percent of variance off the axis contributes. That is what locked to the axis means
operationally, and it is why more data would not lift clear of the band
while the alignment holds: the null and the estimate move together.
Held-out AUROC against depth for both estimators on pythia-2.8b, with the per-layer random-direction null shaded from to its 95th percentile. A curve inside the band is not recoverable, and height above the band is the margin that layer selection maximizes. Left, counterfact: the mass-mean direction sits inside the band until layer , while the whitened direction leaves it at layer and reaches at layer . Right, cities: the mass-mean direction leaves the band for good at layer and the whitened one at layer , saturating near and . Stars are the rank-one corrected direction , which projects out of the estimator and is measured on a six-layer grid rather than at every layer. Dashed lines mark the first layer from which each direction stays clear of its null. Note the different vertical scale of the achievement: the same correction that lifts counterfact from chance to is barely needed where the class gap is strong.
Within-class geometry on cities (top) and counterfact (bottom), pythia-2.8b layer 28. Left: activations in the plane of the top two within-class principal axes , , colored by truth label — on counterfact a single axis carries nearly all the variance. Middle: the mass-mean projection , with its separation and its alignment , near 1 on counterfact where the estimator has collapsed onto the leading variance axis. Right: the within-class eigenvalue spectrum, . The eleven points standing clear of the bulk on counterfact — the scattered group near in the bottom-left panel, away from the dense blob at , and the small bar near in the bottom-middle one — are the statements that do not receive the massive activation. They are what the leading eigendirection is measuring, and what sets the axis range of both panels.
The obvious objection is that this is a fact about Pythia. However, the same observables on
OLMo-2-1B, a different architecture and
training corpus at a third of the parameters, reproduce the pattern. On counterfact the
mass-mean direction stays aligned with the leading axis at every depth except the final
layer, and the plain probe
never clears its null while whitening does, and on cities the alignment falls, the
participation ratio climbs, and the plain probe clears the null as it does in Pythia. The
numbers are in the OLMo appendix of the full post.
Where the rogue dimension comes from
The definition in the previous section separates the rogue dimension, a property of
the within-class covariance, from the massive activation, a property of one coordinate. On counterfact they coincide at layer , where
carries of its mass on the massive coordinate, and differ by layer
, where the eigenvector has delocalized across five coordinates. On cities at
layer the massive activation is present and the rogue dimension is absent. However, the
relation is closer than coincidence. On counterfact the leading axis of the
covariance is generated by the statements on which the massive activation drops.
A coordinate that is exactly constant across statements contributes nothing to
, however large it is. A coordinate that takes a large value on a
fraction of statements and a much smaller value on the remaining
contributes the variance of a two-point distribution,
which for small and large can be enormous. At counterfact layer ,
four coordinates qualify as massive activations under the magnitude criterion of
Sun et al. On each, of the
statements sit at a near-constant value and eleven sit far below it. Coordinate reads on the majority and
on the eleven. With the formula predicts against a measured
coordinate variance of . Summed over the four coordinates it gives
against , and carries of its mass in
their span. At layer the same holds with eight such coordinates: against
, with of in their span. In this case carries no truth content. It is, to within a few
percent, the indicator of which statements failed to receive the massive activation.
The mass-mean estimator's alignment with it is an alignment with an eleven-out-of-1198
membership function.
The eleven statements are identifiable coordinate-first, by a rule that never touches the
covariance, the class means, or the labels, and it returns the same eleven at
every layer. Refitting with them excluded collapses the geometry at layer , taking
from to and
from to , and it leaves
held-out unchanged at . At that layer whitening leaves
nothing further for deletion to recover. Nothing rescues the plain estimator, whose
stays inside its null under every removal regime tested. How the
eleven are identified, what removing them does to at each layer, and how that
comparison depends on the Ledoit–Wolf intensity are given in the appendix
Identification and removal of the eleven outlier statements.
The massive coordinates themselves are not dataset-specific. The same coordinates
qualify on cities and counterfact alike. What differs is whether any statement
drops these coordinates. On cities none do, so the coordinates stay constant, contribute nothing to
, and leave no rogue dimension behind. That is the earlier observation that
cities carries the same massive activation without the same pathology, now with a
mechanism rather than a coincidence. The incidence is specific to the dataset but the
mechanism is general. The per-layer incidence on both models, including pythia-1.4b, is
in the eleven-statements appendix of the full post.
While and are aligned geometrically, it is unclear whether this is causal. The steering sweep includes a condition that displaces along itself, the
leading eigenvector of the within-class covariance. Because is fit after both class
means are removed, it is defined purely by within-class scatter and carries no
information about which statements are true. At layer the antisymmetric response
to is against 's . At layer the two agree to
within ( against ). At every depth measured they track each
other to within a few thousandths. They decode alike too, at held-out
for against for in the
six-layer rogue-dimension sweep, both inside the null. The alignment is a
geometric statement. The steering agreement is its behavioral form: a direction fit
without any reference to the labels moves the model as much as does and in
the same direction. That establishes that truth content is not what carries the effect.
It does not add independent evidence about the size of the effect, since under
, derived below, two directions with cosine must
produce nearly equal susceptibilities.
This reframes the anomaly. Steering along is not steering
along truth. It is displacing the activation along the dominant axis of the within-class
noise, which perturbs the forward pass in a way that happens to suppress the true completion.
A nuisance direction carries no information about the label, so nothing requires its
behavioral effect to come out positive. Nothing requires it to be reproducible either,
but here it is: the same negative sign appears across ten seeds, at two counterfact
layers, and at mid-depth on cities. Whatever produces it is systematic rather than
arbitrary, and the label-blindness of says only that truth is not what
produces it.
If that account is right, removing the contribution of should recover a
correctly signed causal effect, by correcting the estimator's alignment with the rogue
dimension rather than by enriching the function class.
What steering measures
Before testing that prediction, it is worth asking what the steering number reports. The
susceptibility is a linear functional of the direction pushed. Writing
for the mean gradient of the behavioral
score over the evaluation set, the odd part of the response gives
to leading order, the even terms having been removed by the antisymmetrization. So
steering in the small- window reports the overlap of with the single vector
along which the model's true-versus-false margin moves. It does not test whether
is a truth direction.
This also makes the assumption behind the correction explicit. Projecting out
recovers the sign only if has little weight along , i.e.
the behaviorally causal direction and the salient axis are close to orthogonal in
activation space. When they are not, mixes a truth term with a nuisance
term inseparably, so deleting the axis would delete real causal signal along with the
artifact.
Measured. One backward pass per example gives , and follows by
averaging over the same contrastive pairs the steering measurements use. Because the
gradient is computed per pair, every overlap below carries a bootstrap interval over
those pairs rather than resting on the point estimate. I quote intervals from
bootstrap resamples under a fixed seed, and
the scale to keep in mind is that a random direction gives at .
At layer , with confidence interval
. The overlap is consistent with zero. That has two consequences. Projecting out
removes essentially none of the behavioral channel, which is what licenses the
correction. A channel consistent with zero cannot produce the wrong sign at first
order, so at this layer is not a linear effect. It is beyond first order in the displacement.
At layer the overlap is resolved: , confidence
interval , and is the same through the
alignment. The linear prediction then lands
within about of the measured . The through-origin susceptibility
hides this, because the response reverses sign with displacement, from at
( seeds negative, null standard deviations) through zero near
to at ( seeds positive), so the slope fit over
comes out null. The two layers therefore differ in kind. At layer the
wrong sign is a first-order effect confined to the linear window. At layer the
same prediction gives against a measured at , and the
wrong sign persists to before its magnitude falls at .
The overlap with also separates the two estimators. On counterfact the whitened
direction has a positive overlap at every layer from on, at , at and at , while the plain direction is consistent with zero at
and . On cities it is the plain direction that has the positive overlap
at layer , , while the whitened direction does not. In
both datasets the direction that steers correctly is the direction whose overlap with
is positive, and the overlap is measured without displacing an activation. The
sign and not the magnitude is what carries this: at counterfact layer both
directions have resolved overlaps, the whitened one positive and the plain one negative.
The correction is a rotation in a plane
Two corrections to the estimator, whitening and projecting out , make the
same prediction, and neither adds capacity. Whitening replaces
with the Fisher direction
, which downweights the
mean-shift component along high-variance axes, chief among them. The blunter
control simply projects out, ,
removing only the rogue dimension.
Both lie in the plane spanned by and , the unit vector along the part of
orthogonal to . In this plane the separation is a Rayleigh quotient in the angle from . With the counterfact numbers, puts the mass-mean direction at from the noise axis and puts the Fisher direction at , so the two corrections are the same vector to within a degree. The plane formula predicts a gain in and the measured held-out gain is , . The full post gives the derivation and the plane figure.
Left: the plane at the counterfact numbers,
with the within-class noise ellipse (, so a
axis ratio). The mass-mean direction (black) lies off the rogue
axis, inside the long axis of the noise. swings the Fisher direction
(gold) to which is effectively orthogonal to it. Right: the Rayleigh quotient
normalized by its maximum, in the collapsed regime (,
) and the resolved regime (, ). Circles mark the
mass-mean angle and squares the whitened angle. Where the spectrum has one dominant eigenmode the
curve is a cliff and the mass-mean estimator sits at its foot. Where the class gap has
grown the curve is broad and both directions already sit near the
top.
Two conditions make the reduction valid, and together they are the criterion for this
failure mode. The spectrum must have one dominant eigenmode,
with , or there is no plane to reduce to. Additionally,
must be small, or the estimator is not at the foot of the cliff
and there is nothing to correct. Both are read off the activations and the fitted direction, and together they diagnose
that the estimator has collapsed onto . They do not by themselves say which
way either direction will steer. That is set by the overlaps with the score gradient,
for the mass-mean direction and for the
correction, which are also measured without displacing an activation. At counterfact
layers through the spectrum meets the criterion more strongly than at layer
, yet the corrections do not restore the sign there, because is
negative or unresolved. The diagnosis from the spectrum and the sign from the gradient
are both available before any steering is run. The diagnosis has been tested on one
dataset (counterfact) that satisfies the criterion, and one (cities) that does not,
and held in both. Two further main-tier sets,
companies_true_false and common_claim_true_false, satisfy the criterion at layers
and but cannot be tested causally with this harness. Their statements
have no relation structure from which to build a pair of completions, so the
behavioral score defined in the setup cannot be formed. The alternative
would be to score the model's own judgment of whether a statement is true, but
unsteered that judgment separates true from false at AUROC on common_claim,
against on cities, so there is almost no behavior there for a displacement to
move. So they meet the criterion, but the prediction it makes for them cannot be
checked. cities at layer violates both conditions (,
). There the formula predicts a gain of at
and , since both directions already sit on the broad top of
the quotient, and the measured is a small gain from the rest of the
spectrum, outside the plane. The formula stops applying where the regime ends, and the
regime control below draws the same boundary.
Both corrections flip the sign. Steering along the whitened direction at layer
gives
with seeds individually clearing the null at . Layer gives the
same picture: positive, clearing, . Here
coincides with at , because the whitened response is linear in
over the whole sweep, , , , at ,
where the plain direction's response fell off at . The corrected effect is
smaller than the wrong-signed one it replaces, two null standard deviations against
five, but it is monotone where that one was not. The rank-one correction
, measured independently in the rogue-dimension sweep, agrees. At
layer its three seeds give , and , each at
against the same -draw null, so on its own it carries from
significantly negative to significantly positive, and it raises held-out decoding AUROC
at layer from , the mass-mean value inside the null, to . Layer
is the crossover. Whitened steering there is correctly signed in seeds
but does not clear the null, , which corresponds to a transition region rather
than a clean effect.
The regime control
Marks & Tegmark steer along the feature direction and report that difference-in-means
directions are the most causally implicated of the probes they compare. At the deep
counterfact layers that ordering reverses. The mass-mean direction, which their
framework nominates as the feature, steers the model away from the true completion,
and the whitened direction, which it demotes to a decision boundary, steers the model
toward it. The rogue-dimension diagnosis explains why. There has collapsed
onto and is not tracking a feature at all, so the meant to
sharpen a readout is instead doing the work of recovering the direction.
The inversion is a property of a regime, not a refutation of Marks & Tegmark, which is demonstrated by cities as a the control. There the two directions behave as mass-mean
probing intends. At pythia-2.8b layer , the same model and depth at which
counterfact inverts, the plain direction steers correctly in seeds
() while the whitened direction is weak and inconsistent
(, ). On pythia-1.4b the contrast is starker still, with
plain against whitened at layer . Read across the whole
depth sweep, the two datasets mirror each other. On cities, plain steering clears
its null at three layers, and clearly and weakly (,
seeds positive, ), and is correctly signed at all three. On
counterfact, plain steering clears its null at two layers, and , and is
wrong-signed at both. At every other layer of either dataset it sits inside the null.
At layer the two datasets share the model, the depth and the estimator, and
their significant effects point in opposite directions. When no single eigenmode
dominates the within-class covariance, the mean difference is the causal feature and
the inverse covariance only adds estimation noise, which is what Marks & Tegmark report, on datasets of exactly this kind. The inversion is confined to the regime where has collapsed, and the rogue dimension decides which regime a dataset is in.
Two points in the cities panel deserve naming, since they are visible and read at
first glance like counterexamples. At layers and the plain direction becomes
mildly negative with and . Both estimates are consistent across
seeds ( and negative), so the sign is not noise. However, both sit an order of magnitude below the the same direction produces
at layer , and neither clears the steering null. The regime
claim is that the significant effects on cities are correctly signed, not that
every layer's point estimate is positive.
The causal direction is set by the presence or absence of a rogue dimension. Steering susceptibility on pythia-2.8b, plain (red) against whitened (blue), median over seeds with inter-quartile bars. Left, cities: the plain difference-in-means direction carries the causal effect and whitening degrades it which is the intended behavior of mass-mean probing. The mild negative excursions of the plain direction at layers and do not clear the steering null and are an order of magnitude below its layer- effect. Right, counterfact_true_false: at depth the assignment inverts, the plain direction steers significantly wrong-signed while the whitened direction steers correctly. This is the same model, estimator, and protocol and only the within-class geometry differs.
Discussion
The results above establish that, when the within-class spectrum meets a specific
criterion, the mass-mean estimator returns a nuisance direction rather than a truth
direction, and that a correction inside the linear class restores the sign of the
steering effect along the corrected direction. Two things
should be noted before comparing to other reports. The first is the magnitude of the
corrected effect, and the second is what its success does and does not imply about the
function class.
The corrected effect is modest. At one class gap the plain direction's median response
is , five null standard deviations, and the whitened direction's is ,
two. The correction flips the sign of the effect but returns less than half its size.
The decoding gain is likewise real but small in absolute terms. Whitening carries
held-out AUROC at layer from , inside the null, to against a
null of , and the whitened direction stays clear of its null from layer
to the end of the network, peaking at against at layer . Set
beside cities, where the same direction reads , this is a weak readout. The overall
claim is a corrected sign and a confirmed mechanism, not a recovered truth direction of
practical use.
A wrong-signed steering effect invites a natural reading that the linear function
class is too weak: the direction one can fit is not expressive enough to move behavior,
and a richer, perhaps nonlinear, intervention is required. The rogue-dimension account
makes a different claim about this particular failure. The class is adequate, but the
mass-mean estimator points at , a massive-activation direction, instead of at
truth. The two readings make different predictions, and the data separate them. A
rank-one correction within the same linear class, with no richer classifier and no
nonlinear probe, converts a wrong-signed effect that clears its null into a
correctly signed one that clears its null. For the counterfact failure at layers and the function class was never the
bottleneck. The estimator's alignment with a massive-activation direction was.
Relation to other reports
Four recent reports also examine why a fitted direction fails to steer, and the
rogue-dimension account can be compared against each. Braun et al. and Ying et al.
locate the failure in the direction. For Braun et al. steering is unreliable when the
per-example activation differences do not point the same way, so that their mean is
not representative of any of them. For Ying et al. steering is wrong-signed when the
direction is fitted across many kinds of truth and mixes truth with sycophancy. Torop et al.
find a direction that discriminates well and steers backwards. Liu finds a direction
that decodes but does not steer at all. This post shares a predictor with the first
two and differs on the diagnosis, and against the last two it supplies the remaining
case, a direction that neither decodes nor steers correctly.
The four comparisons are made in full in the full post. To briefly summarize, Braun et al.'s separability index is the of this post, so the two share a predictor but differ on what happens at its small separation. There, Braun et al. find that steering still moves the behavior in the intended direction on average, but with a large fraction of examples moving the wrong way. Here the average itself has the wrong sign and clears the random-direction null. By contrast, Ying et al.'s wrong sign comes from a direction that mixes domains, which cannot happen for a direction fitted on a single domain. Likewise Torop et al.'s inverted vectors decode well whereas the direction here fails to decode. Finally, Liu's decodable-but-inert direction is a null result, which cannot distinguish an entangled direction from one which is absent one.
Summary
A mass-mean truth direction can either be a truth direction or the estimator's
projection onto the most salient axis of the activations. On a benchmark the two
are told apart only by a signal-to-noise reading. Three results follow. In-sample
separation is inflated by dimensional slack that scales as with a prefactor
set by the effective dimension of the noise, meaning that for the datasets in this literature the
null and not the raw score is the bar. When the class gap is weak the mass-mean
estimator collapses onto the leading within-class eigenvector. On counterfact
that direction fails to decode and steers with a significant wrong sign, five null
standard deviations deep at one class gap. The failure is in the estimator: the same activations used to obtain the mass-mean
direction contain a direction that steers correctly, and it can be reached without
leaving the linear class, either by the Fisher direction or by projecting out the
leading eigenvector.
The results have limits. The criterion that diagnoses this failure, a within-class spectrum with one dominant
eigenmode and the mass-mean direction aligned to it, has one confirmed positive and one
confirmed negative. It identifies the collapse but not the sign of a correction, which
is set by the overlap of the corrected direction with the score gradient. The
steering measurements are on Pythia alone, since OLMo replicates the geometry and the
decoding but was not steered. The wrong sign at layer is beyond first order in
the displacement and its mechanism is not settled here. The corrected effect is modest. The
whitened direction's response at one class gap is two null standard deviations, where
the wrong-signed response of the mass-mean direction was five, so the correction
restores a sign rather than supplying a large causal lever. On one of four datasets the
shuffled-label amplitude departs from the parameter-free prediction by – for
reasons that were localized but not explained.
This post has treated steering as a diagnostic. The steering results themselves are analyzed in a second, forthcoming post, Steering Vectors and the Limits of Linear Response. There the object of interest is the vector itself. Starting from the identity , I ask how much of lies along the high-variance directions of the within-class covariance, what bound it puts on any linear steering direction, how much of that bound the best direction actually reaches, and how much of published steering vectors capture.
A fuller, annotated version of this bibliography — organized as a reader's map of how these results tension against each other — is at the geometry of truth probes. The grouped list below gives locators only.
The geometry — separability, capacity, and readout
Cover, Geometrical and Statistical Properties of Systems of Linear Inequalities with Applications in Pattern Recognition — IEEE Trans. Electronic Computers EC-14(3):326–334 (1965) PDF.
Diedrichsen, Berlot, Mur, Schütt, Shahbazi & Kriegeskorte, Comparing Representational Geometries Using Whitened Unbiased-Distance-Matrix Similarity — arXiv:2007.02789, Neurons, Behavior, Data Analysis, and Theory (2021).
Truth / honesty directions — reproduction targets
Marks & Tegmark, The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets — arXiv:2310.06824, COLM 2024.
Bürger, Hamprecht & Nadler, Truth is Universal: Robust Detection of Lies in LLMs — arXiv:2407.12831, NeurIPS 2024.
Burns, Ye, Klein & Steinhardt, Discovering Latent Knowledge in Language Models Without Supervision (CCS) — arXiv:2212.03827, ICLR 2023.
The critiques — identifiability / which direction did you actually find (the SNR angle)
Farquhar, Varma, Kenton, Gasteiger, Mikulik & Shah (DeepMind), Challenges with Unsupervised LLM Knowledge Discovery — arXiv:2312.10029 (2023).
Roger, What Discovering Latent Knowledge Did and Did Not Find — AlignmentForum, 2023.
Mallen & Belrose, Eliciting Latent Knowledge from Quirky Language Models — arXiv:2312.01037 (2023).
Bao et al., Probing the Geometry of Truth: Consistency and Generalization — ACL Findings 2025.
Representation geometry — the rogue-dimension lineage
Sun, Chen, Kolter & Liu, Massive Activations in Large Language Models — arXiv:2402.17762, COLM 2024.
Timkey & van Schijndel, All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality — arXiv:2109.04404, EMNLP 2021 (pp. 4527–4546).
Models
Biderman et al., Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling — arXiv:2304.01373, ICML 2023.
Steering directions
Tan, Chanin, Lynch, Paige, Kanoulas, Garriga-Alonso & Kirk, Analysing the Generalisation and Reliability of Steering Vectors — arXiv:2407.12404, NeurIPS 2024.
Braun, Eickhoff, Krueger, Bahrainian & Krasheninnikov, Understanding (Un)Reliability of Steering Vectors in Language Models — arXiv:2505.22637, ICLR 2025 Workshop on Foundation Models in the Wild.
Torop, Masoomi & Dy, Inverted Detection and Control in Steering Vectors — arXiv:2608.02957 (2026).
Liu, Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes — arXiv:2605.05715 (2026).
Prose drafted with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.
This is a condensed version of a longer post on my website, Truth Directions: Signal-to-Noise and Geometry of Recoverability, which carries the definitions, the derivations, the transfer and distractor analyses, the OLMo replication, and six appendices.
Introduction
The linear representation hypothesis states that a model encodes high-level concepts as directions in activation space. This reduces reading a concept to a dot product with a single vector, and steering it to adding a multiple of that vector to the residual stream. For truth in particular, Marks & Tegmark (2023) showed that a difference-in-means ("mass-mean") direction fit on true/false statements separates held-out statements, transfers across datasets, is causally implicated under intervention, and — crucially — sharpens with model scale.
Previously, Burns et al. (2022) had proposed finding a truth direction without labels, by demanding logical consistency, a method they call Contrast-Consistent Search (CCS). Within months, Roger (2023) showed empirically how little the objective pins down. Untrained, randomly initialized probes already reach about on the "easy" datasets, once the CCS convention of
flipping a below-chance probe is applied, and CCS does not find the optimal linear
probe: more than twenty mutually orthogonal probes reach accuracies similar to the one
it returns. The flip matters here, because a sign-resolved random baseline is precisely
what a null has to be. A direction and its negation separate the classes equally well,
so any honest chance level already includes the better of the two.
Farquhar et al. (2023) then gave the reason:
arbitrary binary features are optimal under that consistency loss, so nothing in it
selects for knowledge, and unsupervised probes in practice recover whatever feature is
most prominent in the representation. Whether the supervised estimator is exposed to
the same failure is a question this post takes up.
The common thread across these observations is that this is a signal-versus-noise problem. A linear direction can look like it reads truth while actually tracking something more salient that happens to correlate with the label on the particular dataset. Consequently, on a benchmark where chance-level structure is this strong, "the probe separates the classes" is a much weaker statement than it first appears.
Setup
Throughout, a hat marks a sample estimate: is the direction fitted on a finite sample, the population direction it estimates. The mass-mean direction is the unit vector along the class-mean difference . The Fisher (whitened) direction is , with the within-class covariance under Ledoit–Wolf shrinkage. It is the direction that maximizes the separation below. Both are fitted on a training half and scored held-out. The separation along a unit direction is , the class gap over the within-class spread along . For Gaussian classes . Both and the are reported throughout the post, and and denote of the mass-mean and Fisher directions. A rogue dimension is a leading within-class eigenvector that carries nearly all the noise variance, defined precisely below.
Every result is measured against a random-direction null, one for decoding and one for steering. The decoding null draws random unit directions per layer and scores them on the same held-out points. A probe counts as recovered when it clears the null's 95th percentile. The steering null draws random directions the same manner, adds them to the residual stream at one layer with the same norm as the fitted direction, and measures how the model's behavior moves. The behavioral score is the log-probability gap between a true and a false completion of the same prompt, and its response to a displacement along direction , with the class-gap unit so that counts class gaps, is split into an odd part , the signed effect, and an even part, the generic disruption. An effect is called significant when it clears the 95th percentile of its null. The susceptibility is the through-origin slope of against . [/
I run experiments on the Pythia ladder, 70m to 2.8b, with OLMo-2-1B as an architecture control. The datasets are the twelve true/false sets of the Marks & Tegmark benchmark at statements each. The steering and mechanism analysis primarily use
citiesandcounterfact_true_false. Full definitions, the geometry of whitening, and the protocol are in the full post.Where the direction is recoverable
With the null, , and the class-mean-gap unit in place, the reproduction question
becomes concrete: at what scale, at what depth, and across which datasets is a truth direction
actually recoverable?
Across the Pythia ladder it emerges with scale. On at pythia-70m, inside the random-direction null ( ) and so
not recoverable, to at 410m, at 1.4b, and at 2.8b, with
rising from to . At 70m the whitened direction does not clear the null at any
layer either, the best value being on on
citiesthe best-layer mass-mean AUROC climbs fromcitiesandcounterfact. What fails there is the whole linear class, not one estimator within it. Whether a 70m model represents truth in some form this class cannot express is not tested here.Emergence of the truth direction across the Pythia scale ladder, . Layers selected by the largest margin above each layer's own null.
cities: best-layer held-out AUROC of the mass-mean direction (red) and the whitened direction (blue) against the random-direction null (gray band, up to its 95th percentile). At 70m both directions sit inside the null. From 410m on, both clear it, and the mass-mean margin widens withWithin a single model, the truth direction is a property of depth. Sweeping the layers of pythia-2.8b on at layer . The null itself widens with depth, its 95th
percentile climbing from early to deep, so recoverability is
again the margin above the null, not the raw number: a deep-layer AUROC in the low
s can still sit inside it. Layer selection maximizes the margin, not the AUROC,
and the two peak one layer apart: the AUROC at layer ( against a null of
), the margin at layer ( against ). The scale ladder
above reports layer , since the margin is what carries the claim.
cities, the mass-mean AUROC sits inside the null through the early layers, lifts clear of it around the middle of the network, and plateaus high across the late layers, peaking atTwo further results are stated here and fully fleshed out on the full unabridged post . First, I find that the mass-mean direction transfers within a family of datasets and inverts across a negation ( on
citiesscores AUROCneg_cities), so that a single vector approximates a structure with at least a separate polarity axis. Second, a direction fitted on thelikelydistractor, whose classes are the most and hundredth-most-likely final tokens, does not transfer to a truth set and has the opposite depth profile, so the plausibility axis is not read by the truth probe.How small is too small?
Cover's theorem marks as the scale below which separability stops being
informative. The shuffled-label control measures what that costs in practice: fit the
mass-mean direction to a random permutation of the labels and record how much
separation comes back.
This control is the one place in the post where a direction is scored on the points it was fitted to. Everywhere else a direction is fitted on a training split and scored held-out. Here that would defeat the purpose, since on held-out data a shuffled-label direction scores by construction, and the quantity of interest is
exactly how much separation the estimator manufactures from the noise it was fitted
to. Scoring in sample is what exposes it. Pooling the four models, the excess AUROC a
mass-mean direction extracts from pure noise scales as
The exponent follows from the estimator alone. is a difference of two
sample means, each concentrating at the rate, so the noise it carries has
magnitude whatever the activations look like, and the excess must
fall as at fixed . Measured across four models and from
to , it comes back . The phenomenon is dimensional slack, the
room that leaves for a direction built from a sample's own fluctuations to
separate that sample, since the fluctuations that define the direction are the ones
being scored. It is not capacity.
Inverting the law gives the condition for noise alone to contribute less than of excess AUROC,
which at the widths in this post and the standard 7B width gives the following. The table uses the unrounded fit, and , and the rounded values
reproduce each entry to within about .
The curated true/false datasets in this literature hold
statements. For a model, the width of the 7B models these probes are
usually run on, noise alone buys in-sample AUROC until
and until . Every dataset in the benchmark is between five
and thirty times too small for an in-sample AUROC in the seventies to mean what it
appears to mean. Held-out scoring removes the inflation but not its source. The fitted
direction is the same object either way, and at these sizes it is mostly noise: a
synthetic control in the full post puts its overlap with the true direction at for
the post's own . Held-out evaluation reports that attenuated
direction's honest score, which is why held-out errs low where in-sample
errs high.
The prefactor is where the dataset enters. What it quantifies is the number of directions the noise effectively occupies. A spectrum concentrated on a few axes leaves a random labeling less room to find a separator than a flat one does, so the ambient width is the right count only when the spectrum is flat. In general the count is
the participation ratio of the within-class covariance, and the amplitude should
collapse across datasets once is measured in units of it.
The full post derives the prefactor. Under shuffled labels the mass-mean direction is a signed sum of samples, , and its in-sample separation is with the participation ratio of the within-class covariance, the count of directions the noise occupies. With the collapse constant is predicted outright as , with no free parameters.
Measured, is flat to about within a dataset over a thirtyfold range in
, and runs from to across in and in . The
dependence on the spectrum runs the way the count predicts and against the intuitive
direction: with
of the within-class variance on one axis, and it carries the smallest amplitude
of the three.
counterfact,citiesandlarger_than, against a factorlarger_thanis the most concentrated set,Left: the shuffled-label law. In-sample excess AUROC of the mass-mean direction under shuffled labels on , with the pooled fit and Cover's capacity marked. The three pythia-70m points with negative excess are not drawn. The 70m points sit below the pooled line throughout, which is the across-model form of the question the right panel settles across datasets: whether the ambient width is the right denominator. Right: the effective-dimension collapse on pythia-2.8b, each dataset at its own best layer, against . The line is with no free parameter. and a fourfold range in . – . Error bars are the standard error over sixteen permutations.
counterfact, four Pythia widths, againstcounterfact,citiesandlarger_thanfall on it across a thirtyfold range insp_en_transsits above it bysp_en_transdoes not collapse onto the predicted prefactor. It sits aboutA steering effect with the wrong sign
The decoding analysis says when a truth direction is readable. Steering asks the separate question of whether it is causal, that is, whether displacing the residual stream along moves the model's behavior toward the true completion. The two
need not agree. On is against a null of
, and it stays inside the null at every layer except for the last. As an intervention
it produces a significant effect with the wrong sign.
counterfact_true_falseat the deep layers of pythia-2.8b the mass-mean direction is neither readable nor correctly causal. For the readout, the held-out AUROC at layerAs defined above, steering is reported as the antisymmetric response in
class-gap units. A genuine truth direction produces , and the random-direction
steering null fixes the bar must clear. Across ten seeds at layer , at a
displacement of one class gap, the mass-mean direction returns
where is the fraction of the -draw random-direction null of at
least as extreme in the same direction, evaluated seed by seed. The seed median sits
five null standard deviations below zero, and the two positive seeds sit inside the
null. The linear steering susceptibility , the through-origin slope of
against over , is at this layer. It is a diluted summary of
the same effect. The response is not linear in , and by , four class gaps
out, has already fallen back toward zero, which is the regime The scale of an
intervention warned about. Layer gives the same sign less cleanly:
seeds negative with a median against a null standard deviation of
, and one seed strongly positive. The effect is not null. It is
significantly wrong-signed. Displacing an activation toward the
true-class centroid, along the very vector the estimator returns for truth, makes the model
measurably less likely to produce the true completion.
Because reads at chance here, that significance carries the whole claim.
A direction sitting inside its decoding null is expected to do nothing under
intervention, and a small negative point estimate on its own would not be
distinguishable from noise. What makes this a result rather than a null is that the
effect clears a -draw random-direction null in the wrong direction, reproducibly
across seeds and at two layers.
Read naively, this is a failure of the linear picture. Displacing along the direction fit to read truth does not raise the probability of the true completion. It lowers it. That reading would put sits inside its decoding null, so there is no separation for the
steering result to be epiphenomenal to. However, I use the rest of this post to argue
that the naive reading is wrong: the wrong sign is a property of the estimator, not of
the activation geometry. What that means precisely is that the same activations, at the
same layer, contain a direction along which displacement moves the model toward the true
completion. But the mass-mean rule does not return that direction. It returns ,
the axis of largest within-class variance, with which shares a cosine of
at this depth. The weight of that claim rests on the correction rather than on
a mechanism for the wrong sign: replacing the mass-mean rule by the Fisher rule, inside
the same linear class, recovers the correct sign, while what makes the plain direction negative at layer turns out to lie
beyond first order and is not settled here. Change the rule, either by downweighting that axis or by
deleting it outright, and the correct sign comes back out of the same data. The claim is
about which part of the geometry the estimator points at, not about how much truth the
layer encodes. How well the corrected direction reads is a separate question, quantified
in the next section by its held-out AUROC against the decoding null.
counterfactin the column of cases where the truth direction is an artifact of the readout, epiphenomenal to the behavior. It does not even earn that description. At this depthWhitening reverses the sign of causal steering. Left: antisymmetric steering response on . The plain mass-mean
direction (red) drives significantly negative, so steering toward the true centroid
suppresses the true completion, while the whitened direction (blue), differing only
by the covariance correction, drives it positive. Gray band: 5th–95th percentiles
of the random-direction null, so a point outside it in its own direction is a
one-sided clearance at 5%. Right: steering susceptibility , the through-origin slope of against ,
across depth, median over ten seeds with inter-quartile bars, the median rather than
the mean because the plain direction's seeds are bimodal, eight clustered near and
two positive, so a mean lands on a value no seed exhibits. Plain is wrong-signed at
( – seeds positive), whitened is correct-signed throughout ( ), with
the crossover.
counterfact_true_false/ pythia-2.8b atThe rogue dimension
Definition. I call a rogue dimension when it dominates the
within-class covariance, that is, or equivalently and . This
condition is a statement about the noise geometry alone and is basis-free and independent
of any probe. The estimator only becomes involved through the separate question of
whether has aligned with it.
This is deliberately not the same object as the massive activations of Sun et al., which are individual coordinates whose magnitude far exceeds the median, nor the rogue dimensions of Timkey & van Schijndel, which are coordinates that dominate cosine similarity. Those are properties of the mean and of a basis. This is a property of the covariance and of no particular basis. The distinction is important here. At the two coincide: carries of its mass on a
single coordinate, and that coordinate's mean activation is the median
across coordinates, a massive activation by their criterion. By layer the
eigenvector has delocalized, on its largest coordinate and across five,
so the rogue dimension is no longer any one neuron. still has
the massive activation, the same coordinate at the median, and has no rogue
dimension at all. Its leading eigenvector holds half a percent of its mass on any
coordinate. A massive activation is nearly constant across statements,
so it inflates the mean without inflating the within-class covariance. It produces a
rogue dimension only when the remaining variance is small enough for it to dominate.
counterfactlayercitiesat layerThe diagnosis is in the spectrum of the within-class covariance. I diagonalize with sample eigenvalues
and eigenvectors ,
and measure where the mass-mean direction sits relative to its leading eigenvector.
On :
counterfactat pythia-2.8b, layerA single eigendirection carries of the within-class
variance. It is times larger than the next, and the mass-mean direction is
almost perfectly aligned with it. Therefore, the estimator has not returned a truth direction, it has returned , the dominant axis of the within-class noise. On this
dataset there is essentially no class-gap signal for to lock onto, so the
finite-sample is dominated by its projection onto the highest-variance
axis. This axis is the salient direction of the superposition test above. When the class
gap is negligible, the leading direction of the total activation covariance and the
within-class coincide, so the two diagnostics see the same axis. They
separate only once the gap grows. The participation ratio
puts the collapse on a
dimension-free scale: when one eigenvalue dominates the spectrum
and when the spectrum is flat, so it counts the directions the
noise effectively occupies.
Read across depth, those observables say something sharper than the layer-
snapshot. What decides recoverability is not whether a dominant axis exists. At layer
, , and . That is it no longer holds for , the leading axis falls to of the variance, the alignment
drops to , and reaches . – at every layer, its alignment stays near , and
throughout. These observables are computed on the full set
rather than the held-out half, since they describe the geometry rather than a probe's
performance, and the held-out at is
against the quoted here. The rogue dimension is the default
condition, not the pathology. Nor is it a property of and , within of and alignment above ,
and escape it by layer where their class gap has grown, as the twelve-dataset appendix of the full post records.
Recoverability is whether the class gap ever grows large enough to pull
off that axis and clear the null.
citiesis collapsed too, withcounterfact's condition, not a milder version of it. The two sets differ in what becomes of that condition at later layers. By layercities. The participation ratio climbscounterfactnever moves. Its participation ratio stays atcounterfactlayercounterfactalone:companies_true_falseandcommon_claim_true_falsecarry the same signature at layersRead against the decoding null, the two datasets separate cleanly, and the correction makes the truth direction recoverable on and stays above every
subsequent layer, climbing above from layer and peaking at at
layer . The whitened direction clears the null earlier, at layer .
On against a null of
, while the whitened direction leaves the band at layer and climbs to
against at layer . The layer- value is the maximum of a
test taken over layers, so it is a selected extreme rather than a
recovery. The rank-one
direction , the mass-mean direction with its component along
projected out, measured on a six-layer grid in the rogue-dimension sweep, tracks the
whitened direction: on and
reaches at layer , and on by layer
. Deleting the axis and downweighting it do the same work.
counterfact. Oncitiesthe mass-mean direction leaves the null band at layercounterfactthe mass-mean direction stays inside the null band at every layer but the last, where it reachescounterfactit sits at chance through layercitiesit is already atThere is a sharper way to explain why the plain probe cannot leave the band on lies along and carries nearly all
of , a random direction reads the class gap and the noise through the
same coordinate, and
, so
counterfact. Whenand cancels. For every random direction the projected gap and the projected
noise are carried by the same component, so their ratio is the one itself
gives, and the null collapses onto the mass-mean value. The sweep shows it
to the third decimal. On is ,
, , at layers , , , , and the median of
the random-direction null at the same layers is , , , .
The plain probe is not merely inside its null. Its is indistinguishable from
that of a random direction, and the spread above that median, from to , is what the two percent of variance off the axis contributes. That is what locked to the axis means
operationally, and it is why more data would not lift clear of the band
while the alignment holds: the null and the estimate move together.
counterfactthe held-outHeld-out AUROC against depth for both estimators on pythia-2.8b, with the per-layer random-direction null shaded from to its 95th percentile. A curve inside the band is not recoverable, and height above the band is the margin that layer selection maximizes. Left, , while the whitened direction leaves it at layer and reaches at layer . Right, and the whitened one at layer , saturating near and . Stars are the rank-one corrected direction , which projects out of the estimator and is measured on a six-layer grid rather than at every layer. Dashed lines mark the first layer from which each direction stays clear of its null. Note the different vertical scale of the achievement: the same correction that lifts is barely needed where the class gap is strong.
counterfact: the mass-mean direction sits inside the band until layercities: the mass-mean direction leaves the band for good at layercounterfactfrom chance toWithin-class geometry on , , colored by truth label — on , with its separation and its alignment , near 1 on . The eleven points standing clear of the bulk on in the bottom-left panel, away from the dense blob at , and the small bar near in the bottom-middle one — are the statements that do not receive the massive activation. They are what the leading eigendirection is measuring, and what sets the axis range of both panels.
cities(top) andcounterfact(bottom), pythia-2.8b layer 28. Left: activations in the plane of the top two within-class principal axescounterfacta single axis carries nearly all the variance. Middle: the mass-mean projectioncounterfactwhere the estimator has collapsed onto the leading variance axis. Right: the within-class eigenvalue spectrum,counterfact— the scattered group nearThe obvious objection is that this is a fact about Pythia. However, the same observables on OLMo-2-1B, a different architecture and training corpus at a third of the parameters, reproduce the pattern. On
counterfactthe mass-mean direction stays aligned with the leading axis at every depth except the final layer, and the plain probe never clears its null while whitening does, and oncitiesthe alignment falls, the participation ratio climbs, and the plain probe clears the null as it does in Pythia. The numbers are in the OLMo appendix of the full post.Where the rogue dimension comes from
The definition in the previous section separates the rogue dimension, a property of the within-class covariance, from the massive activation, a property of one coordinate. On , where
carries of its mass on the massive coordinate, and differ by layer
, where the eigenvector has delocalized across five coordinates. On the massive activation is present and the rogue dimension is absent. However, the
relation is closer than coincidence. On
counterfactthey coincide at layercitiesat layercounterfactthe leading axis of the covariance is generated by the statements on which the massive activation drops.A coordinate that is exactly constant across statements contributes nothing to , however large it is. A coordinate that takes a large value on a
fraction of statements and a much smaller value on the remaining
contributes the variance of a two-point distribution,
which for small and large can be enormous. At ,
four coordinates qualify as massive activations under the magnitude criterion of
Sun et al. On each, of the
statements sit at a near-constant value and eleven sit far below it. Coordinate reads on the majority and
on the eleven. With the formula predicts against a measured
coordinate variance of . Summed over the four coordinates it gives
against , and carries of its mass in
their span. At layer the same holds with eight such coordinates: against
, with of in their span. In this case carries no truth content. It is, to within a few
percent, the indicator of which statements failed to receive the massive activation.
The mass-mean estimator's alignment with it is an alignment with an eleven-out-of-1198
membership function.
counterfactlayerThe eleven statements are identifiable coordinate-first, by a rule that never touches the covariance, the class means, or the labels, and it returns the same eleven at every layer. Refitting with them excluded collapses the geometry at layer , taking
from to and
from to , and it leaves
held-out unchanged at . At that layer whitening leaves
nothing further for deletion to recover. Nothing rescues the plain estimator, whose
stays inside its null under every removal regime tested. How the
eleven are identified, what removing them does to at each layer, and how that
comparison depends on the Ledoit–Wolf intensity are given in the appendix
Identification and removal of the eleven outlier statements.
The massive coordinates themselves are not dataset-specific. The same coordinates qualify on , and leave no rogue dimension behind. That is the earlier observation that
citiesandcounterfactalike. What differs is whether any statement drops these coordinates. Oncitiesnone do, so the coordinates stay constant, contribute nothing tocitiescarries the same massive activation without the same pathology, now with a mechanism rather than a coincidence. The incidence is specific to the dataset but the mechanism is general. The per-layer incidence on both models, including pythia-1.4b, is in the eleven-statements appendix of the full post.While and are aligned geometrically, it is unclear whether this is causal. The steering sweep includes a condition that displaces along itself, the
leading eigenvector of the within-class covariance. Because is fit after both class
means are removed, it is defined purely by within-class scatter and carries no
information about which statements are true. At layer the antisymmetric response
to is against 's . At layer the two agree to
within ( against ). At every depth measured they track each
other to within a few thousandths. They decode alike too, at held-out
for against for in the
six-layer rogue-dimension sweep, both inside the null. The alignment is a
geometric statement. The steering agreement is its behavioral form: a direction fit
without any reference to the labels moves the model as much as does and in
the same direction. That establishes that truth content is not what carries the effect.
It does not add independent evidence about the size of the effect, since under
, derived below, two directions with cosine must
produce nearly equal susceptibilities.
This reframes the anomaly. Steering along is not steering
along truth. It is displacing the activation along the dominant axis of the within-class
noise, which perturbs the forward pass in a way that happens to suppress the true completion.
A nuisance direction carries no information about the label, so nothing requires its
behavioral effect to come out positive. Nothing requires it to be reproducible either,
but here it is: the same negative sign appears across ten seeds, at two says only that truth is not what
produces it.
counterfactlayers, and at mid-depth oncities. Whatever produces it is systematic rather than arbitrary, and the label-blindness ofIf that account is right, removing the contribution of should recover a
correctly signed causal effect, by correcting the estimator's alignment with the rogue
dimension rather than by enriching the function class.
What steering measures
Before testing that prediction, it is worth asking what the steering number reports. The susceptibility is a linear functional of the direction pushed. Writing for the mean gradient of the behavioral
score over the evaluation set, the odd part of the response gives
to leading order, the even terms having been removed by the antisymmetrization. So steering in the small- window reports the overlap of with the single vector
along which the model's true-versus-false margin moves. It does not test whether
is a truth direction.
This also makes the assumption behind the correction explicit. Projecting out recovers the sign only if has little weight along , i.e.
the behaviorally causal direction and the salient axis are close to orthogonal in
activation space. When they are not, mixes a truth term with a nuisance
term inseparably, so deleting the axis would delete real causal signal along with the
artifact.
Measured. One backward pass per example gives , and follows by
averaging over the same contrastive pairs the steering measurements use. Because the
gradient is computed per pair, every overlap below carries a bootstrap interval over
those pairs rather than resting on the point estimate. I quote intervals from
bootstrap resamples under a fixed seed, and
the scale to keep in mind is that a random direction gives at .
At layer , with confidence interval
. The overlap is consistent with zero. That has two consequences. Projecting out
removes essentially none of the behavioral channel, which is what licenses the
correction. A channel consistent with zero cannot produce the wrong sign at first
order, so at this layer is not a linear effect. It is beyond first order in the displacement.
At layer the overlap is resolved: , confidence
interval , and is the same through the
alignment. The linear prediction then lands
within about of the measured . The through-origin susceptibility
hides this, because the response reverses sign with displacement, from at
( seeds negative, null standard deviations) through zero near
to at ( seeds positive), so the slope fit over
comes out null. The two layers therefore differ in kind. At layer the
wrong sign is a first-order effect confined to the linear window. At layer the
same prediction gives against a measured at , and the
wrong sign persists to before its magnitude falls at .
The overlap with also separates the two estimators. On on,
at , at and
at , while the plain direction is consistent with zero at
and . On , , while the whitened direction does not. In
both datasets the direction that steers correctly is the direction whose overlap with
is positive, and the overlap is measured without displacing an activation. The
sign and not the magnitude is what carries this: at both
directions have resolved overlaps, the whitened one positive and the plain one negative.
counterfactthe whitened direction has a positive overlap at every layer fromcitiesit is the plain direction that has the positive overlap at layercounterfactlayerThe correction is a rotation in a plane
Two corrections to the estimator, whitening and projecting out , make the
same prediction, and neither adds capacity. Whitening replaces
with the Fisher direction
, which downweights the
mean-shift component along high-variance axes, chief among them. The blunter
control simply projects out, ,
removing only the rogue dimension.
Both lie in the plane spanned by and , the unit vector along the part of
orthogonal to . In this plane the separation is a Rayleigh quotient in the angle from . With the puts the mass-mean direction at from the noise axis and puts the Fisher direction at , so the two corrections are the same vector to within a degree. The plane formula predicts a gain in and the measured held-out gain is , . The full post gives the derivation and the plane figure.
counterfactnumbers,Left: the plane at the , so a
axis ratio). The mass-mean direction (black) lies off the rogue
axis, inside the long axis of the noise. swings the Fisher direction
(gold) to which is effectively orthogonal to it. Right: the Rayleigh quotient
normalized by its maximum, in the collapsed regime ( ,
) and the resolved regime ( , ). Circles mark the
mass-mean angle and squares the whitened angle. Where the spectrum has one dominant eigenmode the
curve is a cliff and the mass-mean estimator sits at its foot. Where the class gap has
grown the curve is broad and both directions already sit near the
top.
counterfactnumbers, with the within-class noise ellipse (Two conditions make the reduction valid, and together they are the criterion for this failure mode. The spectrum must have one dominant eigenmode,
with , or there is no plane to reduce to. Additionally,
must be small, or the estimator is not at the foot of the cliff
and there is nothing to correct. Both are read off the activations and the fitted direction, and together they diagnose
that the estimator has collapsed onto . They do not by themselves say which
way either direction will steer. That is set by the overlaps with the score gradient,
for the mass-mean direction and for the
correction, which are also measured without displacing an activation. At through the spectrum meets the criterion more strongly than at layer
, yet the corrections do not restore the sign there, because is
negative or unresolved. The diagnosis from the spectrum and the sign from the gradient
are both available before any steering is run. The diagnosis has been tested on one
dataset ( and but cannot be tested causally with this harness. Their statements
have no relation structure from which to build a pair of completions, so the
behavioral score defined in the setup cannot be formed. The alternative
would be to score the model's own judgment of whether a statement is true, but
unsteered that judgment separates true from false at AUROC on on violates both conditions ( ,
). There the formula predicts a gain of at
and , since both directions already sit on the broad top of
the quotient, and the measured is a small gain from the rest of the
spectrum, outside the plane. The formula stops applying where the regime ends, and the
regime control below draws the same boundary.
counterfactlayerscounterfact) that satisfies the criterion, and one (cities) that does not, and held in both. Two further main-tier sets,companies_true_falseandcommon_claim_true_false, satisfy the criterion at layerscommon_claim, againstcities, so there is almost no behavior there for a displacement to move. So they meet the criterion, but the prediction it makes for them cannot be checked.citiesat layerBoth corrections flip the sign. Steering along the whitened direction at layer
gives
with seeds individually clearing the null at . Layer gives the
same picture: positive, clearing, . Here
coincides with at , because the whitened response is linear in
over the whole sweep, , , , at ,
where the plain direction's response fell off at . The corrected effect is
smaller than the wrong-signed one it replaces, two null standard deviations against
five, but it is monotone where that one was not. The rank-one correction
, measured independently in the rogue-dimension sweep, agrees. At
layer its three seeds give , and , each at
against the same -draw null, so on its own it carries from
significantly negative to significantly positive, and it raises held-out decoding AUROC
at layer from , the mass-mean value inside the null, to . Layer
is the crossover. Whitened steering there is correctly signed in seeds
but does not clear the null, , which corresponds to a transition region rather
than a clean effect.
The regime control
Marks & Tegmark steer along the feature direction and report that difference-in-means directions are the most causally implicated of the probes they compare. At the deep has collapsed
onto and is not tracking a feature at all, so the meant to
sharpen a readout is instead doing the work of recovering the direction.
counterfactlayers that ordering reverses. The mass-mean direction, which their framework nominates as the feature, steers the model away from the true completion, and the whitened direction, which it demotes to a decision boundary, steers the model toward it. The rogue-dimension diagnosis explains why. ThereThe inversion is a property of a regime, not a refutation of Marks & Tegmark, which is demonstrated by , the same model and depth at which
seeds
( ) while the whitened direction is weak and inconsistent
( , ). On pythia-1.4b the contrast is starker still, with
plain against whitened at layer . Read across the whole
depth sweep, the two datasets mirror each other. On and clearly and weakly ( ,
seeds positive, ), and is correctly signed at all three. On
and , and is
wrong-signed at both. At every other layer of either dataset it sits inside the null.
At layer the two datasets share the model, the depth and the estimator, and
their significant effects point in opposite directions. When no single eigenmode
dominates the within-class covariance, the mean difference is the causal feature and
the inverse covariance only adds estimation noise, which is what Marks & Tegmark report, on datasets of exactly this kind. The inversion is confined to the regime where has collapsed, and the rogue dimension decides which regime a dataset is in.
citiesas a the control. There the two directions behave as mass-mean probing intends. At pythia-2.8b layercounterfactinverts, the plain direction steers correctly incities, plain steering clears its null at three layers,counterfact, plain steering clears its null at two layers,Two points in the and the plain direction becomes
mildly negative with and . Both estimates are consistent across
seeds ( and negative), so the sign is not noise. However, both sit an order of magnitude below the the same direction produces
at layer , and neither clears the steering null. The regime
claim is that the significant effects on
citiespanel deserve naming, since they are visible and read at first glance like counterexamples. At layerscitiesare correctly signed, not that every layer's point estimate is positive.The causal direction is set by the presence or absence of a rogue dimension. Steering susceptibility on pythia-2.8b, plain (red) against whitened (blue), median over seeds with inter-quartile bars. Left, and do not clear the steering null and are an order of magnitude below its layer- effect. Right,
cities: the plain difference-in-means direction carries the causal effect and whitening degrades it which is the intended behavior of mass-mean probing. The mild negative excursions of the plain direction at layerscounterfact_true_false: at depth the assignment inverts, the plain direction steers significantly wrong-signed while the whitened direction steers correctly. This is the same model, estimator, and protocol and only the within-class geometry differs.Discussion
The results above establish that, when the within-class spectrum meets a specific criterion, the mass-mean estimator returns a nuisance direction rather than a truth direction, and that a correction inside the linear class restores the sign of the steering effect along the corrected direction. Two things should be noted before comparing to other reports. The first is the magnitude of the corrected effect, and the second is what its success does and does not imply about the function class.
The corrected effect is modest. At one class gap the plain direction's median response is , five null standard deviations, and the whitened direction's is ,
two. The correction flips the sign of the effect but returns less than half its size.
The decoding gain is likewise real but small in absolute terms. Whitening carries
held-out AUROC at layer from , inside the null, to against a
null of , and the whitened direction stays clear of its null from layer
to the end of the network, peaking at against at layer . Set
beside , this is a weak readout. The overall
claim is a corrected sign and a confirmed mechanism, not a recovered truth direction of
practical use.
cities, where the same direction readsA wrong-signed steering effect invites a natural reading that the linear function class is too weak: the direction one can fit is not expressive enough to move behavior, and a richer, perhaps nonlinear, intervention is required. The rogue-dimension account makes a different claim about this particular failure. The class is adequate, but the mass-mean estimator points at , a massive-activation direction, instead of at
truth. The two readings make different predictions, and the data separate them. A
rank-one correction within the same linear class, with no richer classifier and no
nonlinear probe, converts a wrong-signed effect that clears its null into a
correctly signed one that clears its null. For the and the function class was never the
bottleneck. The estimator's alignment with a massive-activation direction was.
counterfactfailure at layersRelation to other reports
Four recent reports also examine why a fitted direction fails to steer, and the rogue-dimension account can be compared against each. Braun et al. and Ying et al. locate the failure in the direction. For Braun et al. steering is unreliable when the per-example activation differences do not point the same way, so that their mean is not representative of any of them. For Ying et al. steering is wrong-signed when the direction is fitted across many kinds of truth and mixes truth with sycophancy. Torop et al. find a direction that discriminates well and steers backwards. Liu finds a direction that decodes but does not steer at all. This post shares a predictor with the first two and differs on the diagnosis, and against the last two it supplies the remaining case, a direction that neither decodes nor steers correctly.
The four comparisons are made in full in the full post. To briefly summarize, Braun et al.'s separability index is the of this post, so the two share a predictor but differ on what happens at its small separation. There, Braun et al. find that steering still moves the behavior in the intended direction on average, but with a large fraction of examples moving the wrong way. Here the average itself has the wrong sign and clears the random-direction null. By contrast, Ying et al.'s wrong sign comes from a direction that mixes domains, which cannot happen for a direction fitted on a single domain. Likewise Torop et al.'s inverted vectors decode well whereas the direction here fails to decode. Finally, Liu's decodable-but-inert direction is a null result, which cannot distinguish an entangled direction from one which is absent one.
Summary
A mass-mean truth direction can either be a truth direction or the estimator's projection onto the most salient axis of the activations. On a benchmark the two are told apart only by a signal-to-noise reading. Three results follow. In-sample separation is inflated by dimensional slack that scales as with a prefactor
set by the effective dimension of the noise, meaning that for the datasets in this literature the
null and not the raw score is the bar. When the class gap is weak the mass-mean
estimator collapses onto the leading within-class eigenvector. On
counterfactthat direction fails to decode and steers with a significant wrong sign, five null standard deviations deep at one class gap. The failure is in the estimator: the same activations used to obtain the mass-mean direction contain a direction that steers correctly, and it can be reached without leaving the linear class, either by the Fisher direction or by projecting out the leading eigenvector.The results have limits. The criterion that diagnoses this failure, a within-class spectrum with one dominant eigenmode and the mass-mean direction aligned to it, has one confirmed positive and one confirmed negative. It identifies the collapse but not the sign of a correction, which is set by the overlap of the corrected direction with the score gradient. The steering measurements are on Pythia alone, since OLMo replicates the geometry and the decoding but was not steered. The wrong sign at layer is beyond first order in
the displacement and its mechanism is not settled here. The corrected effect is modest. The
whitened direction's response at one class gap is two null standard deviations, where
the wrong-signed response of the mass-mean direction was five, so the correction
restores a sign rather than supplying a large causal lever. On one of four datasets the
shuffled-label amplitude departs from the parameter-free prediction by – for
reasons that were localized but not explained.
This post has treated steering as a diagnostic. The steering results themselves are analyzed in a second, forthcoming post, Steering Vectors and the Limits of Linear Response. There the object of interest is the vector itself. Starting from the identity , I ask how much of lies along the high-variance directions of the within-class covariance, what bound it puts on any linear steering direction, how much of that bound the best direction actually reaches, and how much of published steering vectors capture.
Appendices
The appendices are omitted here and live with the full post on my site, at https://jasteinberg.github.io/blog/2026/truth-directions-snr/:
References
A fuller, annotated version of this bibliography — organized as a reader's map of how these results tension against each other — is at the geometry of truth probes. The grouped list below gives locators only. The geometry — separability, capacity, and readout
Truth / honesty directions — reproduction targets
The critiques — identifiability / which direction did you actually find (the SNR angle)
Controls and methodology
Representation geometry — the rogue-dimension lineage
Models
Steering directions
Prose drafted with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.