The current, incredible performance of AI models is closely related to their compression capabilities (Language Modeling Is Compression (Delétang et al., 2023); Compression Represents Intelligence Linearly (Huang et al., 2024)). This compression is imposed on them by the architectural choices made by engineers. For example, GPT-2 had a vocabulary of 50,257 tokens, yet its “operational space” was only of size 768. In such a space, only 768 directions can be described fully independently, so the model had to find strategies to efficiently store all 50,257 input tokens in this reduced space. In general, we call the phenomenon in which a model stores more features than it has dimensions superposition. However, superposition comes at a cost. At a given step of computation, models utilizing it have neurons that fire for multiple different inputs (so-called polysemantic neurons), which makes them hard to interpret.
Prior work, Toy Models of Superposition (Elhage et al., 2022), studied how superposition forms in two-layer ReLU networks with tied weights. The task was to reconstruct the inputs given to the model (the input features) at its output, after passing them through a narrow bottleneck inside the model. The importance of each feature and the rate at which it occurs (its sparsity) could be varied across experiments, and the quality of the reconstruction was measured with the mean squared error (MSE). The authors suggest that for independent input features of equal importance and sparsity, the best geometries for storing the features may be uniform polytopes.
In this work we study an analogous problem, but with a different model class: a linear encoding layer (the encoder) followed by one or more multilayer perceptrons (MLPs) with a bilinear activation. We treat the stack of MLPs as a single object, the decoder, whose task is to reconstruct the network's input from the geometry produced by the encoding layer.
With this setup we show that the uniform geometries found in Toy Models of Superposition are not the best possible solution to the feature-reconstruction task: non-symmetric geometries yield lower error. We argue that their result is an artifact of the model class they studied, and that deeper, more expressive models do arrive at these asymmetric solutions.
To explain why only sufficiently powerful models find them, we derive two theoretical benchmark decoders: (i) an analytical solution in a simplified case, obtained from the known data-generating process, and (ii) its generalization, approximated from data samples alone. Comparing different models against these benchmarks lets us understand why shallow models default to symmetric solutions while deeper models find the non-symmetric ones.
We hope these results bring us closer to understanding the computations performed by LLMs, since our models and transformer architectures both use a linear encoder as the first step of their computation.
2. Setup
We will be working with the following toy setup:
The input to the models consists of 4 independent features of equal importance. Each feature is active with a small probability (in our experiments we use ), so the distribution of the features can be formally described as
The models consist of a linear encoder layer with a bottleneck compressing the input features to a 2D plane, followed by one or more MLPs with a bilinear activation function (Sharkey, 2023; Pearce et al., 2025):
The task is to minimize the mean squared error (MSE) of the model's reconstruction:
We divide the model into two parts: the linear layer up to the bottleneck (the encoder) and the stack of MLPs after it (the decoder). This division will also help us derive the theoretical reconstruction floor imposed by the architectural constraints (going from a 4D space down to 2D and back to 4D). The model under study and the division point are shown below. The 2D plane at this point is where all the geometries presented throughout this work live.
For completeness, we note that Toy Models of Superposition used a similar architecture on this very task. For 4 features compressed to a 2D space, their architecture would consists of a linear encoder given by a matrix , followed by a decoder that applies the transpose of the same matrix , a bias, and finally a ReLU activation. Because the same weight matrix is used in both layers, we will refer to this case as tied ReLU.
3. Theoretical solution
In this section, we will introduce a simple, symmetric geometry that a theoretical encoder can implement, and provide a closed-form description of a decoder whose goal is to reconstruct the network’s input from the point in the 2D plane that the encoder maps it to.
We will then present two strategies the encoder can employ to lower the reconstruction error, and finally demonstrate how these two techniques can be used in tandem to obtain the true best solution. This solution will serve as the theoretical reconstruction floor against which we will later compare the behavior of the actual models under study.
3.1. Symmetrical antipodal encoding
The simplest way for the encoder to arrange the embeddings of the 4 input features on the 2D plane is as two antipodal pairs. The embeddings of each pair lie on a common line, pointing in opposite directions. The two lines are orthogonal, which allows the pairs to be analyzed independently, and their orientation is otherwise arbitrary. Below we show an example where we have simply chosen the basis axes as the pair lines. Here denotes the embedding of feature , i.e. the 2D vector the encoder maps the -th input feature to. It can be read directly from the encoder matrix as its -th column. There are four such vectors, , one per feature, and the encoder's output (which is also the reading seen by the decoder) is their weighted sum
The embeddings of the input features placed as two antipodal pairs on orthogonal axes.
The two pair lines are orthogonal and placed on the basis axes, so the horizontal coordinate of the reading is simply , regardless of what features 2 and 4 are doing. When reconstructing and we can therefore restrict our attention only to , the position of the reading along the segment between and , with at , at and at the origin.
We will use this line segment to understand the interaction between the introduced encoder and a theoretical decoder through two lenses: (i) knowing the distribution of the features exactly, and (ii) having access only to an "oracle" that gives us samples of the data with no description of the underlying distribution. In both analyses, we assume that the loss function, the MSE defined above, is known.
3.1.1. Closed-form decoder from the known distribution
Knowing the distribution of the input features, we can distinguish four cases that result in a reading on this segment (as a reminder, we use the probability of a feature being active throughout this work):
1. neither feature active, with probability ,
2. feature 1 alone active, with probability ,
3. feature 3 alone active, with probability ,
4. both features 1 and 3 active, with probability .
This encoding method has one huge caveat. In the fourth case, with both features co-active, the reading can land anywhere on the segment (except for the endpoints and themselves). Because of that, even a perfect decoder cannot distinguish this case with full confidence from one in which only a single feature, or none, was active.
The four cases on the segment between and , each represented by a black dot at different position. Example collision of two cases: feature 1 alone and features 1 and 3 together can produce the same reading.
Because of this information loss, the decoder needs to average between different plausible scenarios to give the best possible response (measured with MSE): the best guess at a reading is the average of the guesses each scenario would make, weighted by how likely each scenario is to have produced that reading. The Appendix makes this precise.
Below we present the formula for such a theoretical decoder reconstructing the first feature, when given a reading at position on the segment. Full derivation of this formula is presented in the Appendix, but here we just emphasize one phenomenon: a reading of exactly can only come from the no-feature case, so the decoder outputs there, even though the limits from either side are . This reading is marked with a black dot in the plot below.
The best decoder for on the antipodal pair, as a function of the reading on the segment between and . Open circle: the one-sided limits at ; closed circle: the value at itself.
The decoder for the third feature is simply the mirror image of the curve above, , so we can picture the two decoders on the same plot:
The best decoders for (blue) and (orange) on the antipodal pair, as functions of the reading .
3.1.2. Decoder from data samples alone
The above derivation was possible only because we knew the underlying distribution of the features. Now, let us try to obtain similar results while having access only to an "oracle" that gives us raw samples of the data (which is closer to a real-world situation).
Having samples of the data, we can also visualize them along our line segment:
4096 samples of the pair drawn from the sampling process, plotted as the reading against the value the decoder has to predict.
The above plot can make it more explicit that without the case of co-active features the best possible solution would be to use a ReLU function.
Now, how can we accommodate for the fact that there are situations in which different features are co-active? To do that we need to look at the definition of the loss function we are using.
In our setup the loss is the MSE. Recall that the decoder sees only the position of the reading, and that many different inputs produce the same . For example, with feature 3 inactive, and together with , both give . The decoder must therefore output a single number for all samples with that , and its performance is measured by the squared distance from that number to each sample's true . The number that minimizes the total squared distance to a set of values is their arithmetic mean, so the best decoder outputs the mean of over the samples with reading .
So, the choice of the mean thus comes from the loss, not from the data. If we chose the mean absolute error instead, the best output would be the median of the same set of values, and with a 0–1 loss its most frequent value (the mode).
However, with a finite number samples we cannot take the mean over "all samples with reading " literally, because no two samples have the same reading. Instead we split the line segment into very small intervals (bins), gather the samples whose reading falls into each, and average their . The result is a step function approximating the closed-form curve, and it converges to it as we take more samples and narrower bins. Below we present the result of such an average computed on the data samples shown above:
Binned means of the 4096 samples (100 bins) against the closed-form decoder . The blue dot at is the mean over the samples with reading exactly .
Having introduced the two methods of obtaining the best possible decoder depending on the context we are in, we now move on to the strategies, which can be applied to modify the initial encoder and increase the performance of the model as a whole.
3.2. Asymmetrical antipodal encoding
The first modification is very simple: we make one of the embeddings shorter than the other.
The symmetric antipodal pair and its asymmetric version, in which is shortened and lengthened ().
This has a rather counterintuitive outcome. With this change the prediction error is "poured" into the reconstruction of a single feature; however, measured over both features, the total error turns out to be smaller!
Below we compare the two encodings. In the asymmetric one, is shortened and lengthened so that is times . The first two panels show the closed-form decoders for both features, and the third the MSE these decoders achieve, split into the contributions of the two features.
Symmetric vs asymmetric antipodal pair (): the closed-form decoders for and , and the MSE of the best decoder split into the two features' contributions.
Is it then always better to make the pair more asymmetric? To check this, we sweep the length ratio from to and compute, for each value, the MSE of the best decoder, both in closed form and from binned samples. In the binned case, the segment is divided into bins of equal width and every feature gets its own decoder, which predicts the feature's mean value over the samples that land in the same bin. In both cases, as in the previous figure, we report the sum of the two features' MSEs rather than their average.
The summed MSE of and under the best decoder as a function of the length ratio . Top: the closed-form decoder, in total and split into the two features' contributions. Bottom: the closed form against binned decoders of three resolutions, one decoder per feature on the same bins, their MSEs summed. Triangles mark the minima. The dashed line marks the ratio used above.
The closed-form curve first falls, because the error of the lengthened feature vanishes faster than the error of the shortened one grows, but it flattens out at a ratio of about – (a minimum of at , against for the symmetric pair) and then slowly rises again, toward : beyond this point the shortened feature is essentially unreadable whenever its partner is active, and its error dominates the total. Asymmetry therefore helps only up to a moderate ratio, and the optimum is shallow. A binned decoder has its minimum at the same ratio (– for the three resolutions), but is stricter beyond it: once the short embedding becomes comparable to the width of a bin, the readings of the short feature on its own are no longer resolved, and the MSE climbs steeply, the earlier the coarser the bins (with bins the symmetric pair is better again above a ratio of , with bins above ).
3.3. "Opening" the antipodal pairs
Another modification boils down to slightly tilting one of the embeddings of a pair with respect to its partner. With this strategy, the case of no active features still lands at the origin, and each feature active on its own still lands on its embedding vector, but the co-active case moves off the line into the interior of the parallelogram spanned by the two embeddings.
This allows the decoder to distinguish features that are active on their own from co-active ones, which results in better reconstruction performance. In fact, a decoder of infinite resolution would be able to distinguish the co-activation at an arbitrarily small opening angle, since the co-active readings leave the line as soon as . This does not make the problem trivial, though. Opening a pair does not change the dimensionality of the problem: all four embeddings still live in the same 2D plane, so at most two of them can be linearly independent. A single opened pair, taken alone, would indeed be perfectly decodable, but the other pair shares the plane, and the interior of the parallelogram is also reachable by other combinations of features. The benefit of opening is therefore limited by two things: the interference with the other pair, which we return to at the end of this section, and the resolution of the decoder, which we turn to now.
Trained models have finite resolution. Below we investigate its effects by studying what a binned decoder is actually capable of. We keep the length ratio of both pairs equal to , tilt off the antipode of by , keep the other pair closed, and score every with binned decoders of three resolutions. The bins are now square cells covering the 2D plane of readings, and again every feature gets its own decoder, which predicts the feature's mean value over the samples that land in the same cell. Since all four features are now involved, we report the MSE as defined in the Setup, i.e. averaged over the four features.
The MSE of the binned decoder as a function of the opening angle of the pair , the pair closed, averaged over the four features, for three resolutions, one decoder per feature on the same cells (triangles mark the minima).
The MSE has its minimum at a finite opening, which moves toward zero as the bins get finer. But why does the MSE not stay constant once surpasses this threshold for a given binning? This is because opening one pair makes it interfere with the readings of the other pair. We will study this in more depth in the next section.
3.4. Search for the best strategy
We close this section by investigating what is in fact the best strategy for a linear encoder followed by a decoder that is restricted to no particular function class and is limited only by the resolution of the bins used to compute it.
We do this by a free search over the encoder geometry. For an unconstrained decoder the angle between the two pair lines does not matter, so we may pin two of the embeddings to the axes and let the remaining two be arbitrary vectors in the plane, described by their angle and length. Every geometry is scored by the MSE of a decoder obtained numerically, as in "Decoder from data samples alone" but with two-dimensional bins. The four parameters are optimized by a random search followed by local descent.
Below we present the geometry obtained by the above search along with regions indicating the co-occurrence of various feature pairs.
The geometry found by the search, with a zoom on the origin as an inset: the pair opened by at length ratio , the pair by at . Each shaded parallelogram is the region where the readings of two co-active features land; the two thin orange slivers along the axes belong to the pairs.
The search finds that the two free embeddings settle almost antipodal to the pinned ones, much shorter than them, and slightly tilted, each toward the long embedding of the other pair. The pair is opened by with a length ratio , and the pair by with , reaching an MSE floor of per feature, against for the closed symmetric pairs decoded at the same resolution ( bins).
Surprisingly, the length ratio found by the search is far larger than the – that was optimal for a closed pair, for the binned decoders just as for the closed form. However, this is not a contradiction. The limiting factor in the closed pair case was that the two readings still occupied the same 1D line segment, so whenever two features were co-active, their contributions were mixed together. The moment one of the features is tilted, the co-active case moves into the interior of the parallelogram and can be distinguished by the decoder. But, as mentioned before, this comes at the price of interfering with the features of the second pair, and the best solution of finite resolution accommodates for that by shrinking the feature even more.
To check that this is indeed what sets the ratio, we repeat the asymmetry sweep on the opened geometry: both pairs opened by the angles found by the search ( and ), the length ratio of both swept from to , and every value scored by the same binned decoders as in the previous section. If shrinking the short embeddings really pays for the interference, the minimum should now lie well beyond the – of the closed pair, and should move further out as the bins get finer, since it is the resolution that stops the shrinking. This is what we find: the minimum sits at a ratio of about with bins per axis, with and with . For the finest binning, the MSE stays within of its minimum up to a ratio of , which explains the result obtained by the search.
The MSE of the individual features (middle panel) shows the two sides of the trade-off. At a ratio of 1 all four embeddings still share the same length, and the co-active parallelograms are wide and overlap, which makes the values of the untilted features ambiguous: this interference is the of and . As the tilted embeddings shrink, the parallelograms collapse onto the axes and this error falls to a floor of by a ratio of about . The tilted features and , unlike in the closed pair, hardly pay for shrinking, because the decoder reads them off the line: their error rises to by a ratio of 2, then stays flat until the embeddings become comparable to a bin, and only then climbs steeply. The minimum of the average lies where the long features have gained the most while the short ones are still on their plateau.
We also checked whether the claim from the "Asymmetrical antipodal encoding" section still holds, i.e., whether asymmetry helps within the pair itself, this time with opened embeddings. In the bottom panel we show the sum of the MSEs of the two features of each pair. It turns out that the asymmetry pays off far more than for the closed pair: the MSE of drops from to , and rises again only at the resolution limit.
The MSE of the binned decoders with both pairs opened by the angles found by the search ( and ), as a function of the length ratio. Top: averaged over the four features, for three resolutions (triangles mark the minima). Middle: the MSE of each feature on its own and their average, with bins per axis. Bottom: the MSE summed over the two features of each pair, as in the closed-pair sweep. The dashed line marks the ratio found by the search.
To summarize, the best strategy needs to balance the following three mechanisms:
Making the embeddings of a pair asymmetrical, which pours most of the error into the prediction of one of them, which simultaneously reduces the overall MSE.
Opening the pairs, which makes it possible to distinguish the co-activation of two features from only one of them, or none, being active.
Shortening the tilted embeddings, which minimizes the interference with the features of the second pair.
4. A model's strategy depends on its depth
Knowing what a good strategy is under the general constraints, let us now turn to what the trained models actually do. Below we present representative geometries of models with one and with four bilinear layers.
Representative encoder geometries of trained models with one bilinear layer (left) and four bilinear layers (right). Embeddings of the same antipodal pair share a colour.
To make sure the above results are not an accident, we conducted 20 training runs for each depth, from 1 to 4 bilinear MLP layers, as well as for the tied ReLU model, and checked which strategy each trained model uses. We used sample of size 4,096 and 20,000 training steps in each training.
model
symmetric pairs
asymmetric pairs
asymmetric + opened pairs
other
tied ReLU
19
0
0
1
1 bilinear layer
19
0
0
1
2 bilinear layers
0
16
2
2
3 bilinear layers
0
5
14
1
4 bilinear layers
0
4
10
6
The "other" column gathers the runs whose encoder does not decompose into two antipodal pairs, either because some feature has no roughly opposite partner or because a feature is dropped entirely. It would also collect encoders whose pairs are opened (tilted beyond a small threshold) without being asymmetric, but no run produced such a geometry. Moreover, since opening never appears on its own, the table reports it only jointly with asymmetry, in the "asymmetric + opened pairs" column.
These results clearly show that models with a single bilinear MLP default to symmetric antipodal pairs, and so does the tied ReLU model. Only models with more bilinear MLPs start to use the strategies introduced in the previous section. Why is that? It turns out to boil down to the expressivity of the function class each model can implement.
We can understand this using the benchmark from the previous section.
4.1. The symmetric vs. asymmetric trade-off
Let us first focus on the trade-off between symmetric and asymmetric pairs. We can plot the unconstrained decoders for both choices of the feature embeddings and approximate them with functions of different complexity, corresponding to the model classes we are studying.
The decoder implemented by a model with a single bilinear MLP is a quadratic function. As we can see below, it cannot match the asymmetric benchmark properly, so the model defaults to the symmetric one, which results in a smaller MSE.
A model with a single bilinear MLP, whose decoder is a quadratic in the reading: the best quadratic fit to the unconstrained decoders of the symmetric and the asymmetric pair, and its MSE per feature against the unconstrained floor (dashed).
The decoder of a model with four bilinear layers is far more expressive, because it is a polynomial of degree 16. Such a function can properly approximate both the symmetric and the asymmetric embedding, so the model chooses the asymmetric one, which results in a smaller loss.
A model with four bilinear MLPs, whose decoder is a polynomial of degree 16 in the reading: the same comparison.
Finally, the tied ReLU model faces the same challenge as the model with a single bilinear MLP. It cannot approximate the asymmetric case well, so it also settles on the symmetric solution.
Tied ReLU model: the same comparison.
4.2. Detecting the opening of a pair
We have found that the symmetric vs. asymmetric trade-off comes down to the expressivity of the function class each model can implement. A similar argument explains why some models open their pairs while others do not.
As before, we fix the encoder, this time with one pair slightly opened: is tilted by away from the antipode of , while the other pair stays closed. The readings of co-active features 1 and 3 now fill the thin parallelogram between and , shown in orange below, whereas all other co-active pairs fill the much larger light-blue ones.
One pair opened: is tilted by off the antipode of . The orange parallelogram collects the readings of co-active features 1 and 3. The dashed line is the cut along which we examine the decoder, and the highlighted segment where it crosses the orange region is the strip.
To see what such an encoding demands of the decoder, we cut the plane along a line crossing the orange parallelogram and follow the reconstruction of feature 3 along the cut. Inside the strip, feature 3 is co-active with feature 1, and the decoder should raise its estimate accordingly; the best decoder is therefore a narrow spike over the strip. Outside the strip, its estimate stays near zero, because no pair of co-active features that includes feature 3 can produce a reading on the cut — only the rarer case of three co-active features can place it there with feature 3 among them. We compute the best decoder below as a binned decoder, as in the "Decoder from data samples alone" section. The faint ripple it shows away from the strip is what remains of the triple co-activations, together with the noise of the finite sample.
Below we compare it with the best fits from the function classes of our models. The tied ReLU, quadratic and quartic decoders barely react to the strip at all. The degree-8 decoder of the model with three bilinear MLPs responds with a bump, but a broad one that spills far beyond the strip, so it captures only part of the benefit. This is nevertheless enough for opening to pay off, which is consistent with the table above, where models with three bilinear MLPs are the first to open their pairs in most runs. Only the degree-16 decoder of the model with four bilinear MLPs approximates the spike well. Opening a pair therefore brings no benefit to shallow models, which is why they keep their pairs closed, while deeper models can read the strip and profit from it.
The reconstruction of feature 3 along the cut: the binned decoder (grey) spikes over the strip. The tied ReLU stays at zero, the quadratic and quartic classes barely react, the degree-8 class responds with a broad bump, and only the degree-16 class approximates the spike.
The cut shows only a one-dimensional slice of each decoder. Below we plot the same functions over the whole reading plane, with the cut and the strip drawn on top.
The functions of the previous figure as surfaces over the reading plane.
5. How well do the models approximate the best decoder?
We have given indications of why the models may or may not use different strategies to improve their performance in superposition. But how close do the trained models actually come to these theoretical results?
For this comparison, we retrain the two representative models of the previous section (the same seeds, hence the same initializations) on a much larger sample of inputs instead of the 4,096 used there, keeping the 20,000 training steps.
Below we present the binned decoders for all four features, which the model with a single bilinear MLP is supposed to reconstruct.
Binned decoder ( samples, bins over the reading plane), single bilinear MLP.
Next, we show the best quadratic approximation of these binned decoders.
Best decoder of the class (least-squares quadratic), single bilinear MLP.
Finally, we show the difference between what the model has actually learned and this best decoder of its class. For the model with a single bilinear MLP the difference is surprisingly small!
Trained model minus the best decoder of its class, single bilinear MLP.
The analogous results for the model with four bilinear MLPs are presented below. In this case the difference between the trained model and the best fit in its class is much larger than for the single-MLP model.
Binned decoder, four bilinear MLPs.
Best decoder of the class (least-squares polynomial of degree 16), four bilinear MLPs.
Trained model minus the best decoder of its class, four bilinear MLPs.
To take a closer look, we cut the reading plane of each model along the line of its longest embedding and follow the reconstruction of that feature along the cut. The shallow model traces the best quadratic exactly, while the deeper model visibly departs from the best polynomial of degree 16. Here, too, the difference between the deeper model and the best solution in its class is larger than for the shallow one.
Cut along the line of the longest embedding: the binned decoder, the best decoder of the class and the trained model.
Finally, to quantify the above observations, we define the , i.e., the fraction of the model's error that it could still remove by becoming the best decoder of its class. It is equal to zero when the model already is that decoder.
model
MSE, binned decoder
MSE, best of the class
MSE, trained model
MSE gap
single bilinear MLP, 20,000 steps
0.0061
0.0093
0.0093
0.0%
four bilinear MLPs, 20,000 steps
0.0036
0.0041
0.0060
31.9%
These results underline once more that the shallow model is essentially the best approximation of the binned decoder that its class allows, while the deeper model comes close to it, but not as close as the shallow one.
Is the remaining gap merely the result of a training run that is too short? To check this, we continued the training: starting from the 20,000-step checkpoint of the table above, we trained the model further on the same sample, with a smaller learning rate so that it refines the solution it has already found rather than jumping to a completely new one. The smaller learning rate does not freeze the geometry, though. Over the continuation it keeps slowly drifting within the same strategy, the pairs opening from about and to and and the length ratio of one of them growing from to . As before, each checkpoint is therefore compared with the best decoder of its class for the geometry it has at that moment.
model
MSE, binned decoder
MSE, best of the class
MSE, trained model
MSE gap
four bilinear MLPs, 20,000 steps
0.0036
0.0041
0.0060
31.9%
continued to 150,000 steps
0.0038
0.0040
0.0050
19.9%
continued to 300,000 steps
0.0039
0.0040
0.0049
18.6%
continued to 600,000 steps
0.0039
0.0040
0.0049
18.1%
The MSE gap shrinks, but clearly saturates. Still, one might argue that the culprit is the small learning rate itself: perhaps the model is stuck refining a mediocre solution that a bolder training would escape. To check this, we also retrained the model from scratch, from the same initialization, with the original larger learning rate spread over 150,000, 300,000 and 600,000 steps. Because the learning rate decays over the course of training, stretching the run changes it at every step, so each of these runs follows a completely different trajectory and lands in a different solution. Indeed, the resulting MSE gaps change non-monotonically with the training length — , and — but even the best of them merely matches the at which the continuation saturates, and none improves on it.
The remaining gap of about is therefore not a matter of training length, but of some different phenomenon. We leave it as an open question, since the main purpose of this work was to show that deeper models use asymmetry to their advantage. Nevertheless, understanding why they do not approach the best solutions in their class the way the shallow models do would be a valuable direction for future work.
6. Discussion
The main lesson of this work is that the geometry of superposition is shaped not only by the task, but also by the expressivity of the model at hand. The same data-generating process leads to symmetric antipodal pairs when the network consists of a linear encoder followed by a single bilinear MLP (or when it is the tied ReLU model, which decodes with the transpose of its own encoder matrix followed by a ReLU), while with more bilinear MLPs we obtain non-symmetric geometries. Uniform geometries are therefore not a fact about superposition itself, but a fact about shallow decoders.
Prior work itself hints at this reading. *Toy Models of Superposition* observes deformed, non-uniform geometries only when the features themselves are non-uniform — differing in importance or sparsity, or correlated — while for identical features it finds uniform polytopes. Its authors, however, remark that this structure "seems 'too elegant to be true'" and that "there's a good chance it's at least partly idiosyncratic to the toy model we're investigating". Our results confirm this suspicion and identify the responsible ingredient: identical features alone do not guarantee uniformity — it also takes a decoder too weak to exploit anything better.
If these lessons carry over to real networks, they matter for interpretability. The computation that follows any internal representation of a large model is deep and expressive — far closer to our four-MLP decoder than to a tied ReLU readout — so there is little reason to expect features to be stored as clean, symmetric structures. In particular, we have seen that asymmetry is a useful strategy rather than an accident: two features of equal importance can be embedded with very different norms, so the length of a feature direction need not be a reliable proxy for its importance. The same holds for tilting: an embedding sitting slightly off its expected direction need not be an imprecision of training, but may be a deliberate adjustment that exploits the interplay between the feature embeddings to lower the overall loss.
6.1. Limitations and future work
Our method divides the model into two parts: the first, linear layer is the encoder, and the remaining MLPs form the decoder. We do not yet have a proper understanding of what the individual MLP layers are doing. We believe that feeding the whole geometry output by one layer into the following one, and tracking how it is transformed, would also shed light on the role of the individual layers.
For explanatory reasons, the results shown in this work were limited to a 2D plane at the encoder's output. We would like to check whether our findings survive in higher-dimensional bottlenecks with more features: do the embeddings still organize into asymmetric, slightly tilted antipodal pairs, and do deeper decoders still profit from them? Since such geometries can no longer be simply plotted, this will require replacing pictures with quantitative measures, such as the length ratios and tilts of the pairs and the interference between them.
We found that, contrary to their shallow counterparts, deeper models do not converge to the best possible decoder in their class. We would like to understand precisely why deeper models stop short: for example, whether the obstacle lies in the optimization itself, or in the way the stacked bilinear layers parametrize the polynomials they can in principle express.
7. Summary
We have demonstrated that uniform polytopes are not inherently the best solution to the task of reconstructing independent features of equal importance and sparsity. The earlier indications of their optimality stem from the architectural choices made by the engineers rather than from the task itself.
We have shown two strategies that models can employ to improve on the simple symmetric antipodal embedding of the features: making the embeddings of a pair asymmetric and slightly tilting one of them. Only deeper models can use these strategies, because only their expressivity allows for it.
Finally, we have shown that shallow models are essentially the best possible approximations of the theoretical binned decoders within their class, while deeper models diverge slightly from the best approximation in theirs. This divergence is not an artifact of a training run that is too short: continuing the training of the deeper model shrinks the gap from to about of its error, where it saturates, and retraining it from scratch, with the original larger learning rate annealed over the longer run, changes the gap non-monotonically and does no better. Why deeper models stop short of the best decoders of their class remains an open question.
Code used to generate presented results can be found here.
Acknowledgements
Bartosz Rzepkowski would like to thank Pivotal for their support. This research was carried out during the Pivotal AI Safety Research Fellowship.
Appendix: derivation of the closed-form decoder
The symmetric pair
We derive for the symmetric antipodal pair, where . Recall that each feature is inactive (equal to 0) with probability and otherwise uniformly distributed on , independently of the other.
The best guess at a reading is the average of over all the ways of producing that reading, each way weighted by how likely it is. This is the averaging between scenarios of the main text, and it can be written down exactly. Let stand for the case (neither feature active, feature 1 alone, feature 3 alone, both), for its prior probability and for the density of the readings it produces at . By Bayes' rule, the posterior probability of a case given the reading is
The weight of a case is how often it produces the reading : its prior probability times the density of its readings at . Dividing by the sum of the weights turns them into probabilities. The best guess is then the average of the best guesses of the individual cases, weighted by these posterior probabilities:
The figure below shows these weights and the per-case guesses. One case needs a remark: with neither feature active the reading is exactly , so this case has no density along but a point mass, with the Dirac delta. At it therefore outweighs every other case, and at any it has no weight at all. The rest of this appendix computes the weights and takes the weighted average.
Left: the weight of each case, , i.e. how often it produces the reading ; the case of neither feature active produces only , so its weight is a Dirac delta, , drawn as an arrow in the usual way (its height is not to scale; the number next to it is the probability it carries). Right: the best guess of each case, (for the co-active case, the average over the pairs producing ), and their weighted average, which is the closed-form decoder .
For each case we need , and :
Neither active: . The reading is exactly , so , and . This case alone produces , hence ; the other three cases produce readings .
Feature 1 alone: . The reading is uniform on , so there, and .
Feature 3 alone: . The reading is uniform on , so there, and .
Both active: . Since and are independent and uniform, the pair is spread evenly over the unit square of its possible values (the figure below, left; this is a picture of the input values, not of the reading plane), and the reading is constant along each diagonal line of that square. How often a reading occurs is therefore proportional to the length of its diagonal: it is longest for (the main diagonal from to ) and shrinks linearly to a single corner at (one feature at , the other at ). This gives the triangle on (right).
Left: with both features active, the pair is uniform on the unit square of its possible values, and each reading corresponds to one diagonal line of it. Right: the length of that diagonal, as a function of , is the density of the co-active case.
It remains to find the best guess of the co-active case, , the average of over all the pairs that produce the reading , i.e. over the diagonal . Given , is uniformly distributed over its range on the diagonal, and the mean of a uniform variable is the midpoint of its range. For the diagonal runs from to , so is uniform on with mean . For it runs from to , so is uniform on with mean . In both cases .
With , and of every case in hand, we can now evaluate the weighted average. Only the cases that can produce the reading enter it, which leaves two ranges of to consider.
For the reading can come from "feature 1 alone" or from "both":
For it can come from "feature 3 alone", which contributes , or from "both":
Together with this is the formula of the main text.
The asymmetric pair
The same recipe covers the asymmetric pair, in which the embeddings keep their opposite directions but differ in length. Writing and for the lengths of the short and the long embedding, a reading on the pair line is now
and the symmetric case above is recovered at . The four cases are the same, and only two ingredients of the computation change: the readings of a single feature spread over segments of different lengths, and the density of the co-active case is no longer a triangle.
Neither active: unchanged, the point mass with the guess .
Feature 1 alone: the reading is uniform on , so there, and .
Feature 3 alone: the reading is uniform on , so there, and .
Both active: the pair is again uniform on the unit square, but the lines of constant reading, , are no longer parallel to its diagonal. Near the ends of such a line clips a corner of the square, while in the new, middle range it crosses the square from its left edge to its right edge. This gives three cases. In each of them, the endpoints of the crossing lie on the edges of the square, so they are obtained by fixing the edge's coordinate at or and solving for the other one:
For the line runs from to , so spans , a range of length .
For it runs from to , so spans the whole .
For it runs from to , so spans , a range of length .
Every value of in the span contributes a density of to the reading, because for a fixed , as sweeps , the reading sweeps uniformly over an interval of readings of length . The density is therefore the length of the span divided by , and instead of the triangle we obtain a trapezoid:
Left: with both features active, the pair is uniform on the unit square, and each reading corresponds to one crossing line (here , the ratio used in the main text). Right: the length of that crossing, normalized, is the trapezoidal density of the co-active case.
As before, given the reading, is uniform over its range on the crossing line, with the mean at the midpoint of the spans listed above:
The weighted average then goes through exactly as in the symmetric case, now with three ranges of to consider.
For the reading can come from "feature 3 alone", which contributes , or from "both":
For the same two cases enter, but their densities are now both equal to , so the densities cancel and only the priors remain:
This is the plateau visible in the asymmetric panel of the main text.
For the reading can come from "feature 1 alone" or from "both":
All three reduce to the formulas of the previous section at , with the plateau range collapsing to the single point . As before, the point mass of the inactive case pins . The reconstruction of the other feature needs no separate derivation: swapping the roles of the two features maps onto the same computation with and exchanged and replaced by .
1. Introduction
The current, incredible performance of AI models is closely related to their compression capabilities (Language Modeling Is Compression (Delétang et al., 2023); Compression Represents Intelligence Linearly (Huang et al., 2024)). This compression is imposed on them by the architectural choices made by engineers. For example, GPT-2 had a vocabulary of 50,257 tokens, yet its “operational space” was only of size 768. In such a space, only 768 directions can be described fully independently, so the model had to find strategies to efficiently store all 50,257 input tokens in this reduced space. In general, we call the phenomenon in which a model stores more features than it has dimensions superposition. However, superposition comes at a cost. At a given step of computation, models utilizing it have neurons that fire for multiple different inputs (so-called polysemantic neurons), which makes them hard to interpret.
One might hope to sidestep this by training models wide enough that every feature gets its own dimension (Engineering Monosemanticity in Toy Models (Jermyn et al., 2022)). Besides the cost, it is unclear that this would yield interpretable models, since whether natural features are cleanly separable at all is itself debated (The ‘strong’ feature hypothesis could be wrong (Smith, 2024)). We take a different route and try to understand the phenomenon itself.
Prior work, Toy Models of Superposition (Elhage et al., 2022), studied how superposition forms in two-layer ReLU networks with tied weights. The task was to reconstruct the inputs given to the model (the input features) at its output, after passing them through a narrow bottleneck inside the model. The importance of each feature and the rate at which it occurs (its sparsity) could be varied across experiments, and the quality of the reconstruction was measured with the mean squared error (MSE). The authors suggest that for independent input features of equal importance and sparsity, the best geometries for storing the features may be uniform polytopes.
In this work we study an analogous problem, but with a different model class: a linear encoding layer (the encoder) followed by one or more multilayer perceptrons (MLPs) with a bilinear activation. We treat the stack of MLPs as a single object, the decoder, whose task is to reconstruct the network's input from the geometry produced by the encoding layer.
With this setup we show that the uniform geometries found in Toy Models of Superposition are not the best possible solution to the feature-reconstruction task: non-symmetric geometries yield lower error. We argue that their result is an artifact of the model class they studied, and that deeper, more expressive models do arrive at these asymmetric solutions.
To explain why only sufficiently powerful models find them, we derive two theoretical benchmark decoders: (i) an analytical solution in a simplified case, obtained from the known data-generating process, and (ii) its generalization, approximated from data samples alone. Comparing different models against these benchmarks lets us understand why shallow models default to symmetric solutions while deeper models find the non-symmetric ones.
We hope these results bring us closer to understanding the computations performed by LLMs, since our models and transformer architectures both use a linear encoder as the first step of their computation.
2. Setup
We will be working with the following toy setup:
We divide the model into two parts: the linear layer up to the bottleneck (the encoder) and the stack of MLPs after it (the decoder). This division will also help us derive the theoretical reconstruction floor imposed by the architectural constraints (going from a 4D space down to 2D and back to 4D). The model under study and the division point are shown below. The 2D plane at this point is where all the geometries presented throughout this work live.
For completeness, we note that Toy Models of Superposition used a similar architecture on this very task. For 4 features compressed to a 2D space, their architecture would consists of a linear encoder given by a matrix , followed by a decoder that applies the transpose of the same matrix , a bias, and finally a ReLU activation. Because the same weight matrix is used in both layers, we will refer to this case as tied ReLU.
3. Theoretical solution
In this section, we will introduce a simple, symmetric geometry that a theoretical encoder can implement, and provide a closed-form description of a decoder whose goal is to reconstruct the network’s input from the point in the 2D plane that the encoder maps it to.
We will then present two strategies the encoder can employ to lower the reconstruction error, and finally demonstrate how these two techniques can be used in tandem to obtain the true best solution. This solution will serve as the theoretical reconstruction floor against which we will later compare the behavior of the actual models under study.
3.1. Symmetrical antipodal encoding
The simplest way for the encoder to arrange the embeddings of the 4 input features on the 2D plane is as two antipodal pairs. The embeddings of each pair lie on a common line, pointing in opposite directions. The two lines are orthogonal, which allows the pairs to be analyzed independently, and their orientation is otherwise arbitrary. Below we show an example where we have simply chosen the basis axes as the pair lines. Here denotes the embedding of feature , i.e. the 2D vector the encoder maps the -th input feature to. It can be read directly from the encoder matrix as its -th column. There are four such vectors, , one per feature, and the encoder's output (which is also the reading seen by the decoder) is their weighted sum
The embeddings of the input features placed as two antipodal pairs on orthogonal axes.
The two pair lines are orthogonal and placed on the basis axes, so the horizontal coordinate of the reading is simply , regardless of what features 2 and 4 are doing. When reconstructing and we can therefore restrict our attention only to , the position of the reading along the segment between and , with at , at and at the origin.
We will use this line segment to understand the interaction between the introduced encoder and a theoretical decoder through two lenses: (i) knowing the distribution of the features exactly, and (ii) having access only to an "oracle" that gives us samples of the data with no description of the underlying distribution. In both analyses, we assume that the loss function, the MSE defined above, is known.
3.1.1. Closed-form decoder from the known distribution
Knowing the distribution of the input features, we can distinguish four cases that result in a reading on this segment (as a reminder, we use the probability of a feature being active throughout this work):
1. neither feature active, with probability ,
2. feature 1 alone active, with probability ,
3. feature 3 alone active, with probability ,
4. both features 1 and 3 active, with probability .
This encoding method has one huge caveat. In the fourth case, with both features co-active, the reading can land anywhere on the segment (except for the endpoints and themselves). Because of that, even a perfect decoder cannot distinguish this case with full confidence from one in which only a single feature, or none, was active.
The four cases on the segment between and , each represented by a black dot at different position. Example collision of two cases: feature 1 alone and features 1 and 3 together can produce the same reading.
Because of this information loss, the decoder needs to average between different plausible scenarios to give the best possible response (measured with MSE): the best guess at a reading is the average of the guesses each scenario would make, weighted by how likely each scenario is to have produced that reading. The Appendix makes this precise.
Below we present the formula for such a theoretical decoder reconstructing the first feature, when given a reading at position on the segment. Full derivation of this formula is presented in the Appendix, but here we just emphasize one phenomenon: a reading of exactly can only come from the no-feature case, so the decoder outputs there, even though the limits from either side are . This reading is marked with a black dot in the plot below.
The best decoder for on the antipodal pair, as a function of the reading on the segment between and . Open circle: the one-sided limits at ; closed circle: the value at itself.
The decoder for the third feature is simply the mirror image of the curve above, , so we can picture the two decoders on the same plot:
The best decoders for (blue) and (orange) on the antipodal pair, as functions of the reading .
3.1.2. Decoder from data samples alone
The above derivation was possible only because we knew the underlying distribution of the features. Now, let us try to obtain similar results while having access only to an "oracle" that gives us raw samples of the data (which is closer to a real-world situation).
Having samples of the data, we can also visualize them along our line segment:
4096 samples of the pair drawn from the sampling process, plotted as the reading against the value the decoder has to predict.
The above plot can make it more explicit that without the case of co-active features the best possible solution would be to use a ReLU function.
Now, how can we accommodate for the fact that there are situations in which different features are co-active? To do that we need to look at the definition of the loss function we are using.
In our setup the loss is the MSE. Recall that the decoder sees only the position of the reading, and that many different inputs produce the same . For example, with feature 3 inactive, and together with , both give . The decoder must therefore output a single number for all samples with that , and its performance is measured by the squared distance from that number to each sample's true . The number that minimizes the total squared distance to a set of values is their arithmetic mean, so the best decoder outputs the mean of over the samples with reading .
So, the choice of the mean thus comes from the loss, not from the data. If we chose the mean absolute error instead, the best output would be the median of the same set of values, and with a 0–1 loss its most frequent value (the mode).
However, with a finite number samples we cannot take the mean over "all samples with reading " literally, because no two samples have the same reading. Instead we split the line segment into very small intervals (bins), gather the samples whose reading falls into each, and average their . The result is a step function approximating the closed-form curve, and it converges to it as we take more samples and narrower bins. Below we present the result of such an average computed on the data samples shown above:
Binned means of the 4096 samples (100 bins) against the closed-form decoder . The blue dot at is the mean over the samples with reading exactly .
Having introduced the two methods of obtaining the best possible decoder depending on the context we are in, we now move on to the strategies, which can be applied to modify the initial encoder and increase the performance of the model as a whole.
3.2. Asymmetrical antipodal encoding
The first modification is very simple: we make one of the embeddings shorter than the other.
The symmetric antipodal pair and its asymmetric version, in which is shortened and lengthened ( ).
This has a rather counterintuitive outcome. With this change the prediction error is "poured" into the reconstruction of a single feature; however, measured over both features, the total error turns out to be smaller!
Below we compare the two encodings. In the asymmetric one, is shortened and lengthened so that is times . The first two panels show the closed-form decoders for both features, and the third the MSE these decoders achieve, split into the contributions of the two features.
Symmetric vs asymmetric antipodal pair ( ): the closed-form decoders for and , and the MSE of the best decoder split into the two features' contributions.
Is it then always better to make the pair more asymmetric? To check this, we sweep the length ratio from to and compute, for each value, the MSE of the best decoder, both in closed form and from binned samples. In the binned case, the segment is divided into bins of equal width and every feature gets its own decoder, which predicts the feature's mean value over the samples that land in the same bin. In both cases, as in the previous figure, we report the sum of the two features' MSEs rather than their average.
The summed MSE of and under the best decoder as a function of the length ratio . Top: the closed-form decoder, in total and split into the two features' contributions. Bottom: the closed form against binned decoders of three resolutions, one decoder per feature on the same bins, their MSEs summed. Triangles mark the minima. The dashed line marks the ratio used above.
The closed-form curve first falls, because the error of the lengthened feature vanishes faster than the error of the shortened one grows, but it flattens out at a ratio of about – (a minimum of at , against for the symmetric pair) and then slowly rises again, toward : beyond this point the shortened feature is essentially unreadable whenever its partner is active, and its error dominates the total. Asymmetry therefore helps only up to a moderate ratio, and the optimum is shallow. A binned decoder has its minimum at the same ratio ( – for the three resolutions), but is stricter beyond it: once the short embedding becomes comparable to the width of a bin, the readings of the short feature on its own are no longer resolved, and the MSE climbs steeply, the earlier the coarser the bins (with bins the symmetric pair is better again above a ratio of , with bins above ).
3.3. "Opening" the antipodal pairs
Another modification boils down to slightly tilting one of the embeddings of a pair with respect to its partner. With this strategy, the case of no active features still lands at the origin, and each feature active on its own still lands on its embedding vector, but the co-active case moves off the line into the interior of the parallelogram spanned by the two embeddings.
This allows the decoder to distinguish features that are active on their own from co-active ones, which results in better reconstruction performance. In fact, a decoder of infinite resolution would be able to distinguish the co-activation at an arbitrarily small opening angle, since the co-active readings leave the line as soon as . This does not make the problem trivial, though. Opening a pair does not change the dimensionality of the problem: all four embeddings still live in the same 2D plane, so at most two of them can be linearly independent. A single opened pair, taken alone, would indeed be perfectly decodable, but the other pair shares the plane, and the interior of the parallelogram is also reachable by other combinations of features. The benefit of opening is therefore limited by two things: the interference with the other pair, which we return to at the end of this section, and the resolution of the decoder, which we turn to now.
Trained models have finite resolution. Below we investigate its effects by studying what a binned decoder is actually capable of. We keep the length ratio of both pairs equal to , tilt off the antipode of by , keep the other pair closed, and score every with binned decoders of three resolutions. The bins are now square cells covering the 2D plane of readings, and again every feature gets its own decoder, which predicts the feature's mean value over the samples that land in the same cell. Since all four features are now involved, we report the MSE as defined in the Setup, i.e. averaged over the four features.
The MSE of the binned decoder as a function of the opening angle of the pair , the pair closed, averaged over the four features, for three resolutions, one decoder per feature on the same cells (triangles mark the minima).
The MSE has its minimum at a finite opening, which moves toward zero as the bins get finer. But why does the MSE not stay constant once surpasses this threshold for a given binning? This is because opening one pair makes it interfere with the readings of the other pair. We will study this in more depth in the next section.
3.4. Search for the best strategy
We close this section by investigating what is in fact the best strategy for a linear encoder followed by a decoder that is restricted to no particular function class and is limited only by the resolution of the bins used to compute it.
We do this by a free search over the encoder geometry. For an unconstrained decoder the angle between the two pair lines does not matter, so we may pin two of the embeddings to the axes and let the remaining two be arbitrary vectors in the plane, described by their angle and length. Every geometry is scored by the MSE of a decoder obtained numerically, as in "Decoder from data samples alone" but with two-dimensional bins. The four parameters are optimized by a random search followed by local descent.
Below we present the geometry obtained by the above search along with regions indicating the co-occurrence of various feature pairs.
The geometry found by the search, with a zoom on the origin as an inset: the pair opened by at length ratio , the pair by at . Each shaded parallelogram is the region where the readings of two co-active features land; the two thin orange slivers along the axes belong to the pairs.
The search finds that the two free embeddings settle almost antipodal to the pinned ones, much shorter than them, and slightly tilted, each toward the long embedding of the other pair. The pair is opened by with a length ratio , and the pair by with , reaching an MSE floor of per feature, against for the closed symmetric pairs decoded at the same resolution ( bins).
Surprisingly, the length ratio found by the search is far larger than the – that was optimal for a closed pair, for the binned decoders just as for the closed form. However, this is not a contradiction. The limiting factor in the closed pair case was that the two readings still occupied the same 1D line segment, so whenever two features were co-active, their contributions were mixed together. The moment one of the features is tilted, the co-active case moves into the interior of the parallelogram and can be distinguished by the decoder. But, as mentioned before, this comes at the price of interfering with the features of the second pair, and the best solution of finite resolution accommodates for that by shrinking the feature even more.
To check that this is indeed what sets the ratio, we repeat the asymmetry sweep on the opened geometry: both pairs opened by the angles found by the search ( and ), the length ratio of both swept from to , and every value scored by the same binned decoders as in the previous section. If shrinking the short embeddings really pays for the interference, the minimum should now lie well beyond the – of the closed pair, and should move further out as the bins get finer, since it is the resolution that stops the shrinking. This is what we find: the minimum sits at a ratio of about with bins per axis, with and with . For the finest binning, the MSE stays within of its minimum up to a ratio of , which explains the result obtained by the search.
The MSE of the individual features (middle panel) shows the two sides of the trade-off. At a ratio of 1 all four embeddings still share the same length, and the co-active parallelograms are wide and overlap, which makes the values of the untilted features ambiguous: this interference is the of and . As the tilted embeddings shrink, the parallelograms collapse onto the axes and this error falls to a floor of by a ratio of about . The tilted features and , unlike in the closed pair, hardly pay for shrinking, because the decoder reads them off the line: their error rises to by a ratio of 2, then stays flat until the embeddings become comparable to a bin, and only then climbs steeply. The minimum of the average lies where the long features have gained the most while the short ones are still on their plateau.
We also checked whether the claim from the "Asymmetrical antipodal encoding" section still holds, i.e., whether asymmetry helps within the pair itself, this time with opened embeddings. In the bottom panel we show the sum of the MSEs of the two features of each pair. It turns out that the asymmetry pays off far more than for the closed pair: the MSE of drops from to , and rises again only at the resolution limit.
The MSE of the binned decoders with both pairs opened by the angles found by the search ( and ), as a function of the length ratio. Top: averaged over the four features, for three resolutions (triangles mark the minima). Middle: the MSE of each feature on its own and their average, with bins per axis. Bottom: the MSE summed over the two features of each pair, as in the closed-pair sweep. The dashed line marks the ratio found by the search.
To summarize, the best strategy needs to balance the following three mechanisms:
4. A model's strategy depends on its depth
Knowing what a good strategy is under the general constraints, let us now turn to what the trained models actually do. Below we present representative geometries of models with one and with four bilinear layers.
Representative encoder geometries of trained models with one bilinear layer (left) and four bilinear layers (right). Embeddings of the same antipodal pair share a colour.
To make sure the above results are not an accident, we conducted 20 training runs for each depth, from 1 to 4 bilinear MLP layers, as well as for the tied ReLU model, and checked which strategy each trained model uses. We used sample of size 4,096 and 20,000 training steps in each training.
model
symmetric pairs
asymmetric pairs
asymmetric + opened pairs
other
tied ReLU
19
0
0
1
1 bilinear layer
19
0
0
1
2 bilinear layers
0
16
2
2
3 bilinear layers
0
5
14
1
4 bilinear layers
0
4
10
6
The "other" column gathers the runs whose encoder does not decompose into two antipodal pairs, either because some feature has no roughly opposite partner or because a feature is dropped entirely. It would also collect encoders whose pairs are opened (tilted beyond a small threshold) without being asymmetric, but no run produced such a geometry. Moreover, since opening never appears on its own, the table reports it only jointly with asymmetry, in the "asymmetric + opened pairs" column.
These results clearly show that models with a single bilinear MLP default to symmetric antipodal pairs, and so does the tied ReLU model. Only models with more bilinear MLPs start to use the strategies introduced in the previous section. Why is that? It turns out to boil down to the expressivity of the function class each model can implement.
We can understand this using the benchmark from the previous section.
4.1. The symmetric vs. asymmetric trade-off
Let us first focus on the trade-off between symmetric and asymmetric pairs. We can plot the unconstrained decoders for both choices of the feature embeddings and approximate them with functions of different complexity, corresponding to the model classes we are studying.
The decoder implemented by a model with a single bilinear MLP is a quadratic function. As we can see below, it cannot match the asymmetric benchmark properly, so the model defaults to the symmetric one, which results in a smaller MSE.
A model with a single bilinear MLP, whose decoder is a quadratic in the reading: the best quadratic fit to the unconstrained decoders of the symmetric and the asymmetric pair, and its MSE per feature against the unconstrained floor (dashed).
The decoder of a model with four bilinear layers is far more expressive, because it is a polynomial of degree 16. Such a function can properly approximate both the symmetric and the asymmetric embedding, so the model chooses the asymmetric one, which results in a smaller loss.
A model with four bilinear MLPs, whose decoder is a polynomial of degree 16 in the reading: the same comparison.
Finally, the tied ReLU model faces the same challenge as the model with a single bilinear MLP. It cannot approximate the asymmetric case well, so it also settles on the symmetric solution.
Tied ReLU model: the same comparison.
4.2. Detecting the opening of a pair
We have found that the symmetric vs. asymmetric trade-off comes down to the expressivity of the function class each model can implement. A similar argument explains why some models open their pairs while others do not.
As before, we fix the encoder, this time with one pair slightly opened: is tilted by away from the antipode of , while the other pair stays closed. The readings of co-active features 1 and 3 now fill the thin parallelogram between and , shown in orange below, whereas all other co-active pairs fill the much larger light-blue ones.
One pair opened: is tilted by off the antipode of . The orange parallelogram collects the readings of co-active features 1 and 3. The dashed line is the cut along which we examine the decoder, and the highlighted segment where it crosses the orange region is the strip.
To see what such an encoding demands of the decoder, we cut the plane along a line crossing the orange parallelogram and follow the reconstruction of feature 3 along the cut. Inside the strip, feature 3 is co-active with feature 1, and the decoder should raise its estimate accordingly; the best decoder is therefore a narrow spike over the strip. Outside the strip, its estimate stays near zero, because no pair of co-active features that includes feature 3 can produce a reading on the cut — only the rarer case of three co-active features can place it there with feature 3 among them. We compute the best decoder below as a binned decoder, as in the "Decoder from data samples alone" section. The faint ripple it shows away from the strip is what remains of the triple co-activations, together with the noise of the finite sample.
Below we compare it with the best fits from the function classes of our models. The tied ReLU, quadratic and quartic decoders barely react to the strip at all. The degree-8 decoder of the model with three bilinear MLPs responds with a bump, but a broad one that spills far beyond the strip, so it captures only part of the benefit. This is nevertheless enough for opening to pay off, which is consistent with the table above, where models with three bilinear MLPs are the first to open their pairs in most runs. Only the degree-16 decoder of the model with four bilinear MLPs approximates the spike well. Opening a pair therefore brings no benefit to shallow models, which is why they keep their pairs closed, while deeper models can read the strip and profit from it.
The reconstruction of feature 3 along the cut: the binned decoder (grey) spikes over the strip. The tied ReLU stays at zero, the quadratic and quartic classes barely react, the degree-8 class responds with a broad bump, and only the degree-16 class approximates the spike.
The cut shows only a one-dimensional slice of each decoder. Below we plot the same functions over the whole reading plane, with the cut and the strip drawn on top.
The functions of the previous figure as surfaces over the reading plane.
5. How well do the models approximate the best decoder?
We have given indications of why the models may or may not use different strategies to improve their performance in superposition. But how close do the trained models actually come to these theoretical results?
For this comparison, we retrain the two representative models of the previous section (the same seeds, hence the same initializations) on a much larger sample of inputs instead of the 4,096 used there, keeping the 20,000 training steps.
Below we present the binned decoders for all four features, which the model with a single bilinear MLP is supposed to reconstruct.
Binned decoder ( samples, bins over the reading plane), single bilinear MLP.
Next, we show the best quadratic approximation of these binned decoders.
Best decoder of the class (least-squares quadratic), single bilinear MLP.
Finally, we show the difference between what the model has actually learned and this best decoder of its class. For the model with a single bilinear MLP the difference is surprisingly small!
Trained model minus the best decoder of its class, single bilinear MLP.
The analogous results for the model with four bilinear MLPs are presented below. In this case the difference between the trained model and the best fit in its class is much larger than for the single-MLP model.
Binned decoder, four bilinear MLPs.
Best decoder of the class (least-squares polynomial of degree 16), four bilinear MLPs.
Trained model minus the best decoder of its class, four bilinear MLPs.
To take a closer look, we cut the reading plane of each model along the line of its longest embedding and follow the reconstruction of that feature along the cut. The shallow model traces the best quadratic exactly, while the deeper model visibly departs from the best polynomial of degree 16. Here, too, the difference between the deeper model and the best solution in its class is larger than for the shallow one.
Cut along the line of the longest embedding: the binned decoder, the best decoder of the class and the trained model.
Finally, to quantify the above observations, we define the , i.e., the fraction of the model's error that it could still remove by becoming the best decoder of its class. It is equal to zero when the model already is that decoder.
model
MSE, binned decoder
MSE, best of the class
MSE, trained model
MSE gap
single bilinear MLP, 20,000 steps
0.0061
0.0093
0.0093
0.0%
four bilinear MLPs, 20,000 steps
0.0036
0.0041
0.0060
31.9%
These results underline once more that the shallow model is essentially the best approximation of the binned decoder that its class allows, while the deeper model comes close to it, but not as close as the shallow one.
Is the remaining gap merely the result of a training run that is too short? To check this, we continued the training: starting from the 20,000-step checkpoint of the table above, we trained the model further on the same sample, with a smaller learning rate so that it refines the solution it has already found rather than jumping to a completely new one. The smaller learning rate does not freeze the geometry, though. Over the continuation it keeps slowly drifting within the same strategy, the pairs opening from about and to and and the length ratio of one of them growing from to . As before, each checkpoint is therefore compared with the best decoder of its class for the geometry it has at that moment.
model
MSE, binned decoder
MSE, best of the class
MSE, trained model
MSE gap
four bilinear MLPs, 20,000 steps
0.0036
0.0041
0.0060
31.9%
continued to 150,000 steps
0.0038
0.0040
0.0050
19.9%
continued to 300,000 steps
0.0039
0.0040
0.0049
18.6%
continued to 600,000 steps
0.0039
0.0040
0.0049
18.1%
The MSE gap shrinks, but clearly saturates. Still, one might argue that the culprit is the small learning rate itself: perhaps the model is stuck refining a mediocre solution that a bolder training would escape. To check this, we also retrained the model from scratch, from the same initialization, with the original larger learning rate spread over 150,000, 300,000 and 600,000 steps. Because the learning rate decays over the course of training, stretching the run changes it at every step, so each of these runs follows a completely different trajectory and lands in a different solution. Indeed, the resulting MSE gaps change non-monotonically with the training length — , and — but even the best of them merely matches the at which the continuation saturates, and none improves on it.
The remaining gap of about is therefore not a matter of training length, but of some different phenomenon. We leave it as an open question, since the main purpose of this work was to show that deeper models use asymmetry to their advantage. Nevertheless, understanding why they do not approach the best solutions in their class the way the shallow models do would be a valuable direction for future work.
6. Discussion
The main lesson of this work is that the geometry of superposition is shaped not only by the task, but also by the expressivity of the model at hand. The same data-generating process leads to symmetric antipodal pairs when the network consists of a linear encoder followed by a single bilinear MLP (or when it is the tied ReLU model, which decodes with the transpose of its own encoder matrix followed by a ReLU), while with more bilinear MLPs we obtain non-symmetric geometries. Uniform geometries are therefore not a fact about superposition itself, but a fact about shallow decoders.
Prior work itself hints at this reading. *Toy Models of Superposition* observes deformed, non-uniform geometries only when the features themselves are non-uniform — differing in importance or sparsity, or correlated — while for identical features it finds uniform polytopes. Its authors, however, remark that this structure "seems 'too elegant to be true'" and that "there's a good chance it's at least partly idiosyncratic to the toy model we're investigating". Our results confirm this suspicion and identify the responsible ingredient: identical features alone do not guarantee uniformity — it also takes a decoder too weak to exploit anything better.
If these lessons carry over to real networks, they matter for interpretability. The computation that follows any internal representation of a large model is deep and expressive — far closer to our four-MLP decoder than to a tied ReLU readout — so there is little reason to expect features to be stored as clean, symmetric structures. In particular, we have seen that asymmetry is a useful strategy rather than an accident: two features of equal importance can be embedded with very different norms, so the length of a feature direction need not be a reliable proxy for its importance. The same holds for tilting: an embedding sitting slightly off its expected direction need not be an imprecision of training, but may be a deliberate adjustment that exploits the interplay between the feature embeddings to lower the overall loss.
6.1. Limitations and future work
7. Summary
We have demonstrated that uniform polytopes are not inherently the best solution to the task of reconstructing independent features of equal importance and sparsity. The earlier indications of their optimality stem from the architectural choices made by the engineers rather than from the task itself.
We have shown two strategies that models can employ to improve on the simple symmetric antipodal embedding of the features: making the embeddings of a pair asymmetric and slightly tilting one of them. Only deeper models can use these strategies, because only their expressivity allows for it.
Finally, we have shown that shallow models are essentially the best possible approximations of the theoretical binned decoders within their class, while deeper models diverge slightly from the best approximation in theirs. This divergence is not an artifact of a training run that is too short: continuing the training of the deeper model shrinks the gap from to about of its error, where it saturates, and retraining it from scratch, with the original larger learning rate annealed over the longer run, changes the gap non-monotonically and does no better. Why deeper models stop short of the best decoders of their class remains an open question.
Code used to generate presented results can be found here.
Acknowledgements
Bartosz Rzepkowski would like to thank Pivotal for their support. This research was carried out during the Pivotal AI Safety Research Fellowship.
Appendix: derivation of the closed-form decoder
The symmetric pair
We derive for the symmetric antipodal pair, where . Recall that each feature is inactive (equal to 0) with probability and otherwise uniformly distributed on , independently of the other.
The best guess at a reading is the average of over all the ways of producing that reading, each way weighted by how likely it is. This is the averaging between scenarios of the main text, and it can be written down exactly. Let stand for the case (neither feature active, feature 1 alone, feature 3 alone, both), for its prior probability and for the density of the readings it produces at . By Bayes' rule, the posterior probability of a case given the reading is
The weight of a case is how often it produces the reading : its prior probability times the density of its readings at . Dividing by the sum of the weights turns them into probabilities. The best guess is then the average of the best guesses of the individual cases, weighted by these posterior probabilities:
The figure below shows these weights and the per-case guesses. One case needs a remark: with neither feature active the reading is exactly , so this case has no density along but a point mass, with the Dirac delta. At it therefore outweighs every other case, and at any it has no weight at all. The rest of this appendix computes the weights and takes the weighted average.
Left: the weight of each case, , i.e. how often it produces the reading ; the case of neither feature active produces only , so its weight is a Dirac delta, , drawn as an arrow in the usual way (its height is not to scale; the number next to it is the probability it carries). Right: the best guess of each case, (for the co-active case, the average over the pairs producing ), and their weighted average, which is the closed-form decoder .
For each case we need , and :
Left: with both features active, the pair is uniform on the unit square of its possible values, and each reading corresponds to one diagonal line of it. Right: the length of that diagonal, as a function of , is the density of the co-active case.
It remains to find the best guess of the co-active case, , the average of over all the pairs that produce the reading , i.e. over the diagonal . Given , is uniformly distributed over its range on the diagonal, and the mean of a uniform variable is the midpoint of its range. For the diagonal runs from to , so is uniform on with mean . For it runs from to , so is uniform on with mean . In both cases .
With , and of every case in hand, we can now evaluate the weighted average. Only the cases that can produce the reading enter it, which leaves two ranges of to consider.
The asymmetric pair
The same recipe covers the asymmetric pair, in which the embeddings keep their opposite directions but differ in length. Writing and for the lengths of the short and the long embedding, a reading on the pair line is now
and the symmetric case above is recovered at . The four cases are the same, and only two ingredients of the computation change: the readings of a single feature spread over segments of different lengths, and the density of the co-active case is no longer a triangle.
Every value of in the span contributes a density of to the reading, because for a fixed , as sweeps , the reading sweeps uniformly over an interval of readings of length . The density is therefore the length of the span divided by , and instead of the triangle we obtain a trapezoid:
Left: with both features active, the pair is uniform on the unit square, and each reading corresponds to one crossing line (here , the ratio used in the main text). Right: the length of that crossing, normalized, is the trapezoidal density of the co-active case.
As before, given the reading, is uniform over its range on the crossing line, with the mean at the midpoint of the spans listed above:
The weighted average then goes through exactly as in the symmetric case, now with three ranges of to consider.
This is the plateau visible in the asymmetric panel of the main text.
All three reduce to the formulas of the previous section at , with the plateau range collapsing to the single point . As before, the point mass of the inactive case pins . The reconstruction of the other feature needs no separate derivation: swapping the roles of the two features maps onto the same computation with and exchanged and replaced by .