Summary Neural networks are widely assumed to use superposition to represent more features than they have dimensions. A stronger claim is that they also compute in superposition (CiS), i.e., implement more nonlinear functions than they have neurons (Hänni et al. 2024). CiS remains poorly understood, and until recently there were...
LLMs are commonly assumed to use superposition to represent more features than they have dimensions. The evidence for this is mostly indirect — chiefly the success of SAEs at extracting interpretable directions. A stronger claim is that models also compute in superposition, and for that we have only theoretical evidence....
Note: This is a research update sharing preliminary results as part of ongoing work. Figure 1: Contrastive (difference-of-means, English→Mandarin) feature directions elicit a downstream response at much smaller perturbation magnitudes than SAE directions, which behave similarly to random directions. This holds across multiple models and experimental setups. Summary & Main...