In a recent post, we presented PIRAMID, its leadership and research pillars, and a plan for how they fit together. In this post, we sketch a team-by-team account of progress and targets for the next 6–12 months. We include results to date as evidence of viability: we’ve been a small team, with much of the past year spent building behind the scenes, and we aim to greatly accelerate our progress over the coming year as we expand our efforts and our teams. Like any fundamental scientific effort, none of this is set in stone. We expect some of these bets to need revision and are confident in our ability to reassess and change course as new evidence comes to light.
If you’re interested in collaborating or supporting our work as we expand, please get in touch.
Advancements in Learning Theory
We aim to formulate statistical and mesoscopic theories of feature learning and generalization which provide a level of description between microscopic parameter-level dynamics and macroscopic performance metrics. Our work so far has treated statistical physics (e.g., mean field methods) as a candidate language for statistically describing learned structure, and covariance measures between neurons or weights as candidate mesoscopic objects that track meaningful scales and determine feature relevance during training.
We’ve developed and tested the expressive power of this language in three directions. The first lays the foundation for mean-field theory’s potential and utility. Our most recent pre-print argues that Bayesian mean-field theory is complexity complete, in the sense that it captures the aggregate heuristic bias of finite-width neural networks. In parallel, we have also been distilling mean-field theory for a broader audience, in the hopes that more interpretability practitioners will use and improve this framework. A detailed post on mean field theory and computation in superposition connects field-theoretic descriptions of neural networks with sparse and superposed structure, which are central to empirical interpretability but not yet well integrated into a field-theoretic approach to learning theory.
Second, we are developing a mean-field model of feature learning with nontrivial hidden-layer statistics. Most existing theory, including dynamical mean-field work that tracks evolving kernels and finite-width fluctuations (e.g., Bordelon & Pehlevan, Rubin et al.), treats feature learning through the predictor or through kernels and order parameters that summarize input–output behavior. Instead, our work incorporates hidden-layer structure (characterized by various scales of interaction like non-trivial neuron-neuron covariance, or how localized or distributed a representation is) in the minimal state description. A theory sensitive to the internal structures formed during training could explain why some models learn robust, reusable features while others remain brittle, and could connect representation geometry to monitoring, transfer, and post-training stability. In developing meaningful observables (e.g., similar to susceptibilities), such theories can help us develop tools to more faithfully interpret model internals and guide our understanding of which theoretical limits or toy models hold the most weight in a given setting.
A third line of work aims to develop a theory of error geometry, inspired by past work from Mei and Montanari (e.g., this) and Canatar et al.. Recent progress (to appear soon) shows how correlations between training and test points affect which related examples are learned or missed together in a random-feature model. This project asks a pointwise question that most generalization theory averages away: when a test example is correlated with a specific subset of the training set, how does that correlation structure determine its error? For modern web-scale models, the practically relevant regime is not the idealized setting in which every test point is equally novel. Many important failures occur near pockets of the training distribution, or along directions that are only partially novel. A pointwise theory of error geometry would therefore sharpen more than standard generalization theory: it would give a language for when models interpolate safely, memorize dangerously, or fail on targeted shifts.
Together, these projects move toward a common pipeline: prove statements in idealized stochastic models, identify the induced statistics of features or errors, and test those statistics in realistic networks.
Where this is going. This work depends on the use of toy models and model organisms that act as simplified settings — simple feature-learning models, mean-field models with nontrivial hidden-layer statistics, superposition-like toy models, generalization properties of mean field models, and linear or quadratic models of training dynamics — which can isolate a qualitative or quantitative component of interpretability’s (currently large) theory-practice gap. Our approach is to increase the breadth of our theoretical understanding and its applicability to realistic models by coming up with ‘structural gates’ an idealized model must pass through (incorporate) to improve their utility while avoiding unnecessary complexity. For example, many of our current projects use existing theory at infinite width as an intuition pump with the aim of adding increased complexity in the form of these gates. Designing the gates – criteria for a flavor of difficulty that current theories or tools cannot handle – is a core part of this research and the subject of a post to appear soon.
Our strategy for future work operates along two complementary paths:
Direct Idealization of Toy Tasks (toy model): Consider a toy or formal task which is susceptible to exact analysis and analyze it formally, then see if insights or methods generalize.
Indirect Realizations of Realistic Tasks (model organism): Start from a realistic model — say, GPT-2 small — and construct a theoretically tractable analog that is at least as complex and generalizes comparably, which can be used to reason about the original.
The practical goal is to turn these theories into measurements that are useful for the interpretability and safety of real-world models, reflecting in the following targets:
Target 1: Error geometry in the feature learning regime. The theory described above predicts, a priori, the pattern of correlated successes and failures for a trained random-feature model. We extend this in two directions: making it dynamical (how does the error on a partially correlated test point evolve throughout training, and are there distinct time scales for memorization-like versus generalization-like behavior?) and moving beyond frozen features (in regimes where the effective kernel itself evolves during training, which parts of the random-feature geometry survive and which are replaced by genuine feature-learning effects?) Empirical tests should track, for each held-out point, its similarity profile to the training set, its layer-wise representation trajectory, and its instantaneous prediction error across training time.
Target 2: A better measure of features. These are objects that exist both in data/function space and in activation/parameter space, and which can include distributed, superposed, or circuit-like structure rather than only sparse activation directions. Modeling them requires both an understanding at both a functional level (kernel correction) and a representational level (in terms of neuron marginals and covariances, symmetry breaking and specialization, and lottery ticket-like sub-circuit formation). Defining a zoology of ‘features’, and understanding how and when different forms occur, bears directly on understanding when interpretability is possible at all and which methods can work in a given regime. We think that a good theory of features – defined here as atoms of a representation or the model’s memory/ storage at a given layer – should also couple to a good notion of a circuit (i.e., a decomposition of the model computation into atoms). An advantage of mean field theory techniques is that they inherently have the expressive power to link representational and computational techniques (as in work by Rubin et al. 2023 and Rubin et al. 2025). This suggests a possible interpretability tool that searches for feature maps and sparse “connectomes” between them: low-dimensional functions of layer activations that recover the intermediate structure of a learned computation. A first testbed for this agenda is circuit reconstruction. Given a network trained to implement a known Boolean circuit, modular arithmetic algorithm, or other controlled computation, can we recover the latent gates, feature maps, and parent-child relationships between them from the trained model? This provides a concrete model organism for asking whether mean-field-style summaries can recover mechanistic structure without relying entirely on neuron-level decompositions or SAE-style sparsity assumptions.
Target 3: Signals of Feature Learning. Lastly, we will explore methods to track how features form, change, and interact during training\footnote{This is similar in its aims to the developmental interpretability community.}. What are the various time scales of feature formation, and to what extent is feature learning stationary? Perhaps the most safety-relevant target here, related to silent alignment, is the gap between a feature’s acquisition and its behavioral expression. The practical goals are theory-inspired early-detection measures for such signals, and a taxonomy of classes of sudden change in learning dynamics paired with the appropriate detection techniques.
Interpretability Applications
What does it mean for interpretability to be ambitious? Aiming for a complete understanding of neural network components – in our case, measured by physics-informed faithfulness criteria – is an ambitioustargetin its own right. However, even a partial understanding can be a proxy target that is ambitiousin its application, provided it moves the needle for robust, scalable alignment and is held to the same faithfulness standard. This second class presupposes that models decompose into multiple independent generalizations, organized into basins (in the sense of the loss landscape geometry) that are ‘aligned’ by some suitably rigorous metric. If this is true, interpretability becomes a means of constructively shaping model behavior – steering the model toward more aligned basins during training, e.g., early in RL – rather than only an auditing tool. These training-time interventions are especially relevant for preventing nefarious phenomena like exploration hacking or deceptive alignment, for which post-hoc detection is insufficient.
These targets (both direct and proxy) shape our strategy in two ways:
From the top-down, we build interpretability into model architectures from scratch, designing sparse transformers to make learned computation legible and relatively efficient by construction. Success could eliminate the need for (and cost of) some post-hoc methods. We believe this approach to be neglected, challenging, and very high EV. In contrast to incremental changes made by previous interpretability-motivated architecture choices (e.g., SoLU), we plan to address key structural bottlenecks that prevent the development of more principled design choices.
From the bottom-up, we develop unsupervised tools for reading out and intervening on the internals of existing models, relying on tractable structure encoded in causal feature relationships, activation geometry, and training dynamics rather than assuming sparsity as a primitive. These methods are more likely to yield practical intermediate results on applications such as sandbagging detection, data attribution, alignment faking, and the science of fine-tuning.
Bottom-Up Methods for Scale-Aware Feature Discovery
Our current bottom-up approaches fall under the category of Mechanistically Eliciting Latent Behaviors (MELBO), a class of methods which look for causally important directions in a model’s weight space as a more data-efficient alternative to sparsity-based feature discovery. This approach is conceptually similar to VPD, which likewise optimizes for causal importance under ablations. Our methods search for individual causally important directions, trading the completeness of the learned parameter decomposition for sample efficiency and scalability.
eNTK. This work connects closely with learning theory, and aims to convert the empirical neural tangent kernel (eNTK) into a tool capable of finding weight-space vectors that correlate with features in trained models. Eigendirections of the eNTK are data-space singular vectors of the Jacobian that can be mapped to weight space by applying its transpose. In recent work (presented at the workshop on mechanistic interpretability at ICML 2026), we found that eigendirections of the eNTK track ground-truth or independently defined feature directions in trained neural networks. Related to statistical physics, the scale-aware question here is: which spectral directions of the training dynamics correspond to behaviorally relevant features, and how do those directions emerge over training?
We tested this in two settings:
In a modular arithmetic setting (1L MLP and a 1L Transformer trained on modular addition).the top eNTK eigenspaces align with the Fourier features used by the known learned algorithms. We also found that alignment between the relevant eNTK subspaces and Fourier features evolves during training, with the first derivative of the alignment peaking near the onset of grokking. Since this work was published, we’ve also discovered that ablating the weight-space directions obtained by multiplying the eNTK eigenspace corresponding to Fourier sum modes with the Jacobian causes the model’s accuracy to drop dramatically compared to an equal-sized random ablation. This leads us to believe that the weight-space vectors associated with eNTK eigendirections are causally relevant.
In a pretrained language model setting (Gemma-3-270M), we found that eNTK eigendirections aligned more closely with independently-specified grammar feature directions than a same-budget baseline of PCA on model activations.
CPE. Causal Perturbative Elicitation (CPE) is an unsupervised method that learns rank-1 LoRAs by optimizing for causal importance, by a heuristic generalization of tensor decomposition algorithms. It can be viewed as a follow-up to original work on MELBO and Deep Causal Transcoding (DCT). CPE is essentially a higher-order version of the eNTK approach described above – instead of decomposing just the Jacobian, it implicitly decomposes a weighted sum of differential tensors that are higher order in weights. Taken with the eNTK work, we could ask: when are local, first-order directions enough, and when do we need finite-radius or higher-order perturbations to identify behaviorally meaningful structure?
Current experiments perform CPE on only a middle slice of a transformer, although in principle could be applied to the entire transformer. CPE is meant to back-chain from alignment-relevant problems where sample-efficient feature discovery is needed. With minimal tuning and tiny data-sets (below 64 prompts), CPE can:
recover 85% of password-locked performance on password-locked Llama3-70B.
virtually eliminate alignment-faking on Llama3-70B-based model organisms (increasing min compliance rate across free/paid users from 2% to 92%).
be used to initialize GPT-OSS-20B in an “aligned basin” in a model reward hacking environment, achieving similar oracle performance as training with oracle reward while training with the mis-specified reward.
Comparative Advantages and Near-term Work. How to choose one approach over the other comes down to understanding when the local, first-order directions of the eNTK are enough and when finite-radius, higher-order heuristics like CPE are needed to identify behaviorally meaningful features. The eNTK’s eigendecomposition is the SVD of the Jacobian, which is well-understood and easier to debug than NP-hard tensor decompositions. It can also be constructed from kernels used in mean-field theory, opening a window between interpretability applications and learning theory and making the eNTK a natural candidate for forward chaining from toy model analysis. Going forward, we would like to:
scale the eNTK eigenanalysis to larger models/datasets and compare their performance on feature discovery to SAEs.
compare with toy models of feature learning based on mean-field theory and saddle-to-saddle dynamics, where we can carefully compare the results to higher-order methods.
understand if the way that eNTK-related observables evolve throughout training could be used to build safety-relevant tools such as an early warning system for grokking or a gradient-based data attribution method.
On the other hand, saddle-to-saddle and SLT theories suggest higher-order information is needed for understanding structure in neural networks. CPE searches for meaningful perturbations at a finite non-zero radius around the current network, which may ultimately be more meaningful, and does not require forward-mode autodiff kernels, making it more readily scalable to larger models. Our current work learns rank-1 LoRAs on attention outputs across several layers, softly steering towards orthogonality over the flattened adapters without the need for excessive tuning. In the future, we want to be able to include higher than rank-1 perturbations for increased expressivity, which will require a more compositional measure of diversity than simple flattening. We are also currently studying whether decomposing the loss landscape directly (as opposed to using a purely unsupervised criterion) can aid science-of-fine-tuning-type alignment research. Can we decompose independent generalizations on a real LLM? We expect this to require both a mixture of high-quality model psychology and algorithmic research.
We plan to validate progress for bottom-up approaches by applying these techniques to alignment-relevant tasks. These can be broken into two broad categories:
Model Auditing: Problems like backdoor/sandbagging/deception detection, data attribution, or detection of “off-target fine-tuning effects” like emergent misalignment.
Model Reshaping: Studying whether perturbations in weight-space can reshape model behavior in a constructive way, particularly when standard techniques are “stuck”. For example, alignment faking is an example illustrating scenarios where standard RL may become stuck in sufficiently self-aware models due to exploration hacking, and was a successful application studied in our recent preprint. More ambitious extensions may study behavior reshaping which is more relevant to current-day practice, for instance trying to correct some of the weird generalization behaviors exhibited by gemini models, provided these problems can be reproduced faithfully enough in an open-weights setting.
The bet behind this research agenda is that we can make models sparse by fiat if we guess the right ansatz and are clever about implementation. There’s nothing special about dense models from an expressivity standpoint, as evidenced by prior work on weight-sparse transformers. If anything, myriad work supporting the circuit hypothesis suggest that dense models emulate sparse, at least in an approximate, platonic sense.
The only reason why dense dominates in practice is because it’s well-suited to standard hardware. Previous work estimates that it is around 100-1000x more expensive to train a weight-sparse transformer to the same level of performance as an equivalent dense model. Our recent pre-print is a first step to overcome this hurdle. The central primitive we introduce is a way to generate approximately-orthogonal hashed feature vectors. This trades memory access (slow on GPU) for compute (cheap and plentiful), which gives us the flexibility to score features and convert back to dense representations on the fly. In practice, this lets us scale to ~130k features per layer at only ~1.5x dense training throughput at the 1B parameter scale. Most of the remaining alignment tax shows up as data inefficiency (~5-10x slower than dense). The architecture only induces activation sparsity, but we demonstrate that this already yields more causally relevant directions than post-hoc SAEs. Our goal in the next 6 months is to use some of the same computational primitives (and perhaps new ones) to make the architecture more weight-sparse while staying below the current alignment tax of ~10x dense training costs.
We acknowledge that this problem is hard. In order to make progress, we must be honest about the key underlying bottlenecks, and then try to reason from first principles how to address them one-by-one. We present some of these below to be expanded in an upcoming set of open problems:
Overcoming the alu:mem gap. A fully sparse model must account for the time cost of moving a float from HBM to on-chip, which is around 100-1000x more than performing a FLOP with that same float. Conditionally loading only active weights of a sparse model wins compared to dense only if the active set is truly tiny.
Comp-in-sup feature counting. Various results in the computation in superposition literature suggest that a ~square dense MLP taking inputs in $\mathbb R^{d_\textrm{model}}$ can perform computations on around $O(d_\textrm{model}^2)$ many sparsely-activating features operating in superposition (up to logarithmic factors in the denominator). This suggests that enforcing sparsity in standard basis for the relatively “narrow" values of $d_\textrm{model}$ used in practice (between $1024$ to $4096$) is not enough - the model will likely still operate in a superposition regime to emulate finer-grained features.
Ensuring there are “Enough collisions at init”. In a very wide, sparsely activating network, feature pairs rarely co-fire, burying interaction signal below SGD's noise floor. A natural fix for this is to structure computation hierarchically, starting with a small number of more frequent, coarse-grained features that learn non-trivial interactions, and adding capacity with finer-grained features (from commonly co-occurring subsets of coarse-grained features) as needed.
Maintaining Interpretability in Attention. Softmax attention can mix tokens densely, even over a sparse residual stream, particularly early in training before attention patterns have sharpened. A fully sparse transformer needs additional filtering or sparsity constraints on attention. Though challenging, this problem overlaps with standard capabilities research, with a wealth of strategies to borrow.
Data Models and Validation Methods
For the past several months, our focus has been exploring a data model consisting of hierarchical functions defined on critical mean-field percolation clusters embedded in a high-dimensional data space. The resulting data distribution comprises sparse, low-dimensional fractal clusters with a power-law distribution of cluster sizes. Latent variables modeling a taxonomic hierarchy generate each data point's target value. The data model is analytically tractable with known critical exponents that fix its scaling properties without requiring hyperparameter tuning.
Since our last progress update, we improved the algorithm used to generate the data’s latent hierarchical structure. The previous code created undirected treelike graphs as an approximation of high-dimensional percolation clusters by growing them using a preferential attachment process. Our most recent paper, presented at the 2026 Mechanistic Interpretability Workshop at ICML 2026, replaced this with an exact procedure that leverages a mapping between percolation clusters, random trees, and additive coalescence. The code to generate synthetic datasets based on this model now implements an almost linear-time algorithm to jointly sample a random tree and its hierarchical latent decomposition, efficiently generating graphs with the precise distribution of high-dimensional percolation clusters and enabling data generation at arbitrary scale.
Caption: The percolation data model. (a) Inputs are distributed as self-similar fractal clusters with power-law sizes. (b) Targets are generated by hierarchical latent variables decomposing each cluster.
One important question to ask is: how does a neural network keep track of the ground truth latent hierarchy generated by a percolation dataset? Are these linearly represented within the network’s internal activations? We trained a residual MLP on our synthetic dataset and trained linear probes to regress the dataset’s latent values, grouped by depth in the latent tree. We found that the ground-truth variables are more linearly accessible in the activations of the MLP compared to the raw input, with the performance gap shrinking for larger subtrees. This outcome is consistent with the hypothesis that coarser structure is more accessible from the input geometry. The code for training neural networks on percolation data is available here.
Where this is going. Building on the initial probing experiments, we plan to rigorously verify other hypotheses about neural representations using causal intervention methods. To test hypotheses about `interpretable’ features, we will also compare our latent features with those recovered by sparse autoencoders. We also intend to scale up the data generation code to produce larger datasets with more data points per latent, and to run experiments varying the width, depth, and initialization of networks trained on this data. In addition, we will release public datasets to make the percolation framework accessible to the broader mechanistic interpretability community as a tractable, synthetic sandbox for interpretability experiments.
We are committed to creating better validation methods of empirical or theoretical feature hypotheses, which we think is essential for enabling ambitious interpretability of advanced AI systems. In the next year, we aim to produce a competitive benchmark to assess the performance of interpretability tools. Our goal is to use the properties of tractable but realistic synthetic datasets to design evaluation metrics with more robust faithfulness guarantees than the state of the art, enabling the reliable assessment of a tool’s ability to interpret model internals. Four questions guide future work:
What structural properties of data, instantiated in synthetic datasets, replicate the behavior of deep neural networks trained on natural learning tasks? Example experiments include:
Measuring transfer learning on natural data from pretraining on synthetic data. What structural properties and hyperparameters improve performance?
Studying synthetic models of representational alignment by separately varying the random seeds for the latent variables and embedded graphs in the percolation model.
Do natural datasets have hierarchical structure? If so, how should we measure, model, and interpret it? Example projects include:
Designing an autoencoder to fit a self-similar fractal distribution and training it on natural data.
Looking for evidence of hierarchical structure in data by applying hierarchical clustering methods to sparse autoencoder features.
How does data structure shape the concepts used by intelligent systems? Example directions include:
Defining quantitative metrics to describe how well a neural network reconstructs the latent forest in the percolation model and applying them to evaluate the features reported by interpretability tools.
Understanding how the hierarchical latent variables described by the percolation model relate to existing work on concepts, including natural latents and condensation.
What observable properties of datasets and trained networks differentiate data models? Examples include:
Comparing the percolation kernel spectrum to the power-law spectra observed in real data.
Investigating the scaling laws resulting from a data distribution consisting of a fractal cluster or a set of data manifolds with a power-law size distribution.
Applying the skewed latent hierarchy described by the percolation model to investigate typicality and asymmetrical similarity in learned representations.
Coming Soon
We'll be sharing more about PIRAMID's work, as well as PIAMI (Physics-Informed Ambitious Mech Interp) -- a research program coordinated by PrincInt -- soon.
In a recent post, we presented PIRAMID, its leadership and research pillars, and a plan for how they fit together. In this post, we sketch a team-by-team account of progress and targets for the next 6–12 months. We include results to date as evidence of viability: we’ve been a small team, with much of the past year spent building behind the scenes, and we aim to greatly accelerate our progress over the coming year as we expand our efforts and our teams. Like any fundamental scientific effort, none of this is set in stone. We expect some of these bets to need revision and are confident in our ability to reassess and change course as new evidence comes to light.
If you’re interested in collaborating or supporting our work as we expand, please get in touch.
Advancements in Learning Theory
We aim to formulate statistical and mesoscopic theories of feature learning and generalization which provide a level of description between microscopic parameter-level dynamics and macroscopic performance metrics. Our work so far has treated statistical physics (e.g., mean field methods) as a candidate language for statistically describing learned structure, and covariance measures between neurons or weights as candidate mesoscopic objects that track meaningful scales and determine feature relevance during training.
We’ve developed and tested the expressive power of this language in three directions. The first lays the foundation for mean-field theory’s potential and utility. Our most recent pre-print argues that Bayesian mean-field theory is complexity complete, in the sense that it captures the aggregate heuristic bias of finite-width neural networks. In parallel, we have also been distilling mean-field theory for a broader audience, in the hopes that more interpretability practitioners will use and improve this framework. A detailed post on mean field theory and computation in superposition connects field-theoretic descriptions of neural networks with sparse and superposed structure, which are central to empirical interpretability but not yet well integrated into a field-theoretic approach to learning theory.
Second, we are developing a mean-field model of feature learning with nontrivial hidden-layer statistics. Most existing theory, including dynamical mean-field work that tracks evolving kernels and finite-width fluctuations (e.g., Bordelon & Pehlevan, Rubin et al.), treats feature learning through the predictor or through kernels and order parameters that summarize input–output behavior. Instead, our work incorporates hidden-layer structure (characterized by various scales of interaction like non-trivial neuron-neuron covariance, or how localized or distributed a representation is) in the minimal state description. A theory sensitive to the internal structures formed during training could explain why some models learn robust, reusable features while others remain brittle, and could connect representation geometry to monitoring, transfer, and post-training stability. In developing meaningful observables (e.g., similar to susceptibilities), such theories can help us develop tools to more faithfully interpret model internals and guide our understanding of which theoretical limits or toy models hold the most weight in a given setting.
A third line of work aims to develop a theory of error geometry, inspired by past work from Mei and Montanari (e.g., this) and Canatar et al.. Recent progress (to appear soon) shows how correlations between training and test points affect which related examples are learned or missed together in a random-feature model. This project asks a pointwise question that most generalization theory averages away: when a test example is correlated with a specific subset of the training set, how does that correlation structure determine its error? For modern web-scale models, the practically relevant regime is not the idealized setting in which every test point is equally novel. Many important failures occur near pockets of the training distribution, or along directions that are only partially novel. A pointwise theory of error geometry would therefore sharpen more than standard generalization theory: it would give a language for when models interpolate safely, memorize dangerously, or fail on targeted shifts.
Together, these projects move toward a common pipeline: prove statements in idealized stochastic models, identify the induced statistics of features or errors, and test those statistics in realistic networks.
Where this is going. This work depends on the use of toy models and model organisms that act as simplified settings — simple feature-learning models, mean-field models with nontrivial hidden-layer statistics, superposition-like toy models, generalization properties of mean field models, and linear or quadratic models of training dynamics — which can isolate a qualitative or quantitative component of interpretability’s (currently large) theory-practice gap. Our approach is to increase the breadth of our theoretical understanding and its applicability to realistic models by coming up with ‘structural gates’ an idealized model must pass through (incorporate) to improve their utility while avoiding unnecessary complexity. For example, many of our current projects use existing theory at infinite width as an intuition pump with the aim of adding increased complexity in the form of these gates. Designing the gates – criteria for a flavor of difficulty that current theories or tools cannot handle – is a core part of this research and the subject of a post to appear soon.
Our strategy for future work operates along two complementary paths:
The practical goal is to turn these theories into measurements that are useful for the interpretability and safety of real-world models, reflecting in the following targets:
Target 1: Error geometry in the feature learning regime. The theory described above predicts, a priori, the pattern of correlated successes and failures for a trained random-feature model. We extend this in two directions: making it dynamical (how does the error on a partially correlated test point evolve throughout training, and are there distinct time scales for memorization-like versus generalization-like behavior?) and moving beyond frozen features (in regimes where the effective kernel itself evolves during training, which parts of the random-feature geometry survive and which are replaced by genuine feature-learning effects?) Empirical tests should track, for each held-out point, its similarity profile to the training set, its layer-wise representation trajectory, and its instantaneous prediction error across training time.
Target 2: A better measure of features. These are objects that exist both in data/function space and in activation/parameter space, and which can include distributed, superposed, or circuit-like structure rather than only sparse activation directions. Modeling them requires both an understanding at both a functional level (kernel correction) and a representational level (in terms of neuron marginals and covariances, symmetry breaking and specialization, and lottery ticket-like sub-circuit formation). Defining a zoology of ‘features’, and understanding how and when different forms occur, bears directly on understanding when interpretability is possible at all and which methods can work in a given regime. We think that a good theory of features – defined here as atoms of a representation or the model’s memory/ storage at a given layer – should also couple to a good notion of a circuit (i.e., a decomposition of the model computation into atoms). An advantage of mean field theory techniques is that they inherently have the expressive power to link representational and computational techniques (as in work by Rubin et al. 2023 and Rubin et al. 2025). This suggests a possible interpretability tool that searches for feature maps and sparse “connectomes” between them: low-dimensional functions of layer activations that recover the intermediate structure of a learned computation. A first testbed for this agenda is circuit reconstruction. Given a network trained to implement a known Boolean circuit, modular arithmetic algorithm, or other controlled computation, can we recover the latent gates, feature maps, and parent-child relationships between them from the trained model? This provides a concrete model organism for asking whether mean-field-style summaries can recover mechanistic structure without relying entirely on neuron-level decompositions or SAE-style sparsity assumptions.
Target 3: Signals of Feature Learning. Lastly, we will explore methods to track how features form, change, and interact during training\footnote{This is similar in its aims to the developmental interpretability community.}. What are the various time scales of feature formation, and to what extent is feature learning stationary? Perhaps the most safety-relevant target here, related to silent alignment, is the gap between a feature’s acquisition and its behavioral expression. The practical goals are theory-inspired early-detection measures for such signals, and a taxonomy of classes of sudden change in learning dynamics paired with the appropriate detection techniques.
Interpretability Applications
What does it mean for interpretability to be ambitious? Aiming for a complete understanding of neural network components – in our case, measured by physics-informed faithfulness criteria – is an ambitious target in its own right. However, even a partial understanding can be a proxy target that is ambitious in its application, provided it moves the needle for robust, scalable alignment and is held to the same faithfulness standard. This second class presupposes that models decompose into multiple independent generalizations, organized into basins (in the sense of the loss landscape geometry) that are ‘aligned’ by some suitably rigorous metric. If this is true, interpretability becomes a means of constructively shaping model behavior – steering the model toward more aligned basins during training, e.g., early in RL – rather than only an auditing tool. These training-time interventions are especially relevant for preventing nefarious phenomena like exploration hacking or deceptive alignment, for which post-hoc detection is insufficient.
These targets (both direct and proxy) shape our strategy in two ways:
Bottom-Up Methods for Scale-Aware Feature Discovery
Our current bottom-up approaches fall under the category of Mechanistically Eliciting Latent Behaviors (MELBO), a class of methods which look for causally important directions in a model’s weight space as a more data-efficient alternative to sparsity-based feature discovery. This approach is conceptually similar to VPD, which likewise optimizes for causal importance under ablations. Our methods search for individual causally important directions, trading the completeness of the learned parameter decomposition for sample efficiency and scalability.
eNTK. This work connects closely with learning theory, and aims to convert the empirical neural tangent kernel (eNTK) into a tool capable of finding weight-space vectors that correlate with features in trained models. Eigendirections of the eNTK are data-space singular vectors of the Jacobian that can be mapped to weight space by applying its transpose. In recent work (presented at the workshop on mechanistic interpretability at ICML 2026), we found that eigendirections of the eNTK track ground-truth or independently defined feature directions in trained neural networks. Related to statistical physics, the scale-aware question here is: which spectral directions of the training dynamics correspond to behaviorally relevant features, and how do those directions emerge over training?
We tested this in two settings:
CPE. Causal Perturbative Elicitation (CPE) is an unsupervised method that learns rank-1 LoRAs by optimizing for causal importance, by a heuristic generalization of tensor decomposition algorithms. It can be viewed as a follow-up to original work on MELBO and Deep Causal Transcoding (DCT). CPE is essentially a higher-order version of the eNTK approach described above – instead of decomposing just the Jacobian, it implicitly decomposes a weighted sum of differential tensors that are higher order in weights. Taken with the eNTK work, we could ask: when are local, first-order directions enough, and when do we need finite-radius or higher-order perturbations to identify behaviorally meaningful structure?
Current experiments perform CPE on only a middle slice of a transformer, although in principle could be applied to the entire transformer. CPE is meant to back-chain from alignment-relevant problems where sample-efficient feature discovery is needed. With minimal tuning and tiny data-sets (below 64 prompts), CPE can:
Comparative Advantages and Near-term Work. How to choose one approach over the other comes down to understanding when the local, first-order directions of the eNTK are enough and when finite-radius, higher-order heuristics like CPE are needed to identify behaviorally meaningful features. The eNTK’s eigendecomposition is the SVD of the Jacobian, which is well-understood and easier to debug than NP-hard tensor decompositions. It can also be constructed from kernels used in mean-field theory, opening a window between interpretability applications and learning theory and making the eNTK a natural candidate for forward chaining from toy model analysis. Going forward, we would like to:
On the other hand, saddle-to-saddle and SLT theories suggest higher-order information is needed for understanding structure in neural networks. CPE searches for meaningful perturbations at a finite non-zero radius around the current network, which may ultimately be more meaningful, and does not require forward-mode autodiff kernels, making it more readily scalable to larger models. Our current work learns rank-1 LoRAs on attention outputs across several layers, softly steering towards orthogonality over the flattened adapters without the need for excessive tuning. In the future, we want to be able to include higher than rank-1 perturbations for increased expressivity, which will require a more compositional measure of diversity than simple flattening. We are also currently studying whether decomposing the loss landscape directly (as opposed to using a purely unsupervised criterion) can aid science-of-fine-tuning-type alignment research. Can we decompose independent generalizations on a real LLM? We expect this to require both a mixture of high-quality model psychology and algorithmic research.
We plan to validate progress for bottom-up approaches by applying these techniques to alignment-relevant tasks. These can be broken into two broad categories:
Top-Down Hierarchical Architectures: Scaling Sparse Transformers
The bet behind this research agenda is that we can make models sparse by fiat if we guess the right ansatz and are clever about implementation. There’s nothing special about dense models from an expressivity standpoint, as evidenced by prior work on weight-sparse transformers. If anything, myriad work supporting the circuit hypothesis suggest that dense models emulate sparse, at least in an approximate, platonic sense.
The only reason why dense dominates in practice is because it’s well-suited to standard hardware. Previous work estimates that it is around 100-1000x more expensive to train a weight-sparse transformer to the same level of performance as an equivalent dense model. Our recent pre-print is a first step to overcome this hurdle. The central primitive we introduce is a way to generate approximately-orthogonal hashed feature vectors. This trades memory access (slow on GPU) for compute (cheap and plentiful), which gives us the flexibility to score features and convert back to dense representations on the fly. In practice, this lets us scale to ~130k features per layer at only ~1.5x dense training throughput at the 1B parameter scale. Most of the remaining alignment tax shows up as data inefficiency (~5-10x slower than dense). The architecture only induces activation sparsity, but we demonstrate that this already yields more causally relevant directions than post-hoc SAEs. Our goal in the next 6 months is to use some of the same computational primitives (and perhaps new ones) to make the architecture more weight-sparse while staying below the current alignment tax of ~10x dense training costs.
We acknowledge that this problem is hard. In order to make progress, we must be honest about the key underlying bottlenecks, and then try to reason from first principles how to address them one-by-one. We present some of these below to be expanded in an upcoming set of open problems:
Data Models and Validation Methods
For the past several months, our focus has been exploring a data model consisting of hierarchical functions defined on critical mean-field percolation clusters embedded in a high-dimensional data space. The resulting data distribution comprises sparse, low-dimensional fractal clusters with a power-law distribution of cluster sizes. Latent variables modeling a taxonomic hierarchy generate each data point's target value. The data model is analytically tractable with known critical exponents that fix its scaling properties without requiring hyperparameter tuning.
Since our last progress update, we improved the algorithm used to generate the data’s latent hierarchical structure. The previous code created undirected treelike graphs as an approximation of high-dimensional percolation clusters by growing them using a preferential attachment process. Our most recent paper, presented at the 2026 Mechanistic Interpretability Workshop at ICML 2026, replaced this with an exact procedure that leverages a mapping between percolation clusters, random trees, and additive coalescence. The code to generate synthetic datasets based on this model now implements an almost linear-time algorithm to jointly sample a random tree and its hierarchical latent decomposition, efficiently generating graphs with the precise distribution of high-dimensional percolation clusters and enabling data generation at arbitrary scale.
Caption: The percolation data model. (a) Inputs are distributed as self-similar fractal clusters with power-law sizes. (b) Targets are generated by hierarchical latent variables decomposing each cluster.
One important question to ask is: how does a neural network keep track of the ground truth latent hierarchy generated by a percolation dataset? Are these linearly represented within the network’s internal activations? We trained a residual MLP on our synthetic dataset and trained linear probes to regress the dataset’s latent values, grouped by depth in the latent tree. We found that the ground-truth variables are more linearly accessible in the activations of the MLP compared to the raw input, with the performance gap shrinking for larger subtrees. This outcome is consistent with the hypothesis that coarser structure is more accessible from the input geometry. The code for training neural networks on percolation data is available here.
Where this is going. Building on the initial probing experiments, we plan to rigorously verify other hypotheses about neural representations using causal intervention methods. To test hypotheses about `interpretable’ features, we will also compare our latent features with those recovered by sparse autoencoders. We also intend to scale up the data generation code to produce larger datasets with more data points per latent, and to run experiments varying the width, depth, and initialization of networks trained on this data. In addition, we will release public datasets to make the percolation framework accessible to the broader mechanistic interpretability community as a tractable, synthetic sandbox for interpretability experiments.
We are committed to creating better validation methods of empirical or theoretical feature hypotheses, which we think is essential for enabling ambitious interpretability of advanced AI systems. In the next year, we aim to produce a competitive benchmark to assess the performance of interpretability tools. Our goal is to use the properties of tractable but realistic synthetic datasets to design evaluation metrics with more robust faithfulness guarantees than the state of the art, enabling the reliable assessment of a tool’s ability to interpret model internals. Four questions guide future work:
Coming Soon
We'll be sharing more about PIRAMID's work, as well as PIAMI (Physics-Informed Ambitious Mech Interp) -- a research program coordinated by PrincInt -- soon.