I'm often asked about the differences and similarities between Simplex' and Timaeus' research agendas. The question is natural enough. Both focus on a 'fundamental science' approach to AI alignment. Both organizations base their research agendas on sophisticated mathematical frameworks handed down from a bearded ur-figure (Sumio Watanabe, James Crutchfield).
We may posit the following correspondence
Dan Murfet + Jesse Hoogland : Developmental Interpretability : Singular Learning Theory : Sumio Watanabe
<->
Adam Shai + Paul Riechers : Belief-state Geometry: Computational Mechanics : James Crutchfield
SLT vs CompMech
Round One. Fight!
Weights vs Activations
DevInterp & SLT is about weight space. Belief-state geometry is more about studying activation space.
Activation space is what is already being studied in MechInterp & most 'mainstream' approaches to interpretability. It is concrete and present to the senses. Weight space is much larger, more abstract, harder to measure and sample.
Training vs Inference
SLT is about training. CompMech is about inference.
Both study Bayesian posteriors and updating. For SLT that is the Bayesian posterior on weight space - hence relevant for training. The Belief-State Geometry agenda studies the Mixed State Presentation from CompMech which describes an [idealized] version of in-context learning as the LLM doing Bayesian updating token-by-token as it reads the context.
Caveat. It isn't clear that inference and learning are really fundamentally distinct. It's all just updating/conditioning in Bayesian statistics. Indeed, if one buys the Strong in-Context Learning story the difference between the training/learning of LLMs and inference may be somewhat illusory.
IID vs non-IID Data
Singular Learning Theory has historically been about IID data. CompMech about IID data is trivial.
Caveat. Singular learning theory has been studied for non-IID data but it's fair to say its development is in its early stages.
Parameterization vs Invariant structure
Computational Mechanics is all about the invariant structure. A fundamental objects is the epsilon machine (and its mixed-state presentation) which is obtained by quotienting out all the details of the past (read: the context) that do not matter for predictions of the future (read: continuation)
Singular learning theory would be trivial if not for the weight-parameterization of the loss function. The important bit is that the loss landscape of loss functions on neural-network architectures are singular. That is a feature of the specific parameterization of the underlying function/ probability distribution by the neural network architecture.
Asymptotic vs Exact
Singular Learning Theory's fundamental theorem is Watanabe's Free Energy formula which describes the Bayesian posterior in the large data limit. Computational Mechanics can equally be applied to large as well as small data regimes.
Mechanism vs Behaviourial
Belief-State Geometry posits a model for the underlying internal mechanism of LLMs. The devinterp agenda flags when important structure forms. SLT flags structure forming. CompMech has a mechanistic story of what that structure is.
Bottom-up vs Top-down interpretability
Computational mechanics has a ground-level truth of the data process. By contrast, outside of a few simple examples (Deep Linear Networks, shallow ReLU neural networks) the central quantity of SLT the famed learning coefficient is not known analytically. The learning coefficient is estimated but we know this estimator can be very wrong, both empirically and on first principles.[1]
The belief-state geometry agenda has historically focused on understanding mechanistically small toy models. Perhaps its chief challenge is porting that understanding to large models for which we know the language of HMMs must be inadequate. Some relevant work here: Simplex paper on factored representations. Crutchfield's visionary[2] paper on calculi of emergence.
By contrast, Timaeus moved to large scale experiments quickly. Rather than understanding mechanistically every single detail the point is to understand what features are truly important (as flagged by their learning-theoretic behaviour).
Physics vs Math?
This one is less clear. Daniel Murfet and Sumio Watanabe are mathematicians while Paul Riechers and James Crutchfield are physicists.
However, SLT has close connections with statistical mechanics and Daniel Murfet has an extensive background and track record in mathematical physics. And computational mechanics is unusual within the larger space topics which it overlaps with (quantum physics, complex systems, chaos theory, etc) for its emphasis on precise mathematical definitions and careful conceptual grounding.
Conclusion
I pronounce: Simplex and Timaeus are neither competitors nor natural allies. Their historical domains of study are largely orthogonal. That may change as they start to invade each other's stomping ground.
On the whole, I regard this as a very natural and healthy state. A time-tested way to make progress on very hard problems is breaking them up into separate, mostly orthogonal phenomena and study each phenomena carefully and systematically before. Inshallah, this ultimately paves the way for a paradigm-shifting synergy.
I'm often asked about the differences and similarities between Simplex' and Timaeus' research agendas. The question is natural enough. Both focus on a 'fundamental science' approach to AI alignment. Both organizations base their research agendas on sophisticated mathematical frameworks handed down from a bearded ur-figure (Sumio Watanabe, James Crutchfield).
We may posit the following correspondence
Dan Murfet + Jesse Hoogland : Developmental Interpretability : Singular Learning Theory : Sumio Watanabe
<->
Adam Shai + Paul Riechers : Belief-state Geometry: Computational Mechanics : James Crutchfield
SLT vs CompMech
Round One. Fight!
Weights vs Activations
DevInterp & SLT is about weight space. Belief-state geometry is more about studying activation space.
Activation space is what is already being studied in MechInterp & most 'mainstream' approaches to interpretability. It is concrete and present to the senses. Weight space is much larger, more abstract, harder to measure and sample.
Training vs Inference
SLT is about training. CompMech is about inference.
Both study Bayesian posteriors and updating. For SLT that is the Bayesian posterior on weight space - hence relevant for training. The Belief-State Geometry agenda studies the Mixed State Presentation from CompMech which describes an [idealized] version of in-context learning as the LLM doing Bayesian updating token-by-token as it reads the context.
Caveat. It isn't clear that inference and learning are really fundamentally distinct. It's all just updating/conditioning in Bayesian statistics. Indeed, if one buys the Strong in-Context Learning story the difference between the training/learning of LLMs and inference may be somewhat illusory.
IID vs non-IID Data
Singular Learning Theory has historically been about IID data. CompMech about IID data is trivial.
Caveat. Singular learning theory has been studied for non-IID data but it's fair to say its development is in its early stages.
Parameterization vs Invariant structure
Computational Mechanics is all about the invariant structure. A fundamental objects is the epsilon machine (and its mixed-state presentation) which is obtained by quotienting out all the details of the past (read: the context) that do not matter for predictions of the future (read: continuation)
Singular learning theory would be trivial if not for the weight-parameterization of the loss function. The important bit is that the loss landscape of loss functions on neural-network architectures are singular. That is a feature of the specific parameterization of the underlying function/ probability distribution by the neural network architecture.
Asymptotic vs Exact
Singular Learning Theory's fundamental theorem is Watanabe's Free Energy formula which describes the Bayesian posterior in the large data limit. Computational Mechanics can equally be applied to large as well as small data regimes.
Mechanism vs Behaviourial
Belief-State Geometry posits a model for the underlying internal mechanism of LLMs. The devinterp agenda flags when important structure forms. SLT flags structure forming. CompMech has a mechanistic story of what that structure is.
Bottom-up vs Top-down interpretability
Computational mechanics has a ground-level truth of the data process. By contrast, outside of a few simple examples (Deep Linear Networks, shallow ReLU neural networks) the central quantity of SLT the famed learning coefficient is not known analytically. The learning coefficient is estimated but we know this estimator can be very wrong, both empirically and on first principles.[1]
The belief-state geometry agenda has historically focused on understanding mechanistically small toy models. Perhaps its chief challenge is porting that understanding to large models for which we know the language of HMMs must be inadequate. Some relevant work here: Simplex paper on factored representations. Crutchfield's visionary[2] paper on calculi of emergence.
By contrast, Timaeus moved to large scale experiments quickly. Rather than understanding mechanistically every single detail the point is to understand what features are truly important (as flagged by their learning-theoretic behaviour).
Physics vs Math?
This one is less clear. Daniel Murfet and Sumio Watanabe are mathematicians while Paul Riechers and James Crutchfield are physicists.
However, SLT has close connections with statistical mechanics and Daniel Murfet has an extensive background and track record in mathematical physics. And computational mechanics is unusual within the larger space topics which it overlaps with (quantum physics, complex systems, chaos theory, etc) for its emphasis on precise mathematical definitions and careful conceptual grounding.
Conclusion
I pronounce: Simplex and Timaeus are neither competitors nor natural allies. Their historical domains of study are largely orthogonal. That may change as they start to invade each other's stomping ground.
On the whole, I regard this as a very natural and healthy state. A time-tested way to make progress on very hard problems is breaking them up into separate, mostly orthogonal phenomena and study each phenomena carefully and systematically before. Inshallah, this ultimately paves the way for a paradigm-shifting synergy.
cf Dmitry Vaintrob's Tanh network example.
but not always completely rigourous