Principles of Intelligence (PrincInt, formerly PIBBSS) is launching PIRAMID, an internal research division using the tools and techniques of statistical physics to build scientific foundations for ambitious mechanistic interpretability. PIRAMID’s central premise is that scalable alignment will require more than persuasive ad-hoc explanations of model behavior. It will require interpretability tools that develop alongside a scientific understanding of the structure of data, learning, and representations.
To reflect this, we divide our attention across three synergistic research teams: Advancements in Learning Theory (led by Dmitry Vaintrob), Interpretability Applications (led by Andrew Mack), and Data Models and Validation Methods (led by Ari Brill). Together, they form a loop: theory predicts how structure can be learned and organized in networks, interpretability tools built on these principles help us recover and intervene on that structure, and synthetic datasets with built-in ground truth provide settings in which both theory and tools can be validated. We can think of this as loosely mirroring physics' methodological division of labor, with each group prioritizing theory, empirics, and phenomenology, respectively. This methodological coverage helps to build up a scientific understanding of real-world neural networks that narrows the theory-practice gap.
PIRAMID is part of PrincInt’s larger field-building efforts. Over the past year and a half, we hired a cohort of affiliate researchers to test candidate directions (several of whom went on to form the PIRAMID leadership team), started a series of workshops connecting statistical physics with AI interpretability, and began incubating new academic research groups as part of PIAMI (Physics-Informed Ambitious Mechanistic Interpretability), a coordinated research program of which PIRAMID is one working group. PIAMI is how we plan to stay connected to the communities of expertise this work draws on -- across physi