(This is a demo of the library's capabilities so the sample size is tiny... much more analysis is done in an upcoming paper, for instance we mine hundreds of forks and show that ablating head 5 costs 2.78 logits whereas every other head in the layer costs ≤0.14)
Interactive app demo challenge:
The quickest way to run and reproduce the image state is:
#then run the app at 23m, set Elo to 2400, and input FEN: 4kb1r/p2n1ppp/4q3/4p1B1/4P3/1Q6/PPP2PPP/2KR4 w k - 0 16 chessformer_lens 23m
Try to use move microscope (bottom middle window) and ablate this head (top right button) to determine which head is most causally linked to carrying the stunning queen sacrifice. Bonus points if you can name this legendary game!
This repo's core is one engine with three frontends: engine.py is the interp core (model + hooks + logit lens + head ablation + GAB decomposition + logit/policy across depth). It does all of the analysis for three frontends: a native interactive app (see above), a module to plot mechanistic analyses such as the residual stream across depth, and interactive widgets for notebooks for playful experiments for hypothesis generation. chessformer_lens is built especially for chess transformers that map one token to one square—allowing for beautiful board readable attention patterns.
There is wonderful prior chess-interp work, for instance: McGrath et al. on AlphaZero concepts, Jenner et al. on lookahead in Leela, Karvonen on chess-GPT—but there is no shared general infrastructure for rigorous interpretability work with interactive visualization of chessformer models.
chessformer_lens is built for this, especially inspired by Neel Nanda's transformer_lens library.
The Maia-3 model interpreter is completed; other tokenization schemes will be tackled.
Please don't hesitate to give me feedback or thoughts by email or at davidlitman.com. I hope for this to be a useful and intuitive tool for the community!
Quick interp demo in colab:
Localize knight forks to a single head in Maia-3 with logit-lens and per-head ablation.
https://colab.research.google.com/drive/1YYZBd_SZbjOscRXIqJUbfCaY7rRbEzWx?usp=sharing
(This is a demo of the library's capabilities so the sample size is tiny... much more analysis is done in an upcoming paper, for instance we mine hundreds of forks and show that ablating head 5 costs 2.78 logits whereas every other head in the layer costs ≤0.14)
Interactive app demo challenge:
The quickest way to run and reproduce the image state is:
python3 -m venv .venv && source .venv/bin/activate
pip install git+https://github.com/CSSLab/maia3 #Maia-3 not pip installable yet
pip install "chessformer_lens[all]"
#then run the app at 23m, set Elo to 2400, and input FEN: 4kb1r/p2n1ppp/4q3/4p1B1/4P3/1Q6/PPP2PPP/2KR4 w k - 0 16
chessformer_lens 23m
Try to use move microscope (bottom middle window) and ablate this head (top right button) to determine which head is most causally linked to carrying the stunning queen sacrifice. Bonus points if you can name this legendary game!
--------------------------------------------------------------------------------------------------------
The chessformer_lens library
The github repo is https://github.com/chessformer-lens/chessformer_lens, and it is pip installable.
This repo's core is one engine with three frontends:
engine.pyis the interp core (model + hooks + logit lens + head ablation + GAB decomposition + logit/policy across depth). It does all of the analysis for three frontends: a native interactive app (see above), a module to plot mechanistic analyses such as the residual stream across depth, and interactive widgets for notebooks for playful experiments for hypothesis generation. chessformer_lens is built especially for chess transformers that map one token to one square—allowing for beautiful board readable attention patterns.There is wonderful prior chess-interp work, for instance: McGrath et al. on AlphaZero concepts, Jenner et al. on lookahead in Leela, Karvonen on chess-GPT—but there is no shared general infrastructure for rigorous interpretability work with interactive visualization of chessformer models.
chessformer_lens is built for this, especially inspired by Neel Nanda's transformer_lens library.
The Maia-3 model interpreter is completed; other tokenization schemes will be tackled.
Please don't hesitate to give me feedback or thoughts by email or at davidlitman.com. I hope for this to be a useful and intuitive tool for the community!