Making an LLM interpretability playground
After many months of reading on the interpretability experiments/blogposts by Anthropic, I noticed that despite being more or less capable of understanding the basic ideas expressed in them, I had zero practical experience and many gaps in my mental model of how sparse autoencoders (or interpretability in general) work. In...
Sep 11