A missing lecture in mechanistic interpretability: Feature Attribution and LRP
ML interpretability research has a funny divide. Mechanistic interpretability is the name of a field originated largely by non-traditional researchers, ranging from industry researchers at Anthropic to independent BlueDot-grant researchers to hackers working on fun projects in their free time on Discord. Meanwhile, it is not hard to find the...
Sep 287