Natural Language Transcoders
Describing the computation performed in a stack of transformer layers Anwen Hao, mentored by Adrians Skapars Anthropic’s natural language autoencoders (NLAs) is a promising method to automatically generate explanations of activations. But what if we want to explain the computation that occurs over a stack of layers? To address this...
Aug 1911