When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles
TL;DR Activation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation analysis becomes a conversation with the AO. If someone wanted to audit a model that hides something, like a backdoor or a...
Sep 19