Zero- and mean-ablation disagree about which head drives this SAE feature (Gemma-2-2B induction)
Epistemic status: one model, one layer, three SAE training seeds. The qualitative finding replicates across all three seeds; the effect sizes are from one run and shouldn't be load-bearing without more replication. Code, weights and seeds linked below. Done on a consumer GPU in my spare time. First LW post....
May 201