Stefan Heimersheim and I recently introduced a measurement method that appears to distinguish feature directions from non-feature directions in LLMs, see here for details. tl;dr, We perturb activations into candidate feature directions and fit an
We applied our method to J-lens directions from Anthropic's recent workspace paper in Qwen3.6-27B (for which many of the key results have been replicated, see here for details) and found
Thanks to Neel Nanda for suggesting this experiment.
Experimental details:

Code for reproduction: https://github.com/FranciscoHS/fsec-paper/tree/main/jlens