Reading a Language Model's Thoughts Before It Speaks - Dual Edge: Read/Write
The claim, up front Working with Anthropic's Jacobian lens on a couple of chat models, I found that a model's committed next-token decision is often legible in its residual stream before the model emits anything and that you can use this to catch a prompt-injection roughly one token early. The...
Jul 91