Can deceptive behavior in a LLM be identified and causally manipulated i.e. can we catch when AI Lies?
Abstract I investigated if the deceptive behavior in LLM is identifiable in its internal representations and whether it can be causally manipulated. Using Gemma-2-2B, I found: 1. Truthful and deceptive prompts produced separable residual-stream states. 2. Linear probe evaluation achieved 100% accuracy. * Question-level held-out evaluation achieved 100% accuracy. 3....
Aug 231