x
Can deceptive behavior in a LLM be identified and causally manipulated i.e. can we catch when AI Lies? — LessWrong