small targeted training as a mechanistic “amplifier” for discovering latent/learned deceptive circuitry
what if llm deception is implemented as a conditional computational circuit? a possible way to study this, RL-train a model on environments where deception is instrumentally rewarded, honest → +10 successful deception → +50 caught → −30 the interesting part is that we may not need much deceptive training data....