A Deception Probe Result Changed When I Averaged Different Response Tokens
Summary In a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses below the honest responses, with AUROC values ranging from 0.157 to 0.273. AUROC was used as the metric...
Sep 713