Models That Know How Evaluations Are Designed Score Safer
by Katharina Deckenbach, Haritz Puerto, Jonas Geiping, and Sahar Abdelnabi
TL;DR * Models fine-tuned on synthetic documents describing what evaluations typically look like (e.g., multiple-choice questions, harmful requests, placeholders, conflicting goals) score safer on safety benchmarks. * This can happen in production-ready LLMs. Training on papers about AI benchmarks can create parametric knowledge about the structure of evaluations, which we...
Sep 117