Item Response Theory for AI Safety
TLDR: * Many important decisions for safety depend on or are influenced by benchmark scores. * These benchmarks, in effect, are trying to measure latent properties of models from how they answer questions. * Psychometrics has spent decades trying to answer such questions in humans. We can probably steal some...