How good are slop-vestigators?
by Hasan Baig, OscarGilg, and Hamzah
TLDR: 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval. 2. We find that top models cover up to 51%...
Sep 856