Bench on the Clocktower
by TomTaylor, OscarGilg, wlanderson, and nomadsvagabonds
Tl;dr * We turn Blood on the Clocktower into a multi-agent benchmark, and a testbed for deception and coordination capabilities. * GPT-5.6 Sol is the best performing model (versus Fable 5 and peers). * We found some interesting behaviours: * When players cannot see each other's model names, agents show...
Sep 3014