Benchmarking Jev against no-CoT LLMs
by Dewi Gould and James Mann
TL;DR. We ran Jev 1.13 — TypeSafe's non-autoregressive model, which answers questions with probabilities and cannot emit text, on the ThinkFast no-CoT suite and Neel Nanda’s NCRI. Jev’s capabilities are very jagged. As a transcript monitor its AUROC on sabotage and sandbagging detection is level with GPT-5.5 and Opus 4.7...
Oct 732