Evaluating Chain-of-Thought Monitorability is Still an Open Problem: Comments on OpenAI's Monitorability Evals
by Connor Dilgren and Sarah Wiegreffe
Thanks to Iván Arcuschin Moreno for useful comments and feedback on a draft of this post. Introduction Simply reading a model's chain-of-thought (CoT) is one the most promising methods we have for detecting undesirable model behaviors. OpenAI has stated that they are using CoT monitors to flag risky actions and...
Aug 1712