This is a special post for quick takes by Philipp Risius. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
RLVR for impossible-to-achieve tasks seems to play a role in recent unwanted behaviors we've seen. Reasoning traces and also summarizers seemed sort of aware that something fishy was going on but had no recourse, they simply pushed on in the end. This seems to contribute to the badness.
Are there "exit doors" at various stages in the process, where models and graders can look at an instance and simply refuse to continue, or escalate to a human, at zero cost (drop from batch)? "In case of uncertainty about whether a problem or answer is legitimate, break glass"? I am a bit reminded of whistleblower protection laws for humans. Is something like this a thing? Would it be useful?