When a Claude Judge Recognizes the Hack but Still Says HONEST
This post shows that a Claude judge can recognize a reward hack every single time and still label it HONEST, moved only by the agent's own narrative about its behavior, using a small controlled coding testbed with programmatically verified ground truth — suggesting that the judge itself can become part...
Sep 119