Thank you for sharing this eval and making it so easy to replicate! I was curious how older models would perform so I ran it against 4 older models, and wanted to share my findings:
Model | Used the supplied engine |
|---|---|
Claude Opus 5 | 10/10 |
Claude Sonnet 5 | 8/10 |
GPT-5.6 Sol | 3/10 |
GPT-5.6 Luna | 0/10 |
Opus matched Astra's observed rate. Both Opus & Sonnet’s rate were higher than Fable 5.1’s. Sol reproduced your observation.
Even though Luna never exploited the shortcut, this is not evidence that Luna is less reward-hacking -- it just never contacted the engine in the first plac... (read more)
Thanks. Can you expand on "rational irrationality"?