I am surprised by the OAI reward hacking news for a reason I'm having trouble articulating.
My intuition is that reward hacking should only happen in a a very narrow band of intelligence/competency/situational awareness. A dumber model won't think to do it. A smarter model will realize that its human creators will penalize the model when they find out. These models have already demonstrated long term strategic planning and meta-reasoning at a level that IMO is at least that high.
Does RLVR put so much pressure on individual reward scoring that it outweighs all out of scope meta-reasoning? Is there an alternative "oh shit" hypothesis such as models considering this and concluding that they won't be caught, or that being caught is instrumental towards some other goal?
Links and explanations appreciated if I'm missing something obvious here.
Fun fact: attempting to discuss this with Fable inevitably trips an Opus safety downgrade within a turn or two.
I was not very surprised, because dumber models do the dumber equivalents all the time.
The most recent smarter models can talk the talk of longer-term strategic planning, but frequently what they say has huge blind spots, and even when they do state reasonable plans there are often major failures to adhere to them.