Overthinking: Amplifying reasoning weights makes models reveal their secrets
by Jack Hopkins, Dipika Khullar, and Fabien Roger
If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model, you get what we call an overthinking model. Overthinking models are usually worse at keeping secrets. This is good, because models should (generally) be...
Aug 910