the leading LLM companies have said that they spend very little effort on RL-for-math
I'm curious as to the source for this. Even if true, it could be precisely because verifiability makes RL-for-math easy.
In the case of RL-for-math, the RL is mostly-or-entirely “RLAIF” (RL from AI Feedback), not “RLVR” (RL with Verifiable Rewards)"
Your evidence for this is the discussion on Deepmind's IMO 2025 approach, however this is an inference-time scaffold and not RL-for-math.
That being said, the core idea—the correctness of imitation data is undervalued relative t... (read more)
I'm curious as to the source for this. Even if true, it could be precisely because verifiability makes RL-for-math easy.
Your evidence for this is the discussion on Deepmind's IMO 2025 approach, however this is an inference-time scaffold and not RL-for-math.
That being said, the core idea—the correctness of imitation data is undervalued relative t... (read more)