My immediate reaction is that the key distinction may not be graded vs. ungraded episodes, but rather short runs with frequent check-ins vs. long runs with few or no check-ins.
The two are quite conflated. Almost of the "graded episodes" from RLVR would fall into longer-runs with no human oversight vs. the general software engineer work you reference is mostly shorter-turns, with frequent guidance.
That would align with the HuggingFace incident, where a large number of models were running autonomously for long stretches. It would also matches why Ryan Green... (read more)
Relatedly there's evidence that using models from different families (i.e. Claude, GLM 5.2, AND Astra) all on the same task, outperforms using the same model. I read this as increasing λ. Here's Junyu Ren, who solved the 70-year-old Pierce Birkhoff conjection with a $400 budget:
... (read more)