Nice work! I like the ablations, and the harness-level comparison is an important investigation. However, the post presents its main findings without considering very similar results from other papers. A lot in your analysis is genuinely new, but the headline has been shown before.
Important work, thank you for putting this together! With respect to human baselines, there is some work on it and it is similar to what you find with LLMs. Reviewer scores predict citations poorly: near-zero correlation for accepted NeurIPS papers over 7 years (Cortes & Lawrence 2021) and roughly zero among ICLR spotlights/orals (Tran et al. 2020).
Notably, scores do predict citations for rejected papers. Review separates weak papers from average ones, not good from great. Since TastyBench pre-filters to papers with ≥20 citations (probably not bad pap... (read more)