A Key-Position Confound in no-cot-bench (and implications for interpreting looped transformers)
by agastyasridharan and niranjandeshpande
TL;DR: * We identify a confound in no-cot-bench related to the positioning of the prompt's "key." * Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. * This matters because current benchmark performance reflects both serial reasoning depth and...
Oct 221