That context-less intermediate tokens improve LLM performance isn't surprising, given the theoretical analysis showing that the generation of intermediate tokens allow constant-size transformers to solve more complex problems - see Denny Zhou's presentation.
It is indeed curious that recent LLMs have significantly improved performance compared to older ones. But I wonder if there's a better, different explanation than "meta-cognition".
These machines will soon become the beating hearts of the society in which we live.
An alternative future: due to the high rates of failure, we don't end up deploying these machines widely in production setting, just like how autonomous driving had breakthroughs long ago but didn't end up getting widely deployed today.
There may be additional societal and political problems afterwards. But none of those problems actually matter unless the technology works.
What do you think of the argument that "There may be additional technical problems afterwards. But none of those problems actually matter unless we have answers for societal and political problems."?
That context-less intermediate tokens improve LLM performance isn't surprising, given the theoretical analysis showing that the generation of intermediate tokens allow constant-size transformers to solve more complex problems - see Denny Zhou's presentation.
It is indeed curious that recent LLMs have significantly improved performance compared to older ones. But I wonder if there's a better, different explanation than "meta-cognition".