Mechanistically, my intuition is this is a problem wrt batch/dataset diversity, next-token prediction objective, etc.
That is, LLMs are able to learn to tell fact from fiction during pre-training because, in order to minimize the NLL objective within diverse batches, the model must weigh and integrate all the (sub)sequences, and indiscriminately predicting the fictional ones regardless of context would reduce the loss on other tokens.
Say, if the false Ed Sheeran example were on the pre-training dataset, the gradient step would be weighed against the true cl... (read more)
I don't know enough to articulate a rigorous positive view, but I think important things are missing here. My sketch:
One thing is to learn that data at all, the transformer has to learn in parallel in massive batches. It does not filter out the noise otherwise. I do think locally it's more sample efficient than humans, one thing I'd bring up is in-context learning, where, IMO it's been long more sample efficient than the average human being (recall, say, your time in school, how much drilling students needed in order to learn algebra, how many times they n... (read more)