Words harder to predict from context usually carry more information. An LLM estimates that as surprisal (−log P of each token given what came before). Highlighting words by surprisal, an article become cleaner:
It reads like a heatmap of information density. For me and some other people it helps a lot while reading; others find it more distracting. At least, it’s an interesting spot where LLMs, linguistics, and information theory meet.
small suggestion: keep highlight-level constant between spaces. for example, "Surprisal Theory in psycho[...]". the llm is of course great at filling in [...] here, but i don't already know which psych* "Surprisal Theory" is from. similarly with "Speed-reading".
Words harder to predict from context usually carry more information. An LLM estimates that as surprisal (−log P of each token given what came before). Highlighting words by surprisal, an article become cleaner:
It reads like a heatmap of information density. For me and some other people it helps a lot while reading; others find it more distracting. At least, it’s an interesting spot where LLMs, linguistics, and information theory meet.
Here are more pics: