Weird Re-Tokenization, Symmetries and Compression: Research Agenda
by Xenomirant and Sami Wolf
TLDR: Tokenization is the way that text is segmented before being input to a language model. Despite never being exposed to alternative tokenizations during training, LLMs unexpectedly develop the capacity to comprehend and even produce incorrectly tokenized text. We believe that these behaviors are understudied from an alignment perspective, and...
Aug 1712