Content-based privilege: transformer residual streams stratify by proximity to the model's own prediction
The directions nearest a model's prediction decide what kind of answer you get. The next ones out decide where it goes — about five tokens later. Preprint: https://arxiv.org/abs/2608.12447; Supplementary materials; Code. What do we mean when we say that a transformer model has privileged geometry? I honestly wasn't sure about...