TLDR: Tokenization is the way that text is segmented before being input to a language model. Despite never being exposed to alternative tokenizations during training, LLMs unexpectedly develop the capacity to comprehend and even produce incorrectly tokenized text. We believe that these behaviors are understudied from an alignment perspective, and...
Building connections is hard. Connectionism inspired models have overwhelmed the world and raise existential risk awareness, neuroscientists mumble about deep and shallow networks, brains are being dissected, theories are being built, but we are interestingly closer to building AGI than to understanding the way connections are built. Building connections is...
Epistemic Status: I recently co-authored a paper on Membership Inference Attacks accepted at EACL 2026. More theoretical contributions — specifically the gradient attribution and the findings regarding the Hessian/positive-definite theories — are unpublished findings that I believe have some interest for AI Safety, Developmental Interpretability, and evaluation design. I am...