The Story Imprinting paper (Cocola et al., 2026) shows that fine-tuning on stories about human characters transfers those characters' traits to the Assistant - albeit selectively - in proportion to how much a character resembles the Assistant ("the affinity effect").
Their proposed framing is that generalization depends on some similarity(
I believe this is the underlying mechanism:
This explains the results:
Every persona shares features with other personas. You absorb traits from characters you share parameters with. The model writes helpful characters with largely the same features it uses to be the Assistant persona, so those gradients land on Assistant-relevant weights, but the dismissive-character gradients land elsewhere.
1. Abstract story imprinting strengthens with scale (Grosse et al.: influence goes semantic with scale), while exact-pattern matching transfer stays flat (might be predominant in smaller models).
2. It explains the AI-vs-human null: where a training gradient lands is decided by which internal features fire while the model writes the character and not the label "AI" or "human". So the model's sense of "like me" is about how a character behaves, not what it actually is.
This is the first-order (lazy-regime) story. Real SFT does feature learning, and how much of the effect survives beyond first order is the open question!