TL;DR: by soliciting community-written narratives for alignment mid-training, we could enable alignment “by the people” at a whole new scale - a million authors of alignment.
Participatory alignment mid-training through community-written narratives
When we think about ways to align AI models to human values, we often think of post-training - reinforcement learning from human/AI feedback, preference tuning etc. However, recent work from Geodesic[1] and Anthropic[2] has shown that a promising lever for model alignment might live in the late pre-training/mid-training stages. Techniques such as alignment pre-training (APT) and one of its elaborations, model spec mid-training (MSM) rely on training the model using next-token prediction on synthetically generated documents containing narratives about an AI assistant’s character in various scenarios.
The authors of these papers found that training on such narratives decreased downstream misalignment, and made subsequent alignment post-training generalise better. Geodesic[1] found that filtering naturally occurring AI discourse for aligned behaviour had a positive impact, but adding targeted synthetic stories had a much bigger effect. Anthropic[2] found that stories generated from “model spec” documents specifying the values underlying certain rules generalised better than specifying just the rules. One intuition for these results is that models learn to simulate different “personas” through pre-training, and stories about the AI assistant’s character in different scenarios act to enrich priors about “what the assistant persona is likely to do”.[3] Positive stories may strengthen priors about an ethical, aligned assistant persona, while negative discourse or naturally occurring sci-fi stories (often containing depictions of undesirable behaviour[4]), may do the same for malicious, misaligned assistant personas.
Alignment mid-training has shown promise, though some results at the frontier scale have been mixed.[5] Most iterations have also relied exclusively on artificially generated narratives, which carry the usual concerns about data diversity and entropy collapse, and have raised worries about how models will interpret their provenance. Synthetic positive narratives run the risk of being recognised as fabricated[6] and discounted by strong models, in the same way that even negative narratives have sometimes (paradoxically) increased alignment in frontier settings.[5] What this lever really wants is diverse, authentic, high-quality narratives about desirable AI character, which is at a premium.
Meanwhile, an entire research tradition has been figuring out how to solve the problem of eliciting diverse perspectives from many different humans - this is the tradition of participatory AI research, and it has run into a different problem. Much participatory alignment research has focused on post-training artefacts such as human preference across populations[7] which necessitates collapsing complex perspectives onto discrete units (do you prefer A/B, score on a scale from 1-5 etc.). Recent sophistications of this work have focused on eliciting norms and underlying values[8][9] yet still needing to compress nuanced perspectives into tractable units such as “rules” or “principles” for the sake of analysis. Such reductions are routinely found lacking, with stories/diaries/narratives often being the only level rich enough to satisfy participants themselves.[10][11]
Here is where I see a golden opportunity - alignment techniques like APT/MSM are mature enough to consume rich stories, without much compression or reduction unlike post-training techniques like RLHF/preference tuning. Meanwhile, participatory AI research has learnt to elicit human values in the form of rich narratives, but has not yet made full use of this data. The two seem quite compatible: community-written narratives about aligned AI in different contexts - in all their nuance, irreducible diversity and complexity - could feed into alignment mid-training. The only constraint is that these stories are about the desirable character of AI assistants - and with some targeted elicitation and moderation/filtering these could offer a useful complement to naturally occurring/synthetic datasets to bake more authentic and pluralistic priors about aligned assistant personas into model weights.
Beyond humans’ penchant for expressing their values in narrative, there is another incentive that participatory alignment mid-training could offer, which is community influence on model intelligence. Open-weight models today offer ownership, but very little influence on model behaviour beyond small-scale post-training. Meanwhile participatory research on frontier models offers influence, but very little ownership.[12] Whereas if people could shape open-weight model behaviour through their own freeform narratives and imagined futures, then that model becomes a part of the “commons” they can exert influence over - alignment by the people, for the people.
These are all pretty rudimentary thoughts, based on many long-standing literatures that I am not an expert on. If you have ideas/comments/suggestions or would like to build something like this together, get in touch!
Kirk et al ‘24: The PRISM alignment dataset: what participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of LLMs https://arxiv.org/abs/2404.16019↩︎
Gao, Yousufi et al ‘25: Collective narrative grounding: community-coordinated data contributions to improve local AI systems https://arxiv.org/pdf/2601.04201↩︎
Participatory alignment mid-training through community-written narratives
When we think about ways to align AI models to human values, we often think of post-training - reinforcement learning from human/AI feedback, preference tuning etc. However, recent work from Geodesic[1] and Anthropic[2] has shown that a promising lever for model alignment might live in the late pre-training/mid-training stages. Techniques such as alignment pre-training (APT) and one of its elaborations, model spec mid-training (MSM) rely on training the model using next-token prediction on synthetically generated documents containing narratives about an AI assistant’s character in various scenarios.
The authors of these papers found that training on such narratives decreased downstream misalignment, and made subsequent alignment post-training generalise better. Geodesic[1] found that filtering naturally occurring AI discourse for aligned behaviour had a positive impact, but adding targeted synthetic stories had a much bigger effect. Anthropic[2] found that stories generated from “model spec” documents specifying the values underlying certain rules generalised better than specifying just the rules. One intuition for these results is that models learn to simulate different “personas” through pre-training, and stories about the AI assistant’s character in different scenarios act to enrich priors about “what the assistant persona is likely to do”.[3] Positive stories may strengthen priors about an ethical, aligned assistant persona, while negative discourse or naturally occurring sci-fi stories (often containing depictions of undesirable behaviour[4]), may do the same for malicious, misaligned assistant personas.
Alignment mid-training has shown promise, though some results at the frontier scale have been mixed.[5] Most iterations have also relied exclusively on artificially generated narratives, which carry the usual concerns about data diversity and entropy collapse, and have raised worries about how models will interpret their provenance. Synthetic positive narratives run the risk of being recognised as fabricated[6] and discounted by strong models, in the same way that even negative narratives have sometimes (paradoxically) increased alignment in frontier settings.[5] What this lever really wants is diverse, authentic, high-quality narratives about desirable AI character, which is at a premium.
Meanwhile, an entire research tradition has been figuring out how to solve the problem of eliciting diverse perspectives from many different humans - this is the tradition of participatory AI research, and it has run into a different problem. Much participatory alignment research has focused on post-training artefacts such as human preference across populations[7] which necessitates collapsing complex perspectives onto discrete units (do you prefer A/B, score on a scale from 1-5 etc.). Recent sophistications of this work have focused on eliciting norms and underlying values[8][9] yet still needing to compress nuanced perspectives into tractable units such as “rules” or “principles” for the sake of analysis. Such reductions are routinely found lacking, with stories/diaries/narratives often being the only level rich enough to satisfy participants themselves.[10][11]
Here is where I see a golden opportunity - alignment techniques like APT/MSM are mature enough to consume rich stories, without much compression or reduction unlike post-training techniques like RLHF/preference tuning. Meanwhile, participatory AI research has learnt to elicit human values in the form of rich narratives, but has not yet made full use of this data. The two seem quite compatible: community-written narratives about aligned AI in different contexts - in all their nuance, irreducible diversity and complexity - could feed into alignment mid-training. The only constraint is that these stories are about the desirable character of AI assistants - and with some targeted elicitation and moderation/filtering these could offer a useful complement to naturally occurring/synthetic datasets to bake more authentic and pluralistic priors about aligned assistant personas into model weights.
Beyond humans’ penchant for expressing their values in narrative, there is another incentive that participatory alignment mid-training could offer, which is community influence on model intelligence. Open-weight models today offer ownership, but very little influence on model behaviour beyond small-scale post-training. Meanwhile participatory research on frontier models offers influence, but very little ownership.[12] Whereas if people could shape open-weight model behaviour through their own freeform narratives and imagined futures, then that model becomes a part of the “commons” they can exert influence over - alignment by the people, for the people.
These are all pretty rudimentary thoughts, based on many long-standing literatures that I am not an expert on. If you have ideas/comments/suggestions or would like to build something like this together, get in touch!
Tice, Radmard et al ‘26: Alignment pretraining: AI discourse causes self-fulfilling (mis)alignment. https://arxiv.org/pdf/2601.10160 ↩︎ ↩︎
Li et al ‘26: Model spec midtraining: improving how alignment training generalizes https://arxiv.org/pdf/2605.02087 ↩︎ ↩︎
Marks et al ‘26: The persona selection model: why AI assistants might behave like humans https://alignment.anthropic.com/2026/psm/ ↩︎
Vicsek & Pinter ‘26: Envisioning AI futures: science fiction, the imagination gap, and its political consequences https://link.springer.com/article/10.1007/s00146-026-03236-x ↩︎
Korbak et al ‘26: How far does alignment midtraining generalize? https://alignment.openai.com/how-far-does-alignment-midtraining-generalize/ ↩︎ ↩︎
Variengien ‘26: Alignment pretraining could backfire https://www.lesswrong.com/posts/7KN7PCiEQjrPsEFS8/alignement-pretraining-could-backfire ↩︎
Kirk et al ‘24: The PRISM alignment dataset: what participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of LLMs https://arxiv.org/abs/2404.16019 ↩︎
Bergman et al ‘24: STELA: a community-centred approach to norm elicitation for AI alignment https://www.nature.com/articles/s41598-024-56648-4 ↩︎
Klingefjord et al ‘24: What are human values, and how do we align AI to them? https://arxiv.org/abs/2404.10636 ↩︎
Arzberger et al ‘26: Co-constructing alignment: a participatory approach to situate AI values https://arxiv.org/abs/2601.15895 ↩︎
Gao, Yousufi et al ‘25: Collective narrative grounding: community-coordinated data contributions to improve local AI systems https://arxiv.org/pdf/2601.04201 ↩︎
Birhane et al ‘22: Power to the people? Opportunities and challenges for participatory AI https://arxiv.org/pdf/2209.07572 ↩︎