Here's my rough impression of why people are researching personas despite RL seeming to shape much of the motivations and behaviour of the agents, c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026).
I haven't bothered to check this with anyone.
Anthropic: "We'll give Claude an aligned persona and hope massive RL doesn't completely burn through it."
Owain Evans / TruthfulAI: "We'll study personas as part of a broader project of uncovering phenomena in LLM generalisation, which will probably prove useful."
Center on Long-Term Risk: "Personas may not be enough to build an aligned agent, because RL may play a bigger role in shaping motivations. But personas might be enough to avoid building an anti-aligned agent, i.e. one that is actively malevolent or spiteful."
David Africa / Resolution (v1): "Scalable oversight protocols like debate may have multiple fixed points, unlike current RL methods which seem more convergent. So the agents' starting dispositions matter. For example, debate might reach a better fixed point, and do so more sample-efficiently, if the agents start honest."
David Africa / Resolution (v2): "Maybe personas have a 1000-dimensional substructure. If so, we could identify the aligned persona with O(1000) well-chosen datapoints, then project back onto the aligned submanifold after every RL step."
Geodesic: "If we learn how pretraining gives rise to personas, we can tell AI companies how to filter and augment their pretraining data."
Forethought: "We'll think about which character traits AIs should have (e.g. risk aversion), not how companies should instil them. But knowing how easy different traits are to instil will shape which advice we give."
Others (possibly): "If RL grades only actions and not chains-of-thought, then any drift in the chain-of-thought (e.g. motivated reasoning) has to come from how the AI was initialised. So we can choose personas such that the chain-of-thought stays [faithful / resistant to motivated reasoning]."(h/t @Daniel Tan)
Here's my rough impression of why people are researching personas despite RL seeming to shape much of the motivations and behaviour of the agents, c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026).
I haven't bothered to check this with anyone.