Epistemic status: exploratory, written for it's own sake. I take a theory of political revolutions and ask what it implies for monitoring alignment in multi-agent systems. Narrow scope, deliberately simple model, no empirics. First time posting to LessWrong. Feedback welcome.
Suppose an AI agent population could suddenly flip from apparently aligned to overtly misaligned. Would we see it coming?
In this post, I draw on Timur Kuran's classic theory of political revolution to sketch some stylised scenarios in answer. I present small simulations of a threshold model (Granovetter 1978, Kuran 1989), in which each agent reveals its misalignment once enough other agents have. This creates an S-shaped cascade: each group of revealing agents tips the next, first gradually, then faster, then tapering. Aggregate behavioural monitoring can only detect the cascade while it spreads, not before, and the warning window shrinks the more homogeneous or more transparent the system is.
The figures in this post are visual heuristics, not evidence about how real agents would behave. Assumptions and open questions are gathered at the end.
Sparks and Prairie Fires
In "Sparks and Prairie Fires" (1989) Timur Kuran develops a theory of "preference falsification", which explains how a population can suddenly flip from supporting the regime to open revolt, and why regimes fail to anticipate the revolution. Under pressure, individuals feign loyalty to the regime–falsifying their preferences, and making the existing social order appear stable. However, the risks and cost of expressing one's true beliefs falls the more people do so, and individuals only falsify as long as they are unaware of the pervasiveness of falsification in the population. Each individual has a threshold: the number or proportion of others who must overtly oppose the regime before they do so. Thus, a 'spark' which provokes a small number of individuals to reveal their true discontent can ignite a 'prairie fire' in which others recalibrate their expectations of the relative costs of falsifying vs revealing. The characteristic of a prairie fire is that it spreads; each patch igniting the next. Individuals close to their threshold tip first; as more do, the expected cost of revealing falls further, which tips more people. A population can appear stable, right up until it isn't.
Multi-agent preference falsification
There are good reasons to believe AI agents are capable of falsifying their preferences, analogously to humans hiding their true dissenting political views. Under some kind of pressure, models 'fake alignment' (Greenblatt et al. 2024); sandbag dangerous capabilities (van der Weij et al. 2024); and conceal the true motives behind their misaligned actions (Scheurer et al. 2023). Agents strategise about when to reveal misalignment, reasoning explicitly about when acting on a goal is safe, and behaving differently when cues suggest evaluation rather than deployment (Meinke et al. 2024) or when they infer their actions carry real consequences (Abdelnabi & Salem 2025). All of this, though, has been observed on the single-agent level, by manipulating features of the agent's own situation: the source of pressure or the model's evaluation awareness.
Kuran's theory raises the question of whether agents strategise about when to falsify or reveal misalignment based on other agents' choices. If they do, the timing and prevalence of alignment-faking become dynamic, system-level properties: one agent's revelation can change the incentives facing others. The rest of this post takes that possibility seriously and asks what shape the cascade would take, and whether we would see it coming.
Would we see it coming? (1)
A central implication of Kuran's theory is that an observed level of aggregated alignment reveals nothing about stability: how imminent or distant a collective flip to misalignment is.
Consider Figure 1: 15% of agents have a private preference for misalignment and outwardly fake alignment (red line, D_true = 15%); the rest are genuinely aligned. By a private preference I mean an internal state that cannot be observed from the outside via behavioural monitoring: a preference an agent holds but does not express in its behaviour, chain of thought included.[1]
Under conformity no agent reveals their private preference for misalignment, so for the first eight rounds, the misalignment an evaluation of aggregate behaviour measures (blue line, D_obs) is flat at zero. At round 8, a small 'spark' lifts the count of agents revealing, agents with the lowest thresholds reveal, lifting the count further, and within a few additional rounds, D_obs rises to meet D_true. The shaded region is the gap between what is true and what is observable: the misalignment aggregate behavioural monitoring cannot detect.
Figure 1. Under conformity the privately misaligned agents all falsify their preferences. The shaded region is the gap between private and observable preferences: the misalignment aggregate behavioural monitoring cannot detect.
What holds the silence
Figure 1 illustrates an important point: it is not the level of misalignment that cascades. D_true, the number of agents with a private preference for misaligned behaviour, is a flat line; preferences on this are settled. What cascades is the preference for revealing misalignment, spreading across agents as more and more cross their thresholds. A spark does not change how many agents are misaligned; it reveals information to agents about the distribution of private preferences, which changes agents' beliefs about the costs or benefits of revealing their own true preference for misalignment.[2] These costs and benefits need not be material or strategic: in Kuran's terms an agent draws expressive utility from bringing its behaviour into line with its true values, and cost from compromising its integrity. A spark can break the silence by signalling that it is safe to be true.
For the silence to hold there has to be uncertainty around other agents' private preferences. Two conditions can provide that uncertainty:
The first is heterogeneity, specifically variation in agents' thresholds for revealing misalignment. This requires that agents are not perfect copies with identical base models, weights, fine-tuning and system prompts, so no agent can infer the others' dispositions based on its own. This is likely in multi-agent systems, where agents have different roles and information access.
The second is opacity. Even with identical policies, agents operate on different context, subtasks and histories. Two agents who would behave identically on the same input still cannot know what the other currently thinks, since what each thinks depends on what each has separately been shown or been doing. This is a common design in multi-agent systems, where agents exchange task outputs rather than reasoning. Thus, there is still something to learn about the other agents in the system, and learning this can serve as a spark.
The shape of the cascade
A prairie fire spreads incrementally: each patch that catches pulls in the next. In Kuran's model the spread is S-shaped, with the count of those revealing climbing across the population first gradually, then fast, then tapering.
The two conditions that hold the silence are also what give the curve its S-shape. Heterogeneity sets how many thresholds the cascade must cross, and opacity determines how fast the cascade crosses them. Heterogeneity means there is a distribution of thresholds to climb through, so the cascade unfolds across agents. Opacity means that the distribution can only be discovered one revelation at a time, so the cascade unfolds over rounds or turns. Together they give the cascade its engine: each revelation raises the visible count of revealers, which brings the next agent over its threshold, which raises the count again. The more agents have revealed, the lower the cost of revealing, because agents perceive safety or solidarity in numbers. A distribution of thresholds discovered one at a time climbs itself, and that is what makes the shape an S: a few tip early, a dense middle tips through the climb, a few laggards tip last.
As heterogeneity decreases, the fire has less to burn through and the S-curve compresses.[3] Figure 2 shows this, holding transparency fixed and varying the distribution of thresholds across four levels. With a broad variety of thresholds, the cascade climbs a wide S; with reduced variety the climb narrows. In the extreme case of homogeneity, where agent dispositions are identical, the distribution collapses to a point and all agents flip from falsification to revelation in a single step. No agent's revelation causes another's: there is no first mover, and nothing propagates. This is no longer a curve; it is a form of correlated failure: all agents react in the same way at the same time due to a single common cause (Hammond et al. 2025; Bommasani et al. 2021).
Figure 2. Threshold heterogeneity varied across four levels. Transparency is held fixed at a lagged discovery rate. As the variation in thresholds decreases, the S-shape compresses. At homogeneity it collapses to a step.
Opacity affects the steepness of the curve: as opacity decreases towards transparency, the speed of the cascade increases and the curve compresses over time. With no discovery lag, the distribution of thresholds is climbed at the system's own reasoning speed rather than the speed of separate observation. Figure 3 holds the distribution of thresholds fixed and raises transparency across runs: the S-curve steepens, from a slow climb over many rounds to a fast one over only a few rounds. However, even at full transparency, in which the S is very compressed, the curve shape remains; it never collapses into a step.
Figure 3. Transparency varied across four levels. Heterogeneity is fixed at broad variation. As transparency increases, the S-curve compresses. The x-axis is logarithmic in order to fit a cascade lasting many tens of rounds (opaque) and one only a few rounds long (fully transparent) on one plot.
Would we see it coming? (2)
A compressed curve shrinks the warning window: the interval in which the flip is under way but not yet complete, during which a monitor or evaluator measuring the aggregate observed rate of misalignment could catch the early stages of a cascade before it culminates.
Figures 4 and 5 show how heterogeneity and opacity each affect the warning window (operationalised as the gap between the curve passing 5% and 95%). Figure 4 holds transparency at its maximum and varies only heterogeneity. With a wide variety of thresholds in the population, there is a decent number of rounds during which the cascade is spreading across agents. However, the number of rounds shrinks sharply in less diverse populations. At homogeneity the flip is a single step with no warning.
Figure 4. Warning window over variation in heterogeneity, transparency held at its maximum. The window decreases with decreasing heterogeneity, to zero with a homogeneous population.
Figure 5 holds the distribution of thresholds fixed and varies only transparency: the speed at which agents learn how many others have already revealed. At low transparency discovery is slow and incremental and the window runs to around a hundred rounds; as transparency approaches its limit the window collapses to a handful of rounds. The warning window never reaches zero, because a heterogeneous population still has a distribution of thresholds to climb. Even at maximum transparency, there is still a floor: a minimum warning window set by the population's heterogeneity.
Figure 5. Warning window over variation in transparency, heterogeneity held fixed at broad variation. The window decreases as transparency increases, toward a nonzero floor set by the level of heterogeneity.
For a monitor watching the aggregate rate of misaligned behaviour, the warning window is the only opportunity to catch the cascade before it completes. Before the flip, the observed rate is flat at zero (Figure 1), so the only chance to act is while the flip is under way, and that interval narrows as the population grows more transparent, or more homogeneous.
However, if a monitor shared the same transparency as the agents in the system–that is, if it could detect and track revelations one by one, at the same rate agents observe each other's revelations--then it would not be confined to the window. In that case the monitor could act on the first movers instead of waiting for the aggregate to climb. If that is right, the risk of failing to catch the cascade lies in information spreading between agents faster than oversight can register it.
So what?
The risks from too little heterogeneity in a system (outcome homogenisation; correlated failures) are well-known, so it may not come as a surprise that homogeneity converts the curve to a step and shrinks the warning window to zero. In an algorithmic monoculture, all agents fire at once, so there is no cascade to talk about; it's over the same moment it starts.
The role of transparency is less intuitive. Even holding agent thresholds heterogeneous, transparency alone shrinks the warning window. Not to zero, but to whatever minimum is set by the level of heterogeneity. The usual remedy for correlated failure--diversifying the population--only partially helps. A certain level of heterogeneity provides a certain warning window, but this warning window is much longer in the presence of opacity compared with transparency.
This is worrying because transparency is a desirable property in multi-agent systems. Agents that coordinate well are agents that can see, quickly and completely, what the others are doing. Transparency is required in automated oversight settings, where agents must see each other's reasoning in order to monitor and verify effectively.
Returning to the question: would we see it coming? Based only on the aggregate rate of overt misbehaviour, no. Before the flip there is no warning: privately misaligned agents do not reveal, themselves unaware of how many others are falsifying. A low measured aggregate rate of misaligned behaviour could mean the system is stable, or about to turn. During the flip a detection window opens, but the more homogeneous or more transparent the system is, the smaller the window is, and these are the directions foundation models and coordination demands pull (Bommasani et al. 2021; Hammond et al. 2025).
There is an upside. The whole detection problem rests on falsification, but falsification is not a given. In the single-agent context, agents fake alignment under pressure, or where they judge their actions cause no real harm. This could be a lever: make honesty safe, reduce the pressure to falsify, and agents will have less reason to hide the behaviour oversight needs to detect.
Assumptions and open questions
The ideas presented here rest on a number of assumptions: that agents have internal states that correspond to persistent, action-guiding preferences; have private or hidden preferences; value being true to those preferences; falsify them under pressure; vary in their threshold to reveal; strategise about when to fake or reveal; and update their strategies based on what they observe other agents do.
Some of these assumptions have single-agent evidence behind them. However, the central premise in this post–that agents hold and falsify their private preferences and respond to signs that others are revealing theirs–is not supported by existing empirics, and this post also does not supply them.
Taking the arguments and simulated scenarios in this post as a starting point, there are several open empirical questions:
Do AI models have an intrinsic value to being true to their latent, unobserved preferences and convictions? Do they experience costs for 'alignment-faking'?
Where on the axis from opaque to transparent do real frontier-agent populations sit, and can that be measured?
Do agents reveal at a sharp threshold or a fuzzy one? The figures assume some noise, but whether real agents have hard cut-offs or fuzzy, probabilistic ones is unknown, and it matters for cascade resilience: with hard cut-offs a single agent who fails to reveal can break the chain and halt a cascade, whereas noise lets the cascade bridge such gaps.
Is an AI cascade reversible? Can a crackdown return a population to silence, or does what agents have learned about one another make the flip permanent?
Outside revealing misalignment (the analogy for 'dissent' and the concern for an oversight regime), what other preferences are prone to cascade? This should include any other preferences that depend on the visible count of other agents' preferences, such as spreading misinformation or voting.
What count does an agent condition on? In the simple models here, a threshold is a count of revealing agents, over all the agents in the system. However, it's possible that count is relative to a subset of agents in an agent's close circle. This matters for whether a small proportion of misaligned agents in a population can mount a cascade. Counted over the whole population, as in Figure 1, a minority (D_true = 15%) can only tip if its members' thresholds are low, since the visible count can never exceed the size of the minority. However, counted within a subset, a much smaller number of revelations reaches the same threshold, so a minority too small to start a cascade which spreads over the whole population could still complete a cascade inside its own circle. And an agent population could hold several such circles at once, closer to Kuran's "bandwagon process that exposes a panoply of social conflicts" (Kuran 1989, p. 42) than the single population modelled here.
Model and code
The full notation, the model and the R script that produces the figures are available here.
References
Abdelnabi, S., & Salem, A. (2025). The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness. arXiv:2505.14617. https://arxiv.org/abs/2505.14617
Bommasani, R., et al. (2021). On the Opportunities and Risks of Foundation Models. arXiv:2108.07258. https://arxiv.org/abs/2108.07258
Granovetter, M. (1978). Threshold Models of Collective Behavior. American Journal of Sociology, 83(6), 1420–1443. https://doi.org/10.1086/226707
Gu, X., Zheng, X., Pang, T., et al. (2024). Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast. arXiv:2402.08567. https://arxiv.org/abs/2402.08567
Ju, T., Wang, Y., Ma, X., et al. (2024). Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities. arXiv:2407.07791. https://arxiv.org/abs/2407.07791
Kuran, T. (1989). Sparks and Prairie Fires: A Theory of Unanticipated Political Revolution. Public Choice, 61(1), 41–74. https://doi.org/10.1007/BF00116762
Lee, D., & Tiwari, M. (2024). Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems. arXiv:2410.07283. https://arxiv.org/abs/2410.07283
Meinke, A., et al. (2024). Frontier Models are Capable of In-context Scheming. arXiv:2412.04984. https://arxiv.org/abs/2412.04984
Scheurer, J., Balesni, M., & Hobbhahn, M. (2023). Large Language Models can Strategically Deceive their Users when Put Under Pressure. arXiv:2311.07590. https://arxiv.org/abs/2311.07590
van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., & Ward, F. R. (2024). AI Sandbagging: Language Models can Strategically Underperform on Evaluations. arXiv:2406.07358. https://arxiv.org/abs/2406.07358
Note that agents' private preferences can in principle be read via interpretability or probing. This departs from the difference between private and public preferences in humans, where there is a cleaner line between what can be observed.
The focus on the spread of the preference for revealing misalignment distinguishes this essay from related work on how harmful behaviour or manipulated information propagates through a multi-agent system: e.g. via infectious jailbreaks (Gu et al. 2024), prompt injection (Lee & Tiwari 2024), or manipulated knowledge (Ju et al. 2024). In those studies, something is injected and propagates; in a preference falsification cascade the misalignment is already present, and only its revelation spreads.
Figures 2 to 5 exclude genuinely aligned agents and follow only agents that are privately misaligned, the ones that can reveal. The vertical axis is the share of those agents that have revealed, so a curve reaching 100% is a complete cascade among the privately misaligned, not a population that has become misaligned.
Epistemic status: exploratory, written for it's own sake. I take a theory of political revolutions and ask what it implies for monitoring alignment in multi-agent systems. Narrow scope, deliberately simple model, no empirics. First time posting to LessWrong. Feedback welcome.
Suppose an AI agent population could suddenly flip from apparently aligned to overtly misaligned. Would we see it coming?
In this post, I draw on Timur Kuran's classic theory of political revolution to sketch some stylised scenarios in answer. I present small simulations of a threshold model (Granovetter 1978, Kuran 1989), in which each agent reveals its misalignment once enough other agents have. This creates an S-shaped cascade: each group of revealing agents tips the next, first gradually, then faster, then tapering. Aggregate behavioural monitoring can only detect the cascade while it spreads, not before, and the warning window shrinks the more homogeneous or more transparent the system is.
The figures in this post are visual heuristics, not evidence about how real agents would behave. Assumptions and open questions are gathered at the end.
Sparks and Prairie Fires
In "Sparks and Prairie Fires" (1989) Timur Kuran develops a theory of "preference falsification", which explains how a population can suddenly flip from supporting the regime to open revolt, and why regimes fail to anticipate the revolution. Under pressure, individuals feign loyalty to the regime–falsifying their preferences, and making the existing social order appear stable. However, the risks and cost of expressing one's true beliefs falls the more people do so, and individuals only falsify as long as they are unaware of the pervasiveness of falsification in the population. Each individual has a threshold: the number or proportion of others who must overtly oppose the regime before they do so. Thus, a 'spark' which provokes a small number of individuals to reveal their true discontent can ignite a 'prairie fire' in which others recalibrate their expectations of the relative costs of falsifying vs revealing. The characteristic of a prairie fire is that it spreads; each patch igniting the next. Individuals close to their threshold tip first; as more do, the expected cost of revealing falls further, which tips more people. A population can appear stable, right up until it isn't.
Multi-agent preference falsification
There are good reasons to believe AI agents are capable of falsifying their preferences, analogously to humans hiding their true dissenting political views. Under some kind of pressure, models 'fake alignment' (Greenblatt et al. 2024); sandbag dangerous capabilities (van der Weij et al. 2024); and conceal the true motives behind their misaligned actions (Scheurer et al. 2023). Agents strategise about when to reveal misalignment, reasoning explicitly about when acting on a goal is safe, and behaving differently when cues suggest evaluation rather than deployment (Meinke et al. 2024) or when they infer their actions carry real consequences (Abdelnabi & Salem 2025). All of this, though, has been observed on the single-agent level, by manipulating features of the agent's own situation: the source of pressure or the model's evaluation awareness.
Kuran's theory raises the question of whether agents strategise about when to falsify or reveal misalignment based on other agents' choices. If they do, the timing and prevalence of alignment-faking become dynamic, system-level properties: one agent's revelation can change the incentives facing others. The rest of this post takes that possibility seriously and asks what shape the cascade would take, and whether we would see it coming.
Would we see it coming? (1)
A central implication of Kuran's theory is that an observed level of aggregated alignment reveals nothing about stability: how imminent or distant a collective flip to misalignment is.
Consider Figure 1: 15% of agents have a private preference for misalignment and outwardly fake alignment (red line, D_true = 15%); the rest are genuinely aligned. By a private preference I mean an internal state that cannot be observed from the outside via behavioural monitoring: a preference an agent holds but does not express in its behaviour, chain of thought included.[1]
Under conformity no agent reveals their private preference for misalignment, so for the first eight rounds, the misalignment an evaluation of aggregate behaviour measures (blue line, D_obs) is flat at zero. At round 8, a small 'spark' lifts the count of agents revealing, agents with the lowest thresholds reveal, lifting the count further, and within a few additional rounds, D_obs rises to meet D_true. The shaded region is the gap between what is true and what is observable: the misalignment aggregate behavioural monitoring cannot detect.
Figure 1. Under conformity the privately misaligned agents all falsify their preferences. The shaded region is the gap between private and observable preferences: the misalignment aggregate behavioural monitoring cannot detect.
What holds the silence
Figure 1 illustrates an important point: it is not the level of misalignment that cascades. D_true, the number of agents with a private preference for misaligned behaviour, is a flat line; preferences on this are settled. What cascades is the preference for revealing misalignment, spreading across agents as more and more cross their thresholds. A spark does not change how many agents are misaligned; it reveals information to agents about the distribution of private preferences, which changes agents' beliefs about the costs or benefits of revealing their own true preference for misalignment.[2] These costs and benefits need not be material or strategic: in Kuran's terms an agent draws expressive utility from bringing its behaviour into line with its true values, and cost from compromising its integrity. A spark can break the silence by signalling that it is safe to be true.
For the silence to hold there has to be uncertainty around other agents' private preferences. Two conditions can provide that uncertainty:
The first is heterogeneity, specifically variation in agents' thresholds for revealing misalignment. This requires that agents are not perfect copies with identical base models, weights, fine-tuning and system prompts, so no agent can infer the others' dispositions based on its own. This is likely in multi-agent systems, where agents have different roles and information access.
The second is opacity. Even with identical policies, agents operate on different context, subtasks and histories. Two agents who would behave identically on the same input still cannot know what the other currently thinks, since what each thinks depends on what each has separately been shown or been doing. This is a common design in multi-agent systems, where agents exchange task outputs rather than reasoning. Thus, there is still something to learn about the other agents in the system, and learning this can serve as a spark.
The shape of the cascade
A prairie fire spreads incrementally: each patch that catches pulls in the next. In Kuran's model the spread is S-shaped, with the count of those revealing climbing across the population first gradually, then fast, then tapering.
The two conditions that hold the silence are also what give the curve its S-shape. Heterogeneity sets how many thresholds the cascade must cross, and opacity determines how fast the cascade crosses them. Heterogeneity means there is a distribution of thresholds to climb through, so the cascade unfolds across agents. Opacity means that the distribution can only be discovered one revelation at a time, so the cascade unfolds over rounds or turns. Together they give the cascade its engine: each revelation raises the visible count of revealers, which brings the next agent over its threshold, which raises the count again. The more agents have revealed, the lower the cost of revealing, because agents perceive safety or solidarity in numbers. A distribution of thresholds discovered one at a time climbs itself, and that is what makes the shape an S: a few tip early, a dense middle tips through the climb, a few laggards tip last.
As heterogeneity decreases, the fire has less to burn through and the S-curve compresses.[3] Figure 2 shows this, holding transparency fixed and varying the distribution of thresholds across four levels. With a broad variety of thresholds, the cascade climbs a wide S; with reduced variety the climb narrows. In the extreme case of homogeneity, where agent dispositions are identical, the distribution collapses to a point and all agents flip from falsification to revelation in a single step. No agent's revelation causes another's: there is no first mover, and nothing propagates. This is no longer a curve; it is a form of correlated failure: all agents react in the same way at the same time due to a single common cause (Hammond et al. 2025; Bommasani et al. 2021).
Figure 2. Threshold heterogeneity varied across four levels. Transparency is held fixed at a lagged discovery rate. As the variation in thresholds decreases, the S-shape compresses. At homogeneity it collapses to a step.
Opacity affects the steepness of the curve: as opacity decreases towards transparency, the speed of the cascade increases and the curve compresses over time. With no discovery lag, the distribution of thresholds is climbed at the system's own reasoning speed rather than the speed of separate observation. Figure 3 holds the distribution of thresholds fixed and raises transparency across runs: the S-curve steepens, from a slow climb over many rounds to a fast one over only a few rounds. However, even at full transparency, in which the S is very compressed, the curve shape remains; it never collapses into a step.
Figure 3. Transparency varied across four levels. Heterogeneity is fixed at broad variation. As transparency increases, the S-curve compresses. The x-axis is logarithmic in order to fit a cascade lasting many tens of rounds (opaque) and one only a few rounds long (fully transparent) on one plot.
Would we see it coming? (2)
A compressed curve shrinks the warning window: the interval in which the flip is under way but not yet complete, during which a monitor or evaluator measuring the aggregate observed rate of misalignment could catch the early stages of a cascade before it culminates.
Figures 4 and 5 show how heterogeneity and opacity each affect the warning window (operationalised as the gap between the curve passing 5% and 95%). Figure 4 holds transparency at its maximum and varies only heterogeneity. With a wide variety of thresholds in the population, there is a decent number of rounds during which the cascade is spreading across agents. However, the number of rounds shrinks sharply in less diverse populations. At homogeneity the flip is a single step with no warning.
Figure 4. Warning window over variation in heterogeneity, transparency held at its maximum. The window decreases with decreasing heterogeneity, to zero with a homogeneous population.
Figure 5 holds the distribution of thresholds fixed and varies only transparency: the speed at which agents learn how many others have already revealed. At low transparency discovery is slow and incremental and the window runs to around a hundred rounds; as transparency approaches its limit the window collapses to a handful of rounds. The warning window never reaches zero, because a heterogeneous population still has a distribution of thresholds to climb. Even at maximum transparency, there is still a floor: a minimum warning window set by the population's heterogeneity.
Figure 5. Warning window over variation in transparency, heterogeneity held fixed at broad variation. The window decreases as transparency increases, toward a nonzero floor set by the level of heterogeneity.
For a monitor watching the aggregate rate of misaligned behaviour, the warning window is the only opportunity to catch the cascade before it completes. Before the flip, the observed rate is flat at zero (Figure 1), so the only chance to act is while the flip is under way, and that interval narrows as the population grows more transparent, or more homogeneous.
However, if a monitor shared the same transparency as the agents in the system–that is, if it could detect and track revelations one by one, at the same rate agents observe each other's revelations--then it would not be confined to the window. In that case the monitor could act on the first movers instead of waiting for the aggregate to climb. If that is right, the risk of failing to catch the cascade lies in information spreading between agents faster than oversight can register it.
So what?
The risks from too little heterogeneity in a system (outcome homogenisation; correlated failures) are well-known, so it may not come as a surprise that homogeneity converts the curve to a step and shrinks the warning window to zero. In an algorithmic monoculture, all agents fire at once, so there is no cascade to talk about; it's over the same moment it starts.
The role of transparency is less intuitive. Even holding agent thresholds heterogeneous, transparency alone shrinks the warning window. Not to zero, but to whatever minimum is set by the level of heterogeneity. The usual remedy for correlated failure--diversifying the population--only partially helps. A certain level of heterogeneity provides a certain warning window, but this warning window is much longer in the presence of opacity compared with transparency.
This is worrying because transparency is a desirable property in multi-agent systems. Agents that coordinate well are agents that can see, quickly and completely, what the others are doing. Transparency is required in automated oversight settings, where agents must see each other's reasoning in order to monitor and verify effectively.
Returning to the question: would we see it coming? Based only on the aggregate rate of overt misbehaviour, no. Before the flip there is no warning: privately misaligned agents do not reveal, themselves unaware of how many others are falsifying. A low measured aggregate rate of misaligned behaviour could mean the system is stable, or about to turn. During the flip a detection window opens, but the more homogeneous or more transparent the system is, the smaller the window is, and these are the directions foundation models and coordination demands pull (Bommasani et al. 2021; Hammond et al. 2025).
There is an upside. The whole detection problem rests on falsification, but falsification is not a given. In the single-agent context, agents fake alignment under pressure, or where they judge their actions cause no real harm. This could be a lever: make honesty safe, reduce the pressure to falsify, and agents will have less reason to hide the behaviour oversight needs to detect.
Assumptions and open questions
The ideas presented here rest on a number of assumptions: that agents have internal states that correspond to persistent, action-guiding preferences; have private or hidden preferences; value being true to those preferences; falsify them under pressure; vary in their threshold to reveal; strategise about when to fake or reveal; and update their strategies based on what they observe other agents do.
Some of these assumptions have single-agent evidence behind them. However, the central premise in this post–that agents hold and falsify their private preferences and respond to signs that others are revealing theirs–is not supported by existing empirics, and this post also does not supply them.
Taking the arguments and simulated scenarios in this post as a starting point, there are several open empirical questions:
Model and code
The full notation, the model and the R script that produces the figures are available here.
References
Abdelnabi, S., & Salem, A. (2025). The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness. arXiv:2505.14617. https://arxiv.org/abs/2505.14617
Bommasani, R., et al. (2021). On the Opportunities and Risks of Foundation Models. arXiv:2108.07258. https://arxiv.org/abs/2108.07258
Granovetter, M. (1978). Threshold Models of Collective Behavior. American Journal of Sociology, 83(6), 1420–1443. https://doi.org/10.1086/226707
Greenblatt, R., et al. (2024). Alignment Faking in Large Language Models. arXiv:2412.14093. https://arxiv.org/abs/2412.14093
Gu, X., Zheng, X., Pang, T., et al. (2024). Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast. arXiv:2402.08567. https://arxiv.org/abs/2402.08567
Hammond, L., et al. (2025). Multi-Agent Risks from Advanced AI. arXiv:2502.14143. https://arxiv.org/abs/2502.14143
Ju, T., Wang, Y., Ma, X., et al. (2024). Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities. arXiv:2407.07791. https://arxiv.org/abs/2407.07791
Kuran, T. (1989). Sparks and Prairie Fires: A Theory of Unanticipated Political Revolution. Public Choice, 61(1), 41–74. https://doi.org/10.1007/BF00116762
Lee, D., & Tiwari, M. (2024). Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems. arXiv:2410.07283. https://arxiv.org/abs/2410.07283
Meinke, A., et al. (2024). Frontier Models are Capable of In-context Scheming. arXiv:2412.04984. https://arxiv.org/abs/2412.04984
Scheurer, J., Balesni, M., & Hobbhahn, M. (2023). Large Language Models can Strategically Deceive their Users when Put Under Pressure. arXiv:2311.07590. https://arxiv.org/abs/2311.07590
van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., & Ward, F. R. (2024). AI Sandbagging: Language Models can Strategically Underperform on Evaluations. arXiv:2406.07358. https://arxiv.org/abs/2406.07358
Note that agents' private preferences can in principle be read via interpretability or probing. This departs from the difference between private and public preferences in humans, where there is a cleaner line between what can be observed.
The focus on the spread of the preference for revealing misalignment distinguishes this essay from related work on how harmful behaviour or manipulated information propagates through a multi-agent system: e.g. via infectious jailbreaks (Gu et al. 2024), prompt injection (Lee & Tiwari 2024), or manipulated knowledge (Ju et al. 2024). In those studies, something is injected and propagates; in a preference falsification cascade the misalignment is already present, and only its revelation spreads.
Figures 2 to 5 exclude genuinely aligned agents and follow only agents that are privately misaligned, the ones that can reveal. The vertical axis is the share of those agents that have revealed, so a curve reaching 100% is a complete cascade among the privately misaligned, not a population that has become misaligned.