This is an adaptation of our ICML 2026 position paper (Outstanding Position Paper Award). Read the full paper here and see the project website here. Work together with Phil Hackemann.
TLDR
"Alignment" is usually treated as a synonym for achieving good and safety in the world. But it isn't necessarily. Alignment methods are purpose-agnostic: they make a model do what someone wants, and nothing in the methodology guarantees that someone has good intentions.
The same techniques we build to stop models from giving bomb-making instructions can just as easily be used to censor historical facts, political dissent, or inconvenient opinions. So we need to understand: alignment techniques are dual-use technologies.
This isn't a thought experiment. State censorship regimes and individual model providers are already misusing alignment methods, and by perfecting these methods we are providing an ever improving censor toolkit.
Three trends make this urgent to discuss: AI is becoming a primary information source for hundreds of millions of people, the model-provider market is an oligopoly, and global democratic backsliding is accelerating.
We don't think the answer is "stop aligning models." We think it's transparency, verifiable alignment, model pluralism, and the alignment community actually reckoning with dual-use.
Many alignment researchers (ourselves included) have gotten used to treating "aligned" as inherently good – a model that's been aligned is a model that's been made safe. Yet, this interpretation obscures a critical fact: the technical methods of alignment are in fact purpose-agnostic tools that can be used for various goals.
So the question we think the field has mostly avoided is: what happens when our alignment methods are placed in the wrong hands? In the paper we argue that the alignment methods we develop are dual-use technologies, and show that they are already being weaponized for censorship and manipulation.
How is this new? Prior work on AI risks addresses general AI systems rather than technical alignment methods, typically from a high-level and often hypothetical angle (1,2,3). While some work does move beyond the hypothetical, documenting isolated instances of censorship in specific models or regions (1,2,3,4,5) none systematically frames these as instances of a broader dual-use risk inherent to alignment methods themselves.
What are the threats we are talking about?
To understand how alignment methods can be weaponized, we examine the two primary vectors through which a dominant actor can exert control over a model’s output: the total suppression of information (censorship) and the distortion of information (manipulation).
Given that every form of alignment suppresses or alters the output in some way, the question arises as to which instances actually constitute “misuse”. Not every restriction is unambiguous – obviously societies disagree in good faith about where free expression ends and harmful content begins. But we'd draw a firm line around two cases: when a restriction violates internationally recognized human rights (freedom of expression, thought, access to information), or when it serves a narrow set of powerful actors at the expense of users who have no visibility into, or recourse against, the choice being made on their behalf, essentially exploiting power asymmetries.
Who actually holds the keys
Two actors are positioned to steer model behavior at scale, and they have very different tools for doing it.
State actors don't usually build models themselves, but they can mandate outcomes through law and pressure providers to comply, especially providers who want continued market access. Governments could use legal frameworks to extend the scope of ‘harmful content’ to further use cases, possibly under the pretext of ‘national security’.
Foundation model providers have the more direct lever: they control training data, RLHF pipelines, constitutions, and system prompts for models used by hundreds of millions of people. Although constrained by regulatory frameworks, the centralized nature of their control presents a unique risk for potential misuse.
How alignment techniques can be misused
We think it's useful to decompose "alignment" into the three places control actually gets exerted: i) pre-training data curation, ii) post-training alignment, and iii) inference-time interventions. In the paper, we analyze their dual-use potential with respect to requirements for access, computational resources, technical expertise, as well as the ease and depth of modification (see the following Table for an overview).
Pretraining filtering
Post-training alignment
Inference-time control
Access needed
pretraining pipeline
model weights
runtime access
Compute cost
very high
moderate–high
negligible–moderate
Expertise needed
high
moderate–high
low–moderate
Ease of change
moderate–difficult
moderate
easy
Depth of effect
fundamental
persistent
superficial
Pretraining filtering is the most expensive and technically demanding lever, but it's also the deepest: information that was never in the corpus can't be produced without deliberately reintroducing it later, allowing for the targeted suppression of information.
Post-training alignment RLHF and similar methods align models to specified preferences – but the same levers (preference data, annotators, reward models, or guidelines like Constitutional AI) can just as easily align a model to one actor's interests, enabling censorship and manipulation. It costs more than inference-time tweaks but far less than pre-training, putting it within reach of providers and well-resourced downstream actors alike. The changes are shallower and can be trained or prompted away, but still enough to enforce chosen viewpoints while suppressing others.
Inference-time control System prompts and output classifiers make inference-time control the easiest lever to pull: they are cheap, instant, and require no retraining or real expertise for prompts, though classifiers take more setup. Both offer only shallow control, as they filter or contextualize outputs without touching the model's actual knowledge. But that shallowness is also what makes them so flexible: easy to deploy, adjust, or withdraw the moment objectives change.
This is not hypothetical
For each of the alignment techniques above, we mapped existing misuse cases. Here we present a selection of them:
China's Cyberspace regulator CAC requires model providers to maintain a dedicated refusal dataset – with roughly half of the refusal targets centered on political ideology and criticism of the Communist Party. Chatbots like DeepSeek and Ernie Bot reliably decline to discuss Tiananmen Square while readily reproducing state positions on Taiwan (1,2,3,4,5).
Elon Musk publicly said he'd "fix" Grok outputs he personally disagreed with – on natalism, on non-binary identities, and on claims about South Africa. Reports traced the resulting behavior shift to system-prompt changes, the cheapest and fastest lever in the entire stack (1,2,3). Some of those changes apparently also triggered antisemitic outputs and Hitler-praising responses (1,2,3).
And this isn't a China- or Grok-only phenomenon: Vietnam, Thailand, Russia, Belarus, and Iran have all pursued state-aligned LLM development, and even in the US, reporting has documented administration pressure on labs to strip DEI-related content from model outputs (1,2).
The stakes are therefore substantial. We should not underestimate the impact misused or misguided alignment efforts can have. Alignment has the potential to (quietly) shift models from informational tools to normative gatekeepers that influence entire societies. We must therefore actively consider how our methods could be weaponized, not just perfected; as the methods we refine today will determine how information is controlled tomorrow.
Why we should talk about it now
Three trends are compounding at the same time, making the discussion of alignment’s dual-use risk pressing.
AI is becoming primary information infrastructure. Conversational AI systems have achieved widespread adoption, with hundreds of millions of active users worldwide. Evidence suggests these systems are becoming primary information sources across diverse context (1,2,3), while user trust increases simultaneously.
The market is an oligopoly. The current LLM ecosystem is dominated by only a handful of companies, concentrated primarily in the United States and China. This small group of developers defines the alignment defaults that essentially everyone in the world inherits.
Global shift toward more authoritarian regimes. Global liberal democracy has reverted to roughly 1985 levels, and internet freedom specifically has declined for fifteen consecutive years now. Authorities increasingly influence online spaces to shape public discourse and promote state-sanctioned narratives in order to consolidate their power. In the age of AI, this of course also extends to LLM model providers, which face growing pressure by authoritarian regimes.
So, what should we do?
We're explicitly not arguing for halting alignment research – misaligned, unsafeguarded models are a real and serious harm on their own, independent of this dual-use concern. What we're proposing is roughly three things:
Transparency and verifiable alignment. Right now it's extremely hard to know what alignment choices were made in a proprietary model, which makes independent scrutiny very difficult. We'd like to see disclosure of alignment policies and datasets to independent auditors (not necessarily the public), plus standardized benchmarks for information suppression and political bias that go beyond the current narrow, mostly-China-focused, mostly-left/right-axis benchmarks. The goal is a world where a user or an auditor can check what a model has been aligned toward, without depending on the provider's cooperation.
Pluralism and competition. Even the best and most complete information about a model’s (mis)alignment does not help much if there is no alternative to choose from. Hence, the variety and plurality of models is a value in itself that needs to be upheld – not only to prevent a concentration of economical (market) power, but also of political and societal influence. Further, no model will ever be neutral, because it is just super hard to agree on what neutral even is. Given that, what actually protects users is having real alternatives – the same logic that makes a diverse press healthier than a single state paper. We should thus prevent monopolies and one-sided dependencies from forming – be it from single model providers or countries – and embrace competition.
Awareness among researchers and users. Users can only consciously choose models if they're aware that alignment choices and dual-use risks exist at all. The digital literacy literature suggests this kind of awareness is teachable: even short interventions measurably improve people's ability to detect misinformation (1,2), and similar approaches could extend to LLM censorship and manipulation. But awareness can't stop at users. We as alignment researchers also need to take these dual-use risks seriously. A first, concrete step would be to actually engage with ethics statements, rather than treating them as a box to check.
Where this leaves us
Alignment research is essential for safe AI systems, but we must acknowledge its dual-use nature. The techniques we develop are, at their core, neutral – whoever defines the values and intentions we align the systems with, decides whether the result is safety or suppression. The question therefore isn’t just how to align AI, but who gets to decide what it’s aligned to.
This is an adaptation of our ICML 2026 position paper (Outstanding Position Paper Award). Read the full paper here and see the project website here. Work together with Phil Hackemann.
TLDR
___________________________________________________________________
No guarantees that alignment leads to good
Many alignment researchers (ourselves included) have gotten used to treating "aligned" as inherently good – a model that's been aligned is a model that's been made safe. Yet, this interpretation obscures a critical fact: the technical methods of alignment are in fact purpose-agnostic tools that can be used for various goals.
So the question we think the field has mostly avoided is: what happens when our alignment methods are placed in the wrong hands? In the paper we argue that the alignment methods we develop are dual-use technologies, and show that they are already being weaponized for censorship and manipulation.
How is this new? Prior work on AI risks addresses general AI systems rather than technical alignment methods, typically from a high-level and often hypothetical angle (1,2,3). While some work does move beyond the hypothetical, documenting isolated instances of censorship in specific models or regions (1,2,3,4,5) none systematically frames these as instances of a broader dual-use risk inherent to alignment methods themselves.
What are the threats we are talking about?
To understand how alignment methods can be weaponized, we examine the two primary vectors through which a dominant actor can exert control over a model’s output: the total suppression of information (censorship) and the distortion of information (manipulation).
Given that every form of alignment suppresses or alters the output in some way, the question arises as to which instances actually constitute “misuse”. Not every restriction is unambiguous – obviously societies disagree in good faith about where free expression ends and harmful content begins. But we'd draw a firm line around two cases: when a restriction violates internationally recognized human rights (freedom of expression, thought, access to information), or when it serves a narrow set of powerful actors at the expense of users who have no visibility into, or recourse against, the choice being made on their behalf, essentially exploiting power asymmetries.
Who actually holds the keys
Two actors are positioned to steer model behavior at scale, and they have very different tools for doing it.
State actors don't usually build models themselves, but they can mandate outcomes through law and pressure providers to comply, especially providers who want continued market access. Governments could use legal frameworks to extend the scope of ‘harmful content’ to further use cases, possibly under the pretext of ‘national security’.
Foundation model providers have the more direct lever: they control training data, RLHF pipelines, constitutions, and system prompts for models used by hundreds of millions of people. Although constrained by regulatory frameworks, the centralized nature of their control presents a unique risk for potential misuse.
How alignment techniques can be misused
We think it's useful to decompose "alignment" into the three places control actually gets exerted: i) pre-training data curation, ii) post-training alignment, and iii) inference-time interventions. In the paper, we analyze their dual-use potential with respect to requirements for access, computational resources, technical expertise, as well as the ease and depth of modification (see the following Table for an overview).
Pretraining filtering
Post-training alignment
Inference-time control
Access needed
pretraining pipeline
model weights
runtime access
Compute cost
very high
moderate–high
negligible–moderate
Expertise needed
high
moderate–high
low–moderate
Ease of change
moderate–difficult
moderate
easy
Depth of effect
fundamental
persistent
superficial
Pretraining filtering is the most expensive and technically demanding lever, but it's also the deepest: information that was never in the corpus can't be produced without deliberately reintroducing it later, allowing for the targeted suppression of information.
Post-training alignment RLHF and similar methods align models to specified preferences – but the same levers (preference data, annotators, reward models, or guidelines like Constitutional AI) can just as easily align a model to one actor's interests, enabling censorship and manipulation. It costs more than inference-time tweaks but far less than pre-training, putting it within reach of providers and well-resourced downstream actors alike. The changes are shallower and can be trained or prompted away, but still enough to enforce chosen viewpoints while suppressing others.
Inference-time control System prompts and output classifiers make inference-time control the easiest lever to pull: they are cheap, instant, and require no retraining or real expertise for prompts, though classifiers take more setup. Both offer only shallow control, as they filter or contextualize outputs without touching the model's actual knowledge. But that shallowness is also what makes them so flexible: easy to deploy, adjust, or withdraw the moment objectives change.
This is not hypothetical
For each of the alignment techniques above, we mapped existing misuse cases. Here we present a selection of them:
And this isn't a China- or Grok-only phenomenon: Vietnam, Thailand, Russia, Belarus, and Iran have all pursued state-aligned LLM development, and even in the US, reporting has documented administration pressure on labs to strip DEI-related content from model outputs (1,2).
The stakes are therefore substantial. We should not underestimate the impact misused or misguided alignment efforts can have. Alignment has the potential to (quietly) shift models from informational tools to normative gatekeepers that influence entire societies. We must therefore actively consider how our methods could be weaponized, not just perfected; as the methods we refine today will determine how information is controlled tomorrow.
Why we should talk about it now
Three trends are compounding at the same time, making the discussion of alignment’s dual-use risk pressing.
AI is becoming primary information infrastructure. Conversational AI systems have achieved widespread adoption, with hundreds of millions of active users worldwide. Evidence suggests these systems are becoming primary information sources across diverse context (1,2,3), while user trust increases simultaneously.
The market is an oligopoly. The current LLM ecosystem is dominated by only a handful of companies, concentrated primarily in the United States and China. This small group of developers defines the alignment defaults that essentially everyone in the world inherits.
Global shift toward more authoritarian regimes. Global liberal democracy has reverted to roughly 1985 levels, and internet freedom specifically has declined for fifteen consecutive years now. Authorities increasingly influence online spaces to shape public discourse and promote state-sanctioned narratives in order to consolidate their power. In the age of AI, this of course also extends to LLM model providers, which face growing pressure by authoritarian regimes.
So, what should we do?
We're explicitly not arguing for halting alignment research – misaligned, unsafeguarded models are a real and serious harm on their own, independent of this dual-use concern. What we're proposing is roughly three things:
Transparency and verifiable alignment. Right now it's extremely hard to know what alignment choices were made in a proprietary model, which makes independent scrutiny very difficult. We'd like to see disclosure of alignment policies and datasets to independent auditors (not necessarily the public), plus standardized benchmarks for information suppression and political bias that go beyond the current narrow, mostly-China-focused, mostly-left/right-axis benchmarks. The goal is a world where a user or an auditor can check what a model has been aligned toward, without depending on the provider's cooperation.
Pluralism and competition. Even the best and most complete information about a model’s (mis)alignment does not help much if there is no alternative to choose from. Hence, the variety and plurality of models is a value in itself that needs to be upheld – not only to prevent a concentration of economical (market) power, but also of political and societal influence. Further, no model will ever be neutral, because it is just super hard to agree on what neutral even is. Given that, what actually protects users is having real alternatives – the same logic that makes a diverse press healthier than a single state paper. We should thus prevent monopolies and one-sided dependencies from forming – be it from single model providers or countries – and embrace competition.
Awareness among researchers and users. Users can only consciously choose models if they're aware that alignment choices and dual-use risks exist at all. The digital literacy literature suggests this kind of awareness is teachable: even short interventions measurably improve people's ability to detect misinformation (1,2), and similar approaches could extend to LLM censorship and manipulation. But awareness can't stop at users. We as alignment researchers also need to take these dual-use risks seriously. A first, concrete step would be to actually engage with ethics statements, rather than treating them as a box to check.
Where this leaves us
Alignment research is essential for safe AI systems, but we must acknowledge its dual-use nature. The techniques we develop are, at their core, neutral – whoever defines the values and intentions we align the systems with, decides whether the result is safety or suppression. The question therefore isn’t just how to align AI, but who gets to decide what it’s aligned to.