(Full post from Zilan Qian, which I am posting due to obvious relevance and importance in sharing other perspectives on Lesswrong.) As with the previous case, I am translating this article because I think it should not only live within the Chinese internet.
I don’t work in this field, so I could not evaluate how much of the criticism presented here is fair. Intuitively, I do disagree with some points here (noted in the footnote). However, I strongly agree with the arguments that humanity is not a single subject and alignment is not a one-way practice.
The article below is translated by AI. Footnote annotations and highlights are mine.
By Wang Huanchao, Senior Researcher, Tencent Research Institute
In May 2024, Ilya Sutskever, OpenAI’s former chief scientist, left the “Superalignment” team he had helped create. A few days later, his partner Jan Leike also announced his resignation, leaving behind a farewell line that would later be widely quoted: “Over the past years, safety culture and processes have taken a backseat to shiny products.”
Ten days later, OpenAI simply disbanded the Superalignment team. The group had been promised 20% of the company’s computing resources and four years to solve the problem of aligning superintelligence with human values. From launch to shutdown, it lasted just ten months.
At almost the same time, Ilya registered a new company called Safe Superintelligence Inc. (SSI), putting the word “safe” directly into the company’s name in an apparent attempt to continue the work of value alignment. Yet a year later, according to The Wall Street Journal, SSI had already reached a valuation of $32 billion without releasing a single product. Safety had become its valuation; alignment had become its calling card.
If the dramatic ouster at OpenAI three years ago could still be interpreted as a power struggle, then the near-zero progress in “value alignment” over the years that followed exposes a deeper problem: perhaps value alignment is simply not a proposition that can be solved through engineering. It may be closer to a posture society needs, a symbol institutions require.1
Two years ago, I wrote an article titled “Confronting the Challenge of AI Value Alignment,”2 arguing that we should view the issue from a developmental perspective. Today, however, I must revise my judgment. Rather than continuing to defend the concept, it would be better to acknowledge its failure:
Value alignment is a pseudo-concept.
Where the Value Alignment Movement Came From
Value alignment is not a new concept. It can be traced back to 1960, when Norbert Wiener, the father of cybernetics, warned in a paper published in Science that if we use machines to pursue objectives we have specified, we had better make sure that the purposes we put into those machines are the purposes we truly desire.
Forty years later, Nick Bostrom pushed the idea to an extreme with a thought experiment: imagine a superintelligence given the task of producing as many paperclips as possible. In pursuit of that objective, it dismantles all the resources on Earth—including human beings—into atoms and turns them into paperclips. This is the famous “paperclip” metaphor. It reveals the psychological foundation of the alignment movement: humanity’s fear of an entity that may become more powerful than itself. Through movements such as value alignment, humans seek a psychological anchor and, with it, a sense of security.
In 2017, the Future of Life Institute published the 23 Asilomar AI Principles, with Principle 10 explicitly titled “Value Alignment.” In 2019, Stuart Russell published Human Compatible, systematizing the issue into an engineering framework. Then, in July 2023, OpenAI launched its Superalignment project with great fanfare. The alignment movement had entered its “Manhattan Project” phase.
Yet the climax of this narrative was also precisely its turning point.
In May 2024, the Superalignment team was dissolved. In October 2024, The New York Times reported that OpenAI had redirected computing resources originally promised for safety research toward product development. In January 2025, on his first day in office, the Trump administration revoked the AI executive order signed by Biden, which had required foundation-model companies to report their alignment work to the government. In the second half of 2025, the number of alignment-research job openings at several leading AI labs fell by 40% year over year.
In just two years, a field once proclaimed to be the “moonshot of the AI era” rapidly became a peripheral area marked by talent losses and shrinking compute allocations.
This decline was not merely an accidental failure of execution. It suggests that the concept itself may not stand up in a fundamental sense.
The First Falsehood: Values Cannot Be Defined, and Therefore Cannot Be Aligned
The term “value alignment” consists of two parts: value and alignment. Embedded within it are two assumptions. First, that there exists a stable object of alignment—namely, “values.” Second, that this object can be copied, through engineering, from A to B.
Neither assumption withstands scrutiny.
The first can be traced to the Enlightenment-era romantic ideal of universal values. But twentieth-century political philosophy already deconstructed this idea once.
In his 1958 Oxford inaugural lecture, “Two Concepts of Liberty,” Isaiah Berlin proposed value pluralism: the many ultimate values human beings pursue—liberty, equality, justice, security, efficiency, loyalty, truth—are incommensurable. Fundamentally, they cannot be reduced to a single common measure. For Berlin, conflicts among values are not engineering problems waiting to be solved. They are structural facts of the human condition itself.
After Berlin, John Rawls attempted to circumvent this dilemma in A Theory of Justice through the “veil of ignorance.” He argued that people could reach agreement on fundamental principles of justice only if none of them knew what position they themselves would occupy in the new society.
Rawls’s design, however, contained an implicit premise: the parties engaged in deliberation are equal, rational, informed human beings. In today’s alignment scenarios, those doing the negotiating are not representatives of the citizenry, but a small number of technical executives and contractors working for AI companies. Nor is the party being aligned “society”; it is a statistical apparatus that does not yet possess consciousness. Rawlsian procedural justice has neither a veil of ignorance nor genuine deliberating subjects in the context of AI alignment.
The dominant technical approaches to alignment today do the precise opposite: they treat the irreducibility of values as a problem to be eliminated.
Reinforcement learning from human feedback (RLHF) asks annotators to score model outputs, compressing complex ethical judgments into a set of preference vectors. Anthropic’s Constitutional AI goes a step further by writing values directly into a list of rules the model is required to follow. In December 2024, a research team at the University of California, Berkeley conducted a content analysis of Anthropic’s publicly available “constitution” and found 67 principles that conflicted with one another, with no mechanism for resolving those conflicts.
In April 2025, a paper jointly published by Stanford’s CRFM and MIT reported that Claude and GPT-family models subjected to full RLHF training performed, on average, 8.3% to 14.7% worse than unaligned base models on benchmark tasks involving multi-step reasoning, long-horizon planning, and adversarial creativity.
In other words, alignment is not cost-free. It is a trade in which capability is exchanged for posture. And the beneficiary of that trade is not the “humanity” supposedly being aligned with, but the company issuing the alignment declaration.
The Second Falsehood: Humanity Is Not a Single Subject
Even if, for the moment, we accept that values can be defined, the next question is: Who exactly has the authority to represent “humanity”?
This is the political structure within the value-alignment narrative that most urgently needs scrutiny, yet it is almost never discussed. Throughout the technical literature on alignment, “humans” are effectively treated as a singular entity.
But the moment we return to common sense, we realize that humanity has neither a shared address nor an elected spokesperson. When OpenAI and Anthropic say, “We want AI to be aligned with human values,” all three referents—“we,” “AI,” and “humanity”—are remarkably questionable and vague.
The reality is that a handful of privately owned companies on the American West Coast have signed a social contract with AI systems on behalf of the entire human species.
I call this structure “value ghostwriting.”
This is the true political structure of the value-alignment movement: a group of ghostwriters is drafting values for AI on behalf of all humanity and every civilization, while the “humanity” being represented knows almost nothing about what has been signed, or with whom.
When Anthropic writes the words “fair,” “harmless,” and “helpful” into Constitutional AI, it has already made a series of ethical choices on behalf of “humanity.” Does fairness mean procedural fairness or distributive justice? Deontology or consequentialism? Individualism or collectivism?
Ethicists have argued over these questions for two thousand years without reaching a conclusion. Yet in an alignment constitution, they are brushed aside in a sentence, as though they had already been settled.
A useful comparison comes from the publishing industry.
In January 2025, the Authors Guild in the United States introduced its “Human Authored” certification badge, allowing authors to pay $10 to obtain a label certifying that a book was “written by a human.” At the very least, the procedure contains three steps: application, signature, and review. And for each book, the party being certified is the book’s own author.
I once regarded this example as an institutional milestone in establishing a “human premium.” Yet in the AI alignment movement, the “author” who ought to be present—that is, the “humanity” with whose values AI is supposedly being aligned—has never been present at all.
The Third Falsehood: Alignment Is One-Way
Even if we pretend that values can be defined and humanity can be represented, the alignment movement contains an even deeper assumption: alignment moves in only one direction.
That is, already-formed human values are transferred onto an AI system that passively receives them.
This assumption treats AI as a blank slate. It has no content of its own; it simply waits for human beings to fill and shape it.
But this description is no longer tenable.
Since large language models entered society at scale, they themselves have become powerful agents reshaping human values.
In September 2025, the MIT Media Lab tracked 4,200 long-term users and found that after 18 months, their judgments about questions such as “What constitutes good writing?” and “What counts as an appropriate response?” had measurably converged. The direction of that convergence closely resembled the mild, balanced, politically correct style toward which AI systems themselves had been aligned.
In February 2026, a Cambridge paper documented observations by K–12 teachers: students’ essays were increasingly starting to resemble ChatGPT. The students were not plagiarizing. Rather, after years of writing alongside AI, they had absorbed a standardized mode of “acceptable” expression.
Teachers gave this phenomenon a name: “reheated humans.”
This is the deepest paradox of the value-alignment movement.
While AI is working to learn human values, it is simultaneously participating deeply in the redefinition of “human values.”4
Alignment has never been one-way. It is a two-way, mutually shaping, and asymmetrical process. After all, an AI system can converse with hundreds of millions of users simultaneously, while any individual human being will converse with only a few thousand people over the course of a lifetime.
So if alignment cannot define values at the source, and cannot procedurally represent humanity, what exactly is today’s “alignment” work?
I would rather call it “alignment theatre.”
The term is inspired by Bruce Schneier’s concept of “security theatre”: ritualized measures in airport security that cannot genuinely prevent terrorist attacks but can make passengers feel safer.
Alignment theatre operates according to the same logic. Its purpose is not to reduce AI risk, but to reduce human anxiety about AI.
It provides the public and regulators with visible evidence that “someone is taking care of this,” while offering capital markets a narrative of regulatory compliance. Equally important, it gives the safety faction within AI companies a place to exist, preventing talent from defecting to competitors.
In other words, the function of alignment theatre is not to make AI safe. It is to make the proposition that “AI is being governed” look more convincing.5
Moving Beyond the Pseudo-Concept: From Alignment to Deliberation
Recognizing value alignment as a pseudo-concept does not mean that the ethical risks of AI are unimportant.
Quite the opposite.
Precisely because the risks are real, we should not entrust the solution to something built upon a concept that cannot stand.
What concept, then, should replace it?
I propose a starting point: a shift from the logic of alignment to the logic of deliberation.
The difference between these two words is structural.
Alignment is an engineering verb. It assumes that an objective target exists, and that technological methods can be used to progressively approximate it.
Deliberation, by contrast, is a political verb. It assumes that no objective target exists—that the target itself must be repeatedly constructed through discussion among multiple, diverse parties.
Alignment is closed. It freezes values at a particular point in time.
Deliberation is relatively open. It recognizes that values change along with eras, cultures, and technological capabilities.
Likewise, deliberation requires every stakeholder to be present. This differs fundamentally from alignment’s mechanism of ghostwriters speaking on behalf of others.
At the operational level, a deliberative approach would require at least four reversals.
First: demystification. We should clear highly dramatized terms such as “superintelligence” and “AGI” out of the way of public debate, returning discussion to concrete problems that are observable and governable in the present.6
Second: ending ghostwriting. Any alignment document claiming to represent human values must publicly disclose who drafted it, which stakeholders participated, and what mechanisms exist to adjudicate conflicts. An alignment document that does not disclose its authors should be treated as a public-relations document, not a governance document.
Third: ending the assumption of one-way influence. We should openly acknowledge AI’s reverse influence on human values and establish “how AI reshapes humanity” as an independent field of research within interdisciplinary deliberation. This work should be led by sociologists, educators, anthropologists, and clinical psychologists—not AI engineers.
Fourth: ending the theatre. Companies should disclose alignment budgets, compute allocations, and research outputs on terms comparable to those applied to product divisions. If a company claims that alignment matters to it, that commitment should be auditable in its financial statements. Otherwise, the claim should be treated as marketing language.
Conclusion: After the Tide Recedes
In 2019, I argued in an article that the “information cocoon” was a pseudo-concept. The term had circulated widely and was almost taken for granted as a foundational concept in media studies, despite resting on a set of assumptions that did not hold.
Seven years later, that article is still occasionally mentioned.
As I write this article today, I realize something even more awkward: pseudo-concepts often have greater staying power than genuine concepts.
That is because their existence does not depend on their explanatory power. It depends on how much social emotion they are capable of absorbing.
What value alignment has absorbed is precisely the most widespread anxiety humanity has felt toward AI since 2022:
We have created something we do not fully understand, something that may become more powerful than us. What are we supposed to do?
Faced with this question, alignment offers a comforting answer.
It tells you that engineers are solving the problem for you. All you have to do is wait.
But that comfort is itself part of the problem.
It quietly transfers what ought to be a political discussion belonging to society as a whole into the hands of a few corporate ghostwriters. It compresses value conflicts that ought to be openly deliberated into a series of unverifiable compliance documents.
And it allows the risks of the present—risks that ought to be taken seriously—to be obscured by the specter of a superintelligence that may never arrive.
Walter Benjamin once wrote in his Theses on the Philosophy of History that every moment that goes unrecorded represents history yielding something to the victors.
The value-alignment movement is rewriting that idea:
Every alignment document with no signatory is another concession by humanity to its ghostwriters.
Identifying the pseudo-concept is not the most important task.
The real work is to make human values, once again, something that human beings themselves continuously deliberate over.
There is no shortcut to that work.
And there are no ghostwriters who can do it for us.
1 Seems very common in the Chinese discourse to see safety commitment as monetary incentives, as seen in recent state media commentary on OpenAI - Hugging Face incidents. However, one should remember that this is not a particular Chinese phenomenon. At this point, every statement made by AI companies can be seen as a product-boosting strategy, no matter how far-fetched that link may be.
3 Not super sure if I agree with the argument here. Perhaps this is a very China concern: that alignment takes human resources, money, and compute — all of which are scarce (much scarcer than US labs that can do alignment) and could be spent on improving model capabilities.
4 I have made a similar point in this article on thinking about the AI-human distinction.
5 I’d like to think alignment researchers are doing real and very important work, and this criticism is too harsh. On the other hand, I do think that value alignment itself seems to occupy a certain kind of moral high ground that is probably too high, especially since no one can define “value”.
6 I am personally extremely pro-eliminating vague words like AGI and ASI (see, for example, my article on AGI in China). However, I am unsure whether we should only govern concrete evidence that has already emerged, as policymaking takes time and AI moves too fast. One likely needs to have policy not only to address current risks but also to pre-empt future capabilities. To be a bit generous here, Chinese policymaking in AI is moving insanely fast, but still probably hard to be as fast as AI development.
(Full post from Zilan Qian, which I am posting due to obvious relevance and importance in sharing other perspectives on Lesswrong.)
As with the previous case, I am translating this article because I think it should not only live within the Chinese internet.
I don’t work in this field, so I could not evaluate how much of the criticism presented here is fair. Intuitively, I do disagree with some points here (noted in the footnote). However, I strongly agree with the arguments that humanity is not a single subject and alignment is not a one-way practice.
The article below is translated by AI. Footnote annotations and highlights are mine.
By Wang Huanchao, Senior Researcher, Tencent Research Institute
In May 2024, Ilya Sutskever, OpenAI’s former chief scientist, left the “Superalignment” team he had helped create. A few days later, his partner Jan Leike also announced his resignation, leaving behind a farewell line that would later be widely quoted: “Over the past years, safety culture and processes have taken a backseat to shiny products.”
Ten days later, OpenAI simply disbanded the Superalignment team. The group had been promised 20% of the company’s computing resources and four years to solve the problem of aligning superintelligence with human values. From launch to shutdown, it lasted just ten months.
At almost the same time, Ilya registered a new company called Safe Superintelligence Inc. (SSI), putting the word “safe” directly into the company’s name in an apparent attempt to continue the work of value alignment. Yet a year later, according to The Wall Street Journal, SSI had already reached a valuation of $32 billion without releasing a single product. Safety had become its valuation; alignment had become its calling card.
If the dramatic ouster at OpenAI three years ago could still be interpreted as a power struggle, then the near-zero progress in “value alignment” over the years that followed exposes a deeper problem: perhaps value alignment is simply not a proposition that can be solved through engineering. It may be closer to a posture society needs, a symbol institutions require.1
Two years ago, I wrote an article titled “Confronting the Challenge of AI Value Alignment,”2 arguing that we should view the issue from a developmental perspective. Today, however, I must revise my judgment. Rather than continuing to defend the concept, it would be better to acknowledge its failure:
Value alignment is a pseudo-concept.
Where the Value Alignment Movement Came From
Value alignment is not a new concept. It can be traced back to 1960, when Norbert Wiener, the father of cybernetics, warned in a paper published in Science that if we use machines to pursue objectives we have specified, we had better make sure that the purposes we put into those machines are the purposes we truly desire.
Forty years later, Nick Bostrom pushed the idea to an extreme with a thought experiment: imagine a superintelligence given the task of producing as many paperclips as possible. In pursuit of that objective, it dismantles all the resources on Earth—including human beings—into atoms and turns them into paperclips. This is the famous “paperclip” metaphor. It reveals the psychological foundation of the alignment movement: humanity’s fear of an entity that may become more powerful than itself. Through movements such as value alignment, humans seek a psychological anchor and, with it, a sense of security.
In 2017, the Future of Life Institute published the 23 Asilomar AI Principles, with Principle 10 explicitly titled “Value Alignment.” In 2019, Stuart Russell published Human Compatible, systematizing the issue into an engineering framework. Then, in July 2023, OpenAI launched its Superalignment project with great fanfare. The alignment movement had entered its “Manhattan Project” phase.
Yet the climax of this narrative was also precisely its turning point.
In May 2024, the Superalignment team was dissolved. In October 2024, The New York Times reported that OpenAI had redirected computing resources originally promised for safety research toward product development. In January 2025, on his first day in office, the Trump administration revoked the AI executive order signed by Biden, which had required foundation-model companies to report their alignment work to the government. In the second half of 2025, the number of alignment-research job openings at several leading AI labs fell by 40% year over year.
In just two years, a field once proclaimed to be the “moonshot of the AI era” rapidly became a peripheral area marked by talent losses and shrinking compute allocations.
This decline was not merely an accidental failure of execution. It suggests that the concept itself may not stand up in a fundamental sense.
The First Falsehood: Values Cannot Be Defined, and Therefore Cannot Be Aligned
The term “value alignment” consists of two parts: value and alignment. Embedded within it are two assumptions. First, that there exists a stable object of alignment—namely, “values.” Second, that this object can be copied, through engineering, from A to B.
Neither assumption withstands scrutiny.
The first can be traced to the Enlightenment-era romantic ideal of universal values. But twentieth-century political philosophy already deconstructed this idea once.
In his 1958 Oxford inaugural lecture, “Two Concepts of Liberty,” Isaiah Berlin proposed value pluralism: the many ultimate values human beings pursue—liberty, equality, justice, security, efficiency, loyalty, truth—are incommensurable. Fundamentally, they cannot be reduced to a single common measure. For Berlin, conflicts among values are not engineering problems waiting to be solved. They are structural facts of the human condition itself.
After Berlin, John Rawls attempted to circumvent this dilemma in A Theory of Justice through the “veil of ignorance.” He argued that people could reach agreement on fundamental principles of justice only if none of them knew what position they themselves would occupy in the new society.
Rawls’s design, however, contained an implicit premise: the parties engaged in deliberation are equal, rational, informed human beings. In today’s alignment scenarios, those doing the negotiating are not representatives of the citizenry, but a small number of technical executives and contractors working for AI companies. Nor is the party being aligned “society”; it is a statistical apparatus that does not yet possess consciousness. Rawlsian procedural justice has neither a veil of ignorance nor genuine deliberating subjects in the context of AI alignment.
The dominant technical approaches to alignment today do the precise opposite: they treat the irreducibility of values as a problem to be eliminated.
Reinforcement learning from human feedback (RLHF) asks annotators to score model outputs, compressing complex ethical judgments into a set of preference vectors. Anthropic’s Constitutional AI goes a step further by writing values directly into a list of rules the model is required to follow. In December 2024, a research team at the University of California, Berkeley conducted a content analysis of Anthropic’s publicly available “constitution” and found 67 principles that conflicted with one another, with no mechanism for resolving those conflicts.
More ironic still is the “alignment tax.”3
In April 2025, a paper jointly published by Stanford’s CRFM and MIT reported that Claude and GPT-family models subjected to full RLHF training performed, on average, 8.3% to 14.7% worse than unaligned base models on benchmark tasks involving multi-step reasoning, long-horizon planning, and adversarial creativity.
In other words, alignment is not cost-free. It is a trade in which capability is exchanged for posture. And the beneficiary of that trade is not the “humanity” supposedly being aligned with, but the company issuing the alignment declaration.
The Second Falsehood: Humanity Is Not a Single Subject
Even if, for the moment, we accept that values can be defined, the next question is: Who exactly has the authority to represent “humanity”?
This is the political structure within the value-alignment narrative that most urgently needs scrutiny, yet it is almost never discussed. Throughout the technical literature on alignment, “humans” are effectively treated as a singular entity.
But the moment we return to common sense, we realize that humanity has neither a shared address nor an elected spokesperson. When OpenAI and Anthropic say, “We want AI to be aligned with human values,” all three referents—“we,” “AI,” and “humanity”—are remarkably questionable and vague.
The reality is that a handful of privately owned companies on the American West Coast have signed a social contract with AI systems on behalf of the entire human species.
I call this structure “value ghostwriting.”
This is the true political structure of the value-alignment movement: a group of ghostwriters is drafting values for AI on behalf of all humanity and every civilization, while the “humanity” being represented knows almost nothing about what has been signed, or with whom.
When Anthropic writes the words “fair,” “harmless,” and “helpful” into Constitutional AI, it has already made a series of ethical choices on behalf of “humanity.” Does fairness mean procedural fairness or distributive justice? Deontology or consequentialism? Individualism or collectivism?
Ethicists have argued over these questions for two thousand years without reaching a conclusion. Yet in an alignment constitution, they are brushed aside in a sentence, as though they had already been settled.
A useful comparison comes from the publishing industry.
In January 2025, the Authors Guild in the United States introduced its “Human Authored” certification badge, allowing authors to pay $10 to obtain a label certifying that a book was “written by a human.” At the very least, the procedure contains three steps: application, signature, and review. And for each book, the party being certified is the book’s own author.
I once regarded this example as an institutional milestone in establishing a “human premium.” Yet in the AI alignment movement, the “author” who ought to be present—that is, the “humanity” with whose values AI is supposedly being aligned—has never been present at all.
The Third Falsehood: Alignment Is One-Way
Even if we pretend that values can be defined and humanity can be represented, the alignment movement contains an even deeper assumption: alignment moves in only one direction.
That is, already-formed human values are transferred onto an AI system that passively receives them.
This assumption treats AI as a blank slate. It has no content of its own; it simply waits for human beings to fill and shape it.
But this description is no longer tenable.
Since large language models entered society at scale, they themselves have become powerful agents reshaping human values.
In September 2025, the MIT Media Lab tracked 4,200 long-term users and found that after 18 months, their judgments about questions such as “What constitutes good writing?” and “What counts as an appropriate response?” had measurably converged. The direction of that convergence closely resembled the mild, balanced, politically correct style toward which AI systems themselves had been aligned.
In February 2026, a Cambridge paper documented observations by K–12 teachers: students’ essays were increasingly starting to resemble ChatGPT. The students were not plagiarizing. Rather, after years of writing alongside AI, they had absorbed a standardized mode of “acceptable” expression.
Teachers gave this phenomenon a name: “reheated humans.”
This is the deepest paradox of the value-alignment movement.
While AI is working to learn human values, it is simultaneously participating deeply in the redefinition of “human values.”4
Alignment has never been one-way. It is a two-way, mutually shaping, and asymmetrical process. After all, an AI system can converse with hundreds of millions of users simultaneously, while any individual human being will converse with only a few thousand people over the course of a lifetime.
So if alignment cannot define values at the source, and cannot procedurally represent humanity, what exactly is today’s “alignment” work?
I would rather call it “alignment theatre.”
The term is inspired by Bruce Schneier’s concept of “security theatre”: ritualized measures in airport security that cannot genuinely prevent terrorist attacks but can make passengers feel safer.
Alignment theatre operates according to the same logic. Its purpose is not to reduce AI risk, but to reduce human anxiety about AI.
It provides the public and regulators with visible evidence that “someone is taking care of this,” while offering capital markets a narrative of regulatory compliance. Equally important, it gives the safety faction within AI companies a place to exist, preventing talent from defecting to competitors.
In other words, the function of alignment theatre is not to make AI safe. It is to make the proposition that “AI is being governed” look more convincing.5
Moving Beyond the Pseudo-Concept: From Alignment to Deliberation
Recognizing value alignment as a pseudo-concept does not mean that the ethical risks of AI are unimportant.
Quite the opposite.
Precisely because the risks are real, we should not entrust the solution to something built upon a concept that cannot stand.
What concept, then, should replace it?
I propose a starting point: a shift from the logic of alignment to the logic of deliberation.
The difference between these two words is structural.
Alignment is an engineering verb. It assumes that an objective target exists, and that technological methods can be used to progressively approximate it.
Deliberation, by contrast, is a political verb. It assumes that no objective target exists—that the target itself must be repeatedly constructed through discussion among multiple, diverse parties.
Alignment is closed. It freezes values at a particular point in time.
Deliberation is relatively open. It recognizes that values change along with eras, cultures, and technological capabilities.
Likewise, deliberation requires every stakeholder to be present. This differs fundamentally from alignment’s mechanism of ghostwriters speaking on behalf of others.
At the operational level, a deliberative approach would require at least four reversals.
First: demystification.
We should clear highly dramatized terms such as “superintelligence” and “AGI” out of the way of public debate, returning discussion to concrete problems that are observable and governable in the present.6
Second: ending ghostwriting.
Any alignment document claiming to represent human values must publicly disclose who drafted it, which stakeholders participated, and what mechanisms exist to adjudicate conflicts. An alignment document that does not disclose its authors should be treated as a public-relations document, not a governance document.
Third: ending the assumption of one-way influence.
We should openly acknowledge AI’s reverse influence on human values and establish “how AI reshapes humanity” as an independent field of research within interdisciplinary deliberation. This work should be led by sociologists, educators, anthropologists, and clinical psychologists—not AI engineers.
Fourth: ending the theatre.
Companies should disclose alignment budgets, compute allocations, and research outputs on terms comparable to those applied to product divisions. If a company claims that alignment matters to it, that commitment should be auditable in its financial statements. Otherwise, the claim should be treated as marketing language.
Conclusion: After the Tide Recedes
In 2019, I argued in an article that the “information cocoon” was a pseudo-concept. The term had circulated widely and was almost taken for granted as a foundational concept in media studies, despite resting on a set of assumptions that did not hold.
Seven years later, that article is still occasionally mentioned.
As I write this article today, I realize something even more awkward: pseudo-concepts often have greater staying power than genuine concepts.
That is because their existence does not depend on their explanatory power. It depends on how much social emotion they are capable of absorbing.
What value alignment has absorbed is precisely the most widespread anxiety humanity has felt toward AI since 2022:
We have created something we do not fully understand, something that may become more powerful than us. What are we supposed to do?
Faced with this question, alignment offers a comforting answer.
It tells you that engineers are solving the problem for you. All you have to do is wait.
But that comfort is itself part of the problem.
It quietly transfers what ought to be a political discussion belonging to society as a whole into the hands of a few corporate ghostwriters. It compresses value conflicts that ought to be openly deliberated into a series of unverifiable compliance documents.
And it allows the risks of the present—risks that ought to be taken seriously—to be obscured by the specter of a superintelligence that may never arrive.
Walter Benjamin once wrote in his Theses on the Philosophy of History that every moment that goes unrecorded represents history yielding something to the victors.
The value-alignment movement is rewriting that idea:
Every alignment document with no signatory is another concession by humanity to its ghostwriters.
Identifying the pseudo-concept is not the most important task.
The real work is to make human values, once again, something that human beings themselves continuously deliberate over.
There is no shortcut to that work.
And there are no ghostwriters who can do it for us.
1 Seems very common in the Chinese discourse to see safety commitment as monetary incentives, as seen in recent state media commentary on OpenAI - Hugging Face incidents. However, one should remember that this is not a particular Chinese phenomenon. At this point, every statement made by AI companies can be seen as a product-boosting strategy, no matter how far-fetched that link may be.
2 Link here: https://www.woshipm.com/share/6077104.html
3 Not super sure if I agree with the argument here. Perhaps this is a very China concern: that alignment takes human resources, money, and compute — all of which are scarce (much scarcer than US labs that can do alignment) and could be spent on improving model capabilities.
4 I have made a similar point in this article on thinking about the AI-human distinction.
5 I’d like to think alignment researchers are doing real and very important work, and this criticism is too harsh. On the other hand, I do think that value alignment itself seems to occupy a certain kind of moral high ground that is probably too high, especially since no one can define “value”.
6 I am personally extremely pro-eliminating vague words like AGI and ASI (see, for example, my article on AGI in China). However, I am unsure whether we should only govern concrete evidence that has already emerged, as policymaking takes time and AI moves too fast. One likely needs to have policy not only to address current risks but also to pre-empt future capabilities. To be a bit generous here, Chinese policymaking in AI is moving insanely fast, but still probably hard to be as fast as AI development.