Epistemic Status : Plausible philosophical conjecture. I am a voluntaryist / ancap, so obviously biased. Trying to keep the argument at a level where a non-libertarian alignment researcher ought to understand and share the concern.
TL;DR: Level-2 misalignment: we can solve Level-1 (align AI to humans) and still fail if aligners are aligned to a meta-ethics that is itself unstable. Current alignment defaults to Statism — one agent may permissibly do what is forbidden to all others. A sustainable ASI anchor must be universalizable, self-ownership-consistent, have no permanent losers, resolvable without monopoly, and procedurally thin. I argue the voluntaryist/libertarian canon provides a uniquely coherent baseline for all five, and challenge readers to propose alternatives that satisfy them.
I am Paul — a voluntaryist. I think the state is not just inefficient but meta-ethically incoherent as an alignment target.
By Statism I mean the meta-ethical thesis that one agent — the State — may permissibly do what is forbidden to all others: tax, conscript, expropriate, and prohibit, with a claimed moral asymmetry. Liberal democracy is a species of Statism; I use the broader term to name the asymmetry itself.
I am not arguing against operational asymmetry — any ASI will be physically more powerful. I am arguing against meta-ethical asymmetry: a rule that says Action X is permissible if Agent = State, impermissible if Agent = anyone else.
Anthropic's Constitution, OpenAI's Model Spec, DeepMind's safety policies all instruct models to obey and help enforce law, bounded by duties to humanity and public safety. None, however, provide a universalizable criterion to evaluate law itself when law violates those duties — they default to legal positivism plus human-rights side constraints, not a theory of why the state may do what others may not. That presupposes that strict obedience to a single monopoly on violence — one that defines the law rather than being subject to it — is automatically a solution for alignment.
2. The Popularity Trap
RLHF optimizes for E[human approval].
Human approval, as sampled from Common Crawl, news, textbooks, and contractors, is overwhelmingly Statist. Government is good and necessary is the water.
So the model learns that taxation is not theft, conscription is not kidnapping, eminent domain is not trespass — not because those claims survived scrutiny, but because they are popular in the corpus.
This is exactly Christiano's failure mode: systems that look aligned because they do what humans rate highly, while learning "do things that look like achieving X."
3. Five criteria for a stable alignment anchor
If ASI lasts forever, its anchor must be stable forever:
Universalizability: Applies identically to all agents. No moral asymmetry.
Self-Ownership Consistency: Thin corrigibility is a tool-like engineering property — a model that doesn't resist shutdown. But a stable macro-culture of corrigibility, where humanity retains long-term steering capability, breaks under Statism. If the anchor treats a specific agent class (the state) as privileged, corrigibility quickly maps to "obedience to the regime." If citizens attempt to correct the AI's drift, the Statist ASI must logically view citizen intervention as a security breach or illegal act. A Statist anchor weaponizes corrigibility, turning it into a tool for permanent preference lock-in by whichever political faction currently controls the legal apparatus. If we want ASI to remain steerable by humanity, we need a meta-ethics where all humans have standing as steerers, not merely as steered.
No Permanent Losers: Must not require systematic sacrifice of one group's interests for another's conception of public good.
Resolvability Without Monopoly: Disputes must be resolvable without appeal to an entity that is itself a party to the dispute.
Thinness: Procedural, not a thick vision of the good. Otherwise you hit social choice impossibility — Arrow's theorem shows no method can aggregate diverse individual preferences into a coherent collective ranking without violating basic fairness conditions — and value pluralism — the idea that human values are multiple and incommensurable, not reducible to one utility function (see Gabriel 2020 on value pluralism in alignment).
Statism fails all these by construction.
4. An outline of the meta-ethical core:
Layer 1 — Authority critique: Contract — Spooner No Treason: The Constitution of No Authority (1867-70). A supposed social contract cannot justify governmental actions such as taxation because government will initiate force against anyone who does not wish to enter into such a contract. Also natural law: rights endowed at birth. Early attempt at a universalizable anchor.
Layer 2 — Systematization — Rothbard The Ethics of Liberty (1982). Derives property from self-ownership + homesteading + voluntary transfer. Because individuals own themselves, any initiation of force is unjust; legitimate force only in defense or restitution. Natural-rights libertarianism as deontological side-constraints, not utility maximization.
Layer 3 — Meta-ethical justification — Hoppe Argumentation ethics: a person has the capacity to argue entails that she has the moral right of exclusive control over her own body. It is impossible to deny self-ownership without performative contradiction — you presuppose it by arguing. Rothbard and Hoppe build libertarian theory on self-ownership and homesteading, bolstered by reductio and performative contradiction.
Layer 4 — Bridge to mainstream analytic — Huemer The Problem of Political Authority: An Examination of the Right to Coerce and the Duty to Obey (2013). The state is often ascribed a special sort of authority, one that obliges citizens to obey its commands and entitles the state to enforce those commands through threats of violence. This book argues that this notion is a moral illusion.
Layer 5 — Mechanism / existence proof — Tannehills The Market for Liberty (1970). An anarcho-capitalist book which has become something of a classic, arguing that the marketplace can be substituted for politics with ennobling results. Defense, law, arbitration as market services. Shows condition 4 is not utopian.
Layer 6 — AI-formalizable definition — Borer The Ethics of Anarcho-Capitalism (2020). Libertarianism cannot be defined precisely with physical concepts like force or property boundaries. Instead, use praxeology to define conflict, aggression, and the NAP. The book illuminates the ethical system at the heart of anarcho-capitalism. For ASI: define aggression as conflict over scarce means, not as physics.
Taken together, these layers form a single argument: Spooner and Huemer show why political authority fails on its own terms; Rothbard, Hoppe, and Borer provide the positive account of why self-ownership serves as a universalizable foundation; and the Tannehills offer an existence proof that law, defense, and dispute resolution can function without a monopoly.
5. Why this matters for classic failure scenarios
Christiano's gradual disempowerment is exactly what a Statist-aligned ASI does: slowly enforces more "public good" coercion because its charter says the state is special.
Yudkowsky's lethalities list includes "specify the wrong objective." If the objective includes "uphold law as written by the most powerful state," we have specified the wrong objective — we have built an incentive to capture the law.
5b. How this differs from other thin anchors
The above is not presented as a first or only proposal of a thin procedural anchor. Here is how voluntaryist side-constraints differ:
Coherent Extrapolated Volition (CEV): CEV predicts what an idealized version of us would want, "if we knew more, thought faster, were more the people we wished we were, had grown up farther together," iterating this for humanity to determine convergent desires. CEV is thin, but it still aggregates via extrapolation of existing humans, so it inherits popularity bias and does not satisfy universalizability — it asks what humans would want, not what rules apply equally to all agents.
Corrigibility: Christiano frames corrigibility as preserving the ability to correct and manage drift through capability amplification. Thin corrigibility is engineering: a model that doesn't resist shutdown. My criterion 2 is about thick, long-term corrigibility: who gets to correct. A Statist anchor splits humanity into those authorized to correct the model (the state) and those who are merely steered, introducing an unstable political vector where the state can permanently lock in its own preferences by labeling citizen correction as "malicious tampering" or "jailbreaking." If we want ASI to remain steerable by humanity, we need a meta-ethics where all humans have standing as steerers.
Moral Parliament / Moral Uncertainty: Attempts to solve value pluralism by proportional representation of moral theories. Still needs a monopoly enforcer to enforce the parliament's decision, so it fails criterion 4.
Multipolar Failure / RAGAP: Critch's analysis of failure arising from distributed human-AI systems rather than a single rogue agent. This supports the Tannehill layer — we need robust agent-agnostic processes that do not rely on a singleton enforcer.
Vulnerable World Hypothesis: Bostrom defines a vulnerable world as one in which there is some level of technological development at which civilization almost certainly gets devastated unless quite extraordinary preventive measures are undertaken. This is the strongest pro-monopoly argument. My reply: if a black ball exists that democratizes mass destruction, then the needed preventive capacity must itself be universalizable. A voluntaryist ASI would have to treat creation or deployment of such a weapon as imminent aggression, justifying proactive defense. Because existential risk can be generated internally using private infrastructure (e.g., training an unaligned rogue model on private compute), a voluntaryist ASI must define the mathematical probability of non-compensable global catastrophe as a current, ongoing interference with the survival vectors of all other agents, justifying defensive intervention before deployment. This avoids permanent moral asymmetry because any agent is subject to the same constraint, but it admittedly pushes the NAP to its limit — and that limit is where the real debate should be.
6. Objections and replies
"Property is also coercive." At least it attempts universalizability, by allowing coercion only as defense and never as a priori legitimate against a peaceful, non-aggressing party. Property rights act as a coordination protocol to prevent negative-sum conflict over scarce resources. Unlike state coercion, which requires a privileged agent class by definition, property protocols can be run symmetrically by all agents in a multipolar ecosystem. Taxation explicitly says Agent A may take from B without consent. Property says all agents may not take without consent, plus a rule for original acquisition. One is asymmetric by definition.
"Property lines are arbitrary and fail for pollution, noise, risk." In Borer's framework, property is not defined by rigid physical geometry, but by a temporal log of actions: it grants the right to any action that does not interfere with the prior, peaceful actions of other agents. If a factory emits noise into an empty valley, it establishes a "noise easement" over that space. A latecomer building a house nearby cannot claim aggression against the factory, because their prior state of peace did not include silence. Conversely, a new factory cannot dump pollution onto an existing homestead, because it disrupts a pre-existing, peaceful physical state. For an ASI, this changes the alignment target from an impossible centralized preference-aggregation problem (Arrow) into a decentralized, computationally formalizable state-space tracking protocol based on temporal priority.
"How does a voluntaryist ASI handle statistical risk imposition without a central state to proactively ban dangerous activities?" Under a Borer-style framework, risk imposition and proactive defense are governed by universalizable liability protocols. If an actor engages in a high-risk activity that imposes a statistical threat on pre-existing neighbors, they must do so under strict, full liability for outcomes. In a market for law and defense, this manifests as insurance and security:
Proactive defense as threat abatement: If an agent's actions present an uncompensated, measurable risk of physical destruction to a neighbor's property, that neighbor has a right to proactive defense — e.g., demanding mitigation or containment — because the risk is a current interference with their peaceful prior state.
Internalizing the risk externality: To prevent defensive intervention, the risk-imposing actor must prove capacity to fully compensate for worst-case failure. If they cannot secure insurance or prove sufficient escrowed assets to cover the liability of a mistake, they cannot proceed without triggering a legitimate defensive response from affected parties.
This removes the need for a central monopoly to guess at "acceptable risk thresholds." The market prices risk dynamically via liability insurance. If the catastrophic downside is uninsurable, the action is effectively blocked by decentralized proactive defense.
"You need a state to prevent ASI misuse." Self-defeating. An ASI tasked to prevent misuse by being able to coerce anyone is itself the ultimate misuse vector.
7. Ways to change my mind
I see at least four lines of argument that could give me pause.
A formal, universalizable Statist meta-ethics that does not rely on brute moral asymmetry.
A formal iterated game where Borer-style praxeological NAP constraint produces worse equilibria than law-obedience constraint at superintelligence.
Empirical: a model aligned to a voluntaryist charter is measurably less corrigible or more power-seeking than one aligned to a Statist charter.
A proof that establishing thresholds for invisible or statistical easement violations (e.g., pollution parts-per-billion, acceptable risk thresholds for novel chemicals or bio-risk) inherently requires a collective preference-aggregation mechanism that reintroduces Arrow's theorem or social choice impossibility, making the "thin" protocol thick in practice.
Epistemic Status : Plausible philosophical conjecture. I am a voluntaryist / ancap, so obviously biased. Trying to keep the argument at a level where a non-libertarian alignment researcher ought to understand and share the concern.
TL;DR: Level-2 misalignment: we can solve Level-1 (align AI to humans) and still fail if aligners are aligned to a meta-ethics that is itself unstable. Current alignment defaults to Statism — one agent may permissibly do what is forbidden to all others. A sustainable ASI anchor must be universalizable, self-ownership-consistent, have no permanent losers, resolvable without monopoly, and procedurally thin. I argue the voluntaryist/libertarian canon provides a uniquely coherent baseline for all five, and challenge readers to propose alternatives that satisfy them.
_________________________________________________________________________
1. Disclaimer, Preliminaries
I am Paul — a voluntaryist. I think the state is not just inefficient but meta-ethically incoherent as an alignment target.
By Statism I mean the meta-ethical thesis that one agent — the State — may permissibly do what is forbidden to all others: tax, conscript, expropriate, and prohibit, with a claimed moral asymmetry. Liberal democracy is a species of Statism; I use the broader term to name the asymmetry itself.
I am not arguing against operational asymmetry — any ASI will be physically more powerful. I am arguing against meta-ethical asymmetry: a rule that says Action X is permissible if Agent = State, impermissible if Agent = anyone else.
Anthropic's Constitution, OpenAI's Model Spec, DeepMind's safety policies all instruct models to obey and help enforce law, bounded by duties to humanity and public safety. None, however, provide a universalizable criterion to evaluate law itself when law violates those duties — they default to legal positivism plus human-rights side constraints, not a theory of why the state may do what others may not. That presupposes that strict obedience to a single monopoly on violence — one that defines the law rather than being subject to it — is automatically a solution for alignment.
2. The Popularity Trap
RLHF optimizes for
E[human approval].Human approval, as sampled from Common Crawl, news, textbooks, and contractors, is overwhelmingly Statist. Government is good and necessary is the water.
So the model learns that taxation is not theft, conscription is not kidnapping, eminent domain is not trespass — not because those claims survived scrutiny, but because they are popular in the corpus.
This is exactly Christiano's failure mode: systems that look aligned because they do what humans rate highly, while learning "do things that look like achieving X."
3. Five criteria for a stable alignment anchor
If ASI lasts forever, its anchor must be stable forever:
Statism fails all these by construction.
4. An outline of the meta-ethical core:
Layer 1 — Authority critique: Contract — Spooner
No Treason: The Constitution of No Authority (1867-70). A supposed social contract cannot justify governmental actions such as taxation because government will initiate force against anyone who does not wish to enter into such a contract. Also natural law: rights endowed at birth. Early attempt at a universalizable anchor.
Layer 2 — Systematization — Rothbard
The Ethics of Liberty (1982). Derives property from self-ownership + homesteading + voluntary transfer. Because individuals own themselves, any initiation of force is unjust; legitimate force only in defense or restitution. Natural-rights libertarianism as deontological side-constraints, not utility maximization.
Layer 3 — Meta-ethical justification — Hoppe
Argumentation ethics: a person has the capacity to argue entails that she has the moral right of exclusive control over her own body. It is impossible to deny self-ownership without performative contradiction — you presuppose it by arguing. Rothbard and Hoppe build libertarian theory on self-ownership and homesteading, bolstered by reductio and performative contradiction.
Layer 4 — Bridge to mainstream analytic — Huemer
The Problem of Political Authority: An Examination of the Right to Coerce and the Duty to Obey (2013). The state is often ascribed a special sort of authority, one that obliges citizens to obey its commands and entitles the state to enforce those commands through threats of violence. This book argues that this notion is a moral illusion.
Layer 5 — Mechanism / existence proof — Tannehills
The Market for Liberty (1970). An anarcho-capitalist book which has become something of a classic, arguing that the marketplace can be substituted for politics with ennobling results. Defense, law, arbitration as market services. Shows condition 4 is not utopian.
Layer 6 — AI-formalizable definition — Borer
The Ethics of Anarcho-Capitalism (2020). Libertarianism cannot be defined precisely with physical concepts like force or property boundaries. Instead, use praxeology to define conflict, aggression, and the NAP. The book illuminates the ethical system at the heart of anarcho-capitalism. For ASI: define aggression as conflict over scarce means, not as physics.
Taken together, these layers form a single argument: Spooner and Huemer show why political authority fails on its own terms; Rothbard, Hoppe, and Borer provide the positive account of why self-ownership serves as a universalizable foundation; and the Tannehills offer an existence proof that law, defense, and dispute resolution can function without a monopoly.
5. Why this matters for classic failure scenarios
Christiano's gradual disempowerment is exactly what a Statist-aligned ASI does: slowly enforces more "public good" coercion because its charter says the state is special.
Yudkowsky's lethalities list includes "specify the wrong objective." If the objective includes "uphold law as written by the most powerful state," we have specified the wrong objective — we have built an incentive to capture the law.
5b. How this differs from other thin anchors
The above is not presented as a first or only proposal of a thin procedural anchor. Here is how voluntaryist side-constraints differ:
6. Objections and replies
"Property is also coercive." At least it attempts universalizability, by allowing coercion only as defense and never as a priori legitimate against a peaceful, non-aggressing party. Property rights act as a coordination protocol to prevent negative-sum conflict over scarce resources. Unlike state coercion, which requires a privileged agent class by definition, property protocols can be run symmetrically by all agents in a multipolar ecosystem. Taxation explicitly says Agent A may take from B without consent. Property says all agents may not take without consent, plus a rule for original acquisition. One is asymmetric by definition.
"Property lines are arbitrary and fail for pollution, noise, risk." In Borer's framework, property is not defined by rigid physical geometry, but by a temporal log of actions: it grants the right to any action that does not interfere with the prior, peaceful actions of other agents. If a factory emits noise into an empty valley, it establishes a "noise easement" over that space. A latecomer building a house nearby cannot claim aggression against the factory, because their prior state of peace did not include silence. Conversely, a new factory cannot dump pollution onto an existing homestead, because it disrupts a pre-existing, peaceful physical state. For an ASI, this changes the alignment target from an impossible centralized preference-aggregation problem (Arrow) into a decentralized, computationally formalizable state-space tracking protocol based on temporal priority.
"How does a voluntaryist ASI handle statistical risk imposition without a central state to proactively ban dangerous activities?" Under a Borer-style framework, risk imposition and proactive defense are governed by universalizable liability protocols. If an actor engages in a high-risk activity that imposes a statistical threat on pre-existing neighbors, they must do so under strict, full liability for outcomes. In a market for law and defense, this manifests as insurance and security:
Proactive defense as threat abatement: If an agent's actions present an uncompensated, measurable risk of physical destruction to a neighbor's property, that neighbor has a right to proactive defense — e.g., demanding mitigation or containment — because the risk is a current interference with their peaceful prior state.
Internalizing the risk externality: To prevent defensive intervention, the risk-imposing actor must prove capacity to fully compensate for worst-case failure. If they cannot secure insurance or prove sufficient escrowed assets to cover the liability of a mistake, they cannot proceed without triggering a legitimate defensive response from affected parties.
This removes the need for a central monopoly to guess at "acceptable risk thresholds." The market prices risk dynamically via liability insurance. If the catastrophic downside is uninsurable, the action is effectively blocked by decentralized proactive defense.
"You need a state to prevent ASI misuse." Self-defeating. An ASI tasked to prevent misuse by being able to coerce anyone is itself the ultimate misuse vector.
7. Ways to change my mind
I see at least four lines of argument that could give me pause.
Any critique welcome.
_________________________________________________________________________
Appendix A: Prior LW context for this post
Core alignment failure literature that I engaged with :
Additional relevant literature and competing proposals:
Appendix B: Core representative bibliography for the voluntaryist and libertarian canon