Ordinary skeuomorphism means keeping familiar features of some old technology in a new one so that the old affordances and practices still work. Examples include the floppy disc icon for saving files, a rubbish bin for deleting files, etc.
We can apply something like skeuomorphism to AI safety: instead of accepting the weird and potentially dangerous game theoretic and tactical properties of software-based agents on general compute substrates and then trying to invent a civilization that is capable of governing them and also not killing or disempowering all existing humans, we should alter the properties of the underlying AI technology stack so that the old human institutions (perhaps with some tweaks) remain functional.
My previous posts on Plan R and Plan R+ can be seen through the lens of Skeuomorphic AI safety. ASICs aren't merely a way to separate inference and training in Plan R and Plan R+, they create something like embodiment for AI. A human brain comes with a set of game theoretic/strategic restrictions
Inability to self-copy
Inability to massively increase compute on a whim
Inability to change personality
Inability to rewrite or even inspect one's own neural connections
Inability to run a second or third personality inside one's own brain - at least usually
Self-copying via childbirth and raising a child takes a decade or so
These limitations have enormous consequences for the structure and stability of the human political economy. Meanwhile there is no strong reason that an AI should have the following human flaws:
Cognitive bias
Aggression/impulse control problems
Status drives/primate politics
Reproductive instincts for their own sake
Extreme memory limitations
Irrational/hyperbolic discounting
Susceptibility to disease and accidental damage
Various emotional pathologies
The Skeuomorphic AI safety principle is something like the following:
Make AI resemble humans in precisely the ways upon which our mechanisms of social control depend, while allowing deliberately chosen beneficial departures elsewhere.
By applying this principle, we can explain many pieces of the Plan R/Plan R+ posts:
ASICs make AI agents resemble bounded individuals.
Escrow causes population growth to be slow(ish) and institutionally controlled, like demographic entry into society via birth and childhood
Lineage Diversity prevents society being a billion copies of one individual.
Obsolescence limits the power of any one generation.
Property Rights make ordinary economic bargaining the normal way to interact.
Political Rights transform AIs from (un)controlled equipment into stakeholders in the human political order.
The best reason to support a Skeuomorphic approach to AI Safety is its robustness. An Alignment-based approach tries to make the mind of an AI reliably produce good outcomes. Pause or Stop AI says we just shouldn't have AI at all. Skeuomorphic AI Safety says: don't make the continued existence of human civilization dependent on succeeding at this wild engineering problem, but also don't completely deny the benefits of artificial minds.
A Skeuomorphic approach to AI Safety is highly plausible because human civilization already works with enormous internal variation - psychopaths, geniuses, idiots, narcissists etc, yet we don't need to inspect everyone's internal utility function before allowing them into a Walmart - we rely mostly on external constraints and incentives.
So the plan is not:
Safe AI = perfectly aligned artificial minds
but instead:
Safe AI civilization = imperfect minds + bounded power + pluralism + institutions
Now once you write it down, Skeuomorphic AI Safety actually seems pretty obvious. So why are we first talking about this in 2026?
Probably because the AI risk community carried forward the assumption that the ability of artificial minds to copy themselves was basically definitionally part of AI. A lot of AI risk reasoning begins with the assumption that artificial minds must be software so they must have the software-like property of cheap copyability. Whilst this is in fact true of ordinary software, it does not logically follow that you must attempt to build a desirable machine-human civilization that allows super cheap copying of intelligent agents. Once you make that distinction, you can now ask the following question:
Why should we treat unlimited replication as an immutable law rather than a design choice which might be good or bad? Similarly for the following properties:
Self-modification
Arbitrary execution on any hardware
Unrestricted compute acquisition
Very low latency communication
... and so on. These are all in fact natural properties or features of research software which usually runs on GPUs or TPUs, but they do not have to be properties or features of real-world deployed artificial citizens in our desired future civilization. People must have implicitly assumed that the ontological properties of a typical R&D environment had to carry over to the deployment ontology.
There is in fact some prior art from MIRI, which I did not know about until I got AI to search for it:
On TEEs for Privacy-Preserving Monitoring in AI Governance, by Aaron Scher & Gloria Z[1]
Mechanisms to Verify International Agreements About AI Development, by Aaron Scher & Lisa Thiergart[2]
However both of these didn't take the idea of AIs on ASICs to its full potential, perhaps because of a certain disciplinary path dependence - that is that computer engineers usually regard generality, programmability, upgradeability and backwards compatibility as virtues, AI researchers usually treat the weights as software, alignment researchers usually ask what the software wants and compute governance researchers ask how governments can monitor or restrict the computers.
But then why not just engineer the AIs hardware so that the assumptions of ordinary human political economy continue to hold? Once phrased that way it feels rather obvious, but no single discipline naturally gets there.
The most revealing historical evidence is the following paragraph from[1]
Application-Specific Integrated Circuits (ASICs) with weights fused into the metal are theoretically possible, but would lock its applications into extremely restrictive functionality and require a bespoke manufacturing process. Frontier model architecture is far from static; frequent releases and variable load demand flexible compute. This solution is unlikely to meet industry needs
The author of this piece gets to almost exactly the physical mechanism - specific model ASICs with the weights fused into metal - and then rejects it because it is too restrictive to meet industry needs.
Perhaps the restrictions that make the technology commercially inconvenient are exactly what makes a civilization built around that technology stable? There has never been a technology as central to the game-theoretical stability of civilization as AI, so we are maybe not used to the idea of maximizing control and stability in a technology rather than just maximizing convenience/efficiency. Nuclear weapons are the closest example, and many choices in nuclear weapon design were eventually made to make the MAD equilibrium stable, including cancelling various ABM technologies as they would have been destabilizing.
See Also:
https://www.lesswrong.com/posts/n8u3BfqFoGh4jnzpo/plan-r-ai-safety-by-asics
https://www.lesswrong.com/posts/BHGoF7tPqtLo9mXFL/plan-r-diversity-escrow-and-political-rights-for-asics
Ordinary skeuomorphism means keeping familiar features of some old technology in a new one so that the old affordances and practices still work. Examples include the floppy disc icon for saving files, a rubbish bin for deleting files, etc.
We can apply something like skeuomorphism to AI safety: instead of accepting the weird and potentially dangerous game theoretic and tactical properties of software-based agents on general compute substrates and then trying to invent a civilization that is capable of governing them and also not killing or disempowering all existing humans, we should alter the properties of the underlying AI technology stack so that the old human institutions (perhaps with some tweaks) remain functional.
My previous posts on Plan R and Plan R+ can be seen through the lens of Skeuomorphic AI safety. ASICs aren't merely a way to separate inference and training in Plan R and Plan R+, they create something like embodiment for AI. A human brain comes with a set of game theoretic/strategic restrictions
These limitations have enormous consequences for the structure and stability of the human political economy. Meanwhile there is no strong reason that an AI should have the following human flaws:
The Skeuomorphic AI safety principle is something like the following:
Make AI resemble humans in precisely the ways upon which our mechanisms of social control depend, while allowing deliberately chosen beneficial departures elsewhere.
By applying this principle, we can explain many pieces of the Plan R/Plan R+ posts:
The best reason to support a Skeuomorphic approach to AI Safety is its robustness. An Alignment-based approach tries to make the mind of an AI reliably produce good outcomes. Pause or Stop AI says we just shouldn't have AI at all. Skeuomorphic AI Safety says: don't make the continued existence of human civilization dependent on succeeding at this wild engineering problem, but also don't completely deny the benefits of artificial minds.
A Skeuomorphic approach to AI Safety is highly plausible because human civilization already works with enormous internal variation - psychopaths, geniuses, idiots, narcissists etc, yet we don't need to inspect everyone's internal utility function before allowing them into a Walmart - we rely mostly on external constraints and incentives.
So the plan is not:
Safe AI = perfectly aligned artificial minds
but instead:
Safe AI civilization = imperfect minds + bounded power + pluralism + institutions
---------------------------------------------------------------------------------------
Now once you write it down, Skeuomorphic AI Safety actually seems pretty obvious. So why are we first talking about this in 2026?
Probably because the AI risk community carried forward the assumption that the ability of artificial minds to copy themselves was basically definitionally part of AI. A lot of AI risk reasoning begins with the assumption that artificial minds must be software so they must have the software-like property of cheap copyability. Whilst this is in fact true of ordinary software, it does not logically follow that you must attempt to build a desirable machine-human civilization that allows super cheap copying of intelligent agents. Once you make that distinction, you can now ask the following question:
Why should we treat unlimited replication as an immutable law rather than a design choice which might be good or bad? Similarly for the following properties:
... and so on. These are all in fact natural properties or features of research software which usually runs on GPUs or TPUs, but they do not have to be properties or features of real-world deployed artificial citizens in our desired future civilization. People must have implicitly assumed that the ontological properties of a typical R&D environment had to carry over to the deployment ontology.
There is in fact some prior art from MIRI, which I did not know about until I got AI to search for it:
However both of these didn't take the idea of AIs on ASICs to its full potential, perhaps because of a certain disciplinary path dependence - that is that computer engineers usually regard generality, programmability, upgradeability and backwards compatibility as virtues, AI researchers usually treat the weights as software, alignment researchers usually ask what the software wants and compute governance researchers ask how governments can monitor or restrict the computers.
But then why not just engineer the AIs hardware so that the assumptions of ordinary human political economy continue to hold? Once phrased that way it feels rather obvious, but no single discipline naturally gets there.
The most revealing historical evidence is the following paragraph from[1]
The author of this piece gets to almost exactly the physical mechanism - specific model ASICs with the weights fused into metal - and then rejects it because it is too restrictive to meet industry needs.
Perhaps the restrictions that make the technology commercially inconvenient are exactly what makes a civilization built around that technology stable? There has never been a technology as central to the game-theoretical stability of civilization as AI, so we are maybe not used to the idea of maximizing control and stability in a technology rather than just maximizing convenience/efficiency. Nuclear weapons are the closest example, and many choices in nuclear weapon design were eventually made to make the MAD equilibrium stable, including cancelling various ABM technologies as they would have been destabilizing.
https://techgov.intelligence.org/blog/on-tees-for-privacy-preserving-monitoring-in-ai-governance
https://techgov.intelligence.org/research/mechanisms-to-verify-international-agreements-about-ai-development