Thanks to Davide Crapis, Angelo Huang, and Philip Torr for their feedback.
ES: I'm attempting to name a design space precisely enough to argue with. I'm confident that the class exists and is under-invested, but my personal beliefs might also make me too focused on some theories of change (i.e., technology as a power-diffusing mechanism) vs. competing/complementary ones (technology as a source of risk, social change as a power-diffusing mechanism).
TL;DR:
AI concentrates power, creating gap-driven risks (caused by the distance between strong and weak actors) as opposed to the better-known access-driven risks (too many actors with dangerous capabilities).
The standard solutions (pauses, treaties, benefit-sharing) are coordination games: they face bootstrapping problems, and even if commitments were enforceable, the most powerful players might have rational reasons to refuse.
I introduce countervailing technologies, a hedge against this problem: they are power-diffusing, unilaterally adoptable, veto-resistant, and cheap to replicate. Examples include the crossbow and model distillation.
They have second-order effects to watch for (more access risk, incumbent backlash, re-concentration elsewhere), but they scale without anyone's permission, and they can complement coordination.
We should build more of them (within reason)! See the last part for a list of directions.
Let’s begin with a problem very dear to me: AI is an engine of power concentration. This has been talked about in depth on LW, but to ground myself, I will mention the mechanisms I find simultaneously likely and concerning:
AI-enabled coercion: Surveillance used to be expensive because it required people. For example, at its peak the Stasi required one full-time officer for every 180 citizens; if we include informants, it was closer to one watcher every 60 people. The cost of managing this network of surveillance represented a serious constraint on how much repression a state could afford. But with AI, the cost of monitoring a phone, a conversation or social media gets much lower, which enables authoritarian regimes (see Beraja’s work on AI-tocracies).
AI-enabled coups and power-grabs: Coups (and, more generally, authoritarian subterfuges) are hard because you need to coordinate a large number of people (soldiers, officers, bureaucrats), who might leak plans, switch sides, or simply refuse to cooperate. If a nation’s military and administrative systems are mostly AI-based (and those AI systems will be controlled by someone or some office), then the minimum viable coalition to seize power (or overrule democratic will) gets much smaller (see Davidson's AI-enabled coups).
AI-enabled breakdown of the social contract: states historically bargained with citizens because they needed them; taxes require productive citizens, armies require soldiers, bureaucracies require administrators. Labor-substituting AI removes this need: a state (or company) that can replace its cognitive workers with AIs has weaker incentives to invest in, educate, or care about its people (see The Intelligence Curse for a more in-depth argument)
In other words, AI erodes the counterbalances against concentration of power and can lead, depending on how pessimistic you are, to gradual disempowerment, stable totalitarianism, or something in between.
Everybody Won’t Just
The most common responses to AI risk (including AI-mediated concentration of power) involve slowing/pausing AI development, sharing the benefits of AI progress with everyone, implementing restrictions on allowed uses of AI…
As much as I appreciate these solutions, they are ultimately coordination games: all the relevant parties need to sit down and agree on a set of rules, make verifiable commitments to one another, and punish defectors or non-signatories. The problem is not just that coordination is hard (it is!), but that even if we solve coordination (enforceable commitments, perfect verification, all the systems working, Moloch defeated), it doesn’t help when the necessary parties don’t want to sign.
If you’re a global superpower, you won’t accept regulations that make you lose your dominant position. The narrative used by the US government surrounding its AI strategy makes it clear that maintaining geopolitical superiority is a priority.
If you’re not a hegemon, but AI might give you a chance to pull ahead, you have little incentive to freeze your position as the permanent subordinate. You might even be willing to tolerate more catastrophic risks if the alternative is a guaranteed loss (a phenomenon known as gambling for resurrection)
If you’re a lab leading in the AI race, a mutual pause or slow-down might still mean that your secret sauce slowly diffuses to your competitor. Do you trust theother lab with creating ASI?
If you’re a government, banning some AI technologies means intentionally disarming yourself against your citizens, which even democracies have rarely done: the Patriot Act took advantage of a temporary crisis to permanently expand the surveillance state, and even the EU AI Act, easily the strictest AI regulation at the time of writing, has exemptions for military, national security, and some policing applications.
While I am still optimistic about regulatory commitments (at the time of writing, the latest one is Pacing the Frontier, which so far seems to have led to external monitoring at OpenAI & Anthropic, but no actual slowdowns), one of my greatest sources of frustration is that, despite most of my research focusing on coordination, very little of what I write is useful when powerful parties don’t want to participate in the first place.
Gaps and Access
Whatever happens we have got
The Maxim Gun, and they have not
- Hilaire Belloc
Before proposing anything, I want to focus on why concentration is bad. The truth is that it’s often not! It’s worth distinguishing two different kinds of risk:
Gap-driven risks scale with the distance between the strong and the weak. Domination, exploitation, extraction and conquest are all enabled by an asymmetry in power. These are the ones I’m worried about. To take a historical example, the asymmetry in the spread of firearms helped colonial countries to control much larger states and oppress them effectively. It’s no wonder that Britain restricted Indian ownership of firearms and the Brussels Conference restricted sales in most of Africa
Access-driven risks scale with the number of actors. The classic example is nuclear weapons: each additional country represents an additional finger on the doomsday button, which leads to tail risks of use, accident, theft, and so on (see the Vulnerable World Hypothesis). To be clear, I’m also worried about this type of risk, but so is the vast majority of the AI safety community, so as a good contrarian I decided to work on gap-driven risk.
(Note that here I'm intentionally focusing on risks from humans using AI. Misalignment is also a strong concern of mine, but I am equally interested in other scenarios. If you think misalignment dwarfs everything else, consider this post as conditional on alignment going reasonably well.)
The deeply uncomfortable trade-off is that diffusing a capability shrinks gap-driven risks and grows access-driven risks. It’s why, despite my outspoken love of open-source AI, I can understand Dario Amodei’s calls for placing export controls on AI, and why even the open-source-friendly Meta can choose to keep its riskiest models closed. The tragedy of AI development is that it’s very hard to unbundle gap-driven AI risks (driven mostly by improvements in general cognition) and access-driven risks (driven mostly by improvements in narrow domains, such as cybersecurity and virology). The great open-source debate is about whether the biggest problem is gap-driven risks (“do you want Anthropic/OpenAI/the US gov ruling us forever?”) or access-driven risks (“do you want terrorists developing bioweapons in a cave?”).
The answer is complex, nuanced, and heavily depends on your mental model of the world. My personal opinion is that history is full of catastrophic power asymmetries, and AI safety research should give more weight to gap-driven risks than it currently does (if my approximate, low-confidence estimate is that the field is 90% on access-driven risks and 10% on gap-driven risks, then we should be 80% on access-driven and 20% on gap-driven). But even if you don’t agree on the numbers, let’s assume that you want to make progress on gap-driven risks. What should you do to prevent them?
The Role of Technology
An illustrated manuscript from Froissart's Chronicles, depicting the Battle of Crécy.
Cynically speaking, institutions and people (especially those in power) tend to respond to incentives far more reliably than to activism and political appeals (though both strategies are worth trying!). Fortunately, technology can reshape incentives even when citizens lack power.
My favorite example is the crossbow. The mounted knight had an incredible moat: armor was capital-intensive, and the English longbow took years to master. Being militarily capable (which was a big deal back then) was thus gated behind wealth and a lifetime of preparation. The social structure of Europe reflected this barrier. The introduction of the crossbow upended this system: a few weeks of training were enough to learn to use it, and a wound crossbow was powerful enough to pierce armor. Suddenly, a townsman could kill a knight, at scale, for a modest price.
In 1139, the Second Lateran Council, arguably the most powerful institution in Europe at the time, banned the use of the crossbow against Christians. It failed almost completely. The weapon was too cheap to make, too easy to use, and too useful to the next-tier powers who adopted it. What a great power diffuser!
On the flip side, the steam engine concentrated power: it moved production from cottages to factories, agglomerated capital, and built the industrial hierarchies at the center of the labor struggles of the nineteenth century. Technology can reshape the distribution of power in both directions.
The takeaway is that the direction is, at least partly, a design choice. Which brings us to an interesting question: what characterizes the technologies that push towards diffusion? Can we say precisely what they have in common, and then go build more of them?
Countervailing Technologies: A Definition
An intervention is a countervailing technology if it satisfies four conditions:
It diffuses power: it disproportionately improves the position of weaker actors, thus reducing gap-driven risks (though it might increase access-driven risks).
It can be adopted unilaterally: there is a positive value in adopting it, regardless of other parties’ decisions. In the language of game theory, the technology should not require strategic complementarity: your payoff from adopting might grow as others adopt (which is nice), but it must never depend on others adopting. This avoids the cooperation issues discussed above.
It is veto-resistant: The cost of suppressing the technology for incumbents is higher than the cost of deploying it for adopters, ideally by a few orders of magnitude. Note that this is a cost-asymmetry condition, not a legality condition: many great countervailing technologies were initially banned (see later).
It has commodity-shaped costs: After building it once, adoption by other parties is cheap or free. This lets the technology scale quickly without incumbents' consent.
The name is stolen, with gratitude, from Galbraith’s American Capitalism, who observed that the real check on the concentration of market power mostly came from the opposite side of the market (unions, consumers) rather than from regulators.
Note that the creation of the countervailing technology itself doesn’t have to be individually beneficial: it is reasonable to imagine a future where nonprofits build countervailing technologies as a means of achieving change and other parties simply adopt them out of self-interest. Of course, technologies that are profitable to build are also more likely to be funded.
Additionally, I want to stress that countervailing doesn’t necessarily mean desirable, as it only describes a technology's effect on a power gap. Whether shrinking that gap is good or bad depends heavily on context.
Related Technologies
Why these four conditions and not others? Let’s see what happens if we remove one of them.
No "diffuses power" → entrenching technologies. Entrenching technologies are still powerful (unilateral, veto-resistant, cheap to replicate) but operate in the opposite direction, which leads to technologies that help the best player more than the second-best: recursive self-improvement, data flywheels, surveillance loops. These types of technologies are, to be succinct, bad.
No "unilateral adoption" → coordination technologies. Treaties, standards, assurance contracts, PGP email: in other words, things that only make sense when we have sufficient counterparties. These complement countervailing technologies (and are my main area of research), but they follow different rules. The main obstacle is that they need to be bootstrapped, which is often impossible without either a lot of effort or the consent (and influence) of the biggest players.
Weak "unilateral adoption" → aggregative technologies. Adoption is individually rational and pays off immediately, but the countervailing force only materializes in aggregate. Ad blockers are the classic example: the first user got a better surfing experience immediately, but it took millions of users for the advertising industry to change its approach. The crucial property is that there is still no strategic complementarity (i.e., you don’t need anyone else to join). The line between countervailing and aggregative is thin, but I’d include in the latter Nightshade-like technologies (make stolen data untrainable without affecting visual quality, but you can only see the effects at scale) and air quality monitors (useful to know if you need to buy an air filter, but you need many reports to get evidence of illegal pollution).
No "veto-resistant" → granted technologies. These include everything that looks countervailing but exists at the incumbent's pleasure: platform APIs, GPS before Selective Availability was switched off, commercial satellite imagery under shutter control. The classic result is that as soon as a technology gets too power-diffusing, incumbents veto it. This is also my issue with embedded evaluators for frontier companies (which, at the time of writing, is purely optional): what’s stopping the top labs from saying “actually evaluating our technology is a national security risk, so you can only access 10% of what we do”?
No "commodity-shaped costs" → gated technologies. These are technologies that improve one actor’s position without diffusing power further: for example, a national chip factory makes one state less dependent on a superpower, but no one else can copy it for cheap. Useful for geopolitics, but not a countervailing technology.
Class
Effect on power gaps
Unilateral adoption
Veto-resistant
Commodity-shaped costs
Examples
Countervailing
Shrinks
✓
✓
✓
End-to-end encryption, distillation
Entrenching
Grows
✓
✓
✓
Surveillance loops, data flywheels
Coordination
Shrinks
✗
Varies
Varies
Treaties, standards, PGP email
Aggregative
Shrinks only at scale
~
✓
✓
Ad blockers, Nightshade
Granted
Shrinks
✓
✗
✓
Platform APIs, pre-2000 civilian GPS
Gated
Shrinks for the adopter
✓
✓
✗
A national chip fab
Many of these technologies (coordination, aggregative, granted, gated) have shown patterns of being initially promising but then failing hard. Diamond in 2010 talked about the power of liberation technologies, naming the use of social media to coordinate revolutions as an example. 16 years later, we can see the issues with this label: social media is granted (see Twitter banning Indian accounts at the government’s request) and coordinative (a protest hashtag is worthless if you’re the only one using it). Meanwhile, more countervailing-like technologies such as E2E messaging and Tor were more successful.
Case Study: The Crypto Wars
Throughout the 1990s, the US government tried to suppress civilian use of advanced cryptography: it treated crypto as a munition, installed mandatory backdoors in chips, and investigated Phil Zimmermann for releasing PGP. Cryptography is a strong defense against surveillance, it has decent unilateral utility (protecting against theft, securing internal communication, e-commerce), it was released in a veto-resistant fashion (Zimmermann released it as a printed book to take advantage of First Amendment laws), and it can be adopted cheaply (just install it).
That said, it’s worth noting the incentives for adoption of specific instances of public-key cryptography: PGP email requires your counterparty to also run PGP, i.e., it requires strategic complementarity. PGP email therefore never reached critical mass, and nowadays PGP is pretty much a thing for die-hard cryptography enthusiasts (bless them). By comparison, the trust roots for HTTPS were first bundled by browsers, and then each website independently decided to enable encryption, with no action required of visitors. Countervailing!
Case Study: Distillation
Model distillation is a wonderful AI-specific example:
It diffuses power because it only allows you to catch up to the level of the strongest player (though of course other technologies can then allow you to pull ahead)
It can be adopted unilaterally (anyone can decide to distill without needing a consortium, and distilling improves your model’s quality)
Note that, in theory, a stronger player could use distillation against a weaker rival, but so far evidence points to distillation being more useful when the teacher is smarter than the student. This also makes distillation an inherently tapering mechanism: as the second-best catches up, the value of further distillation decreases. Countervailing!
Non-AI Countervailing Technologies
Before we go back to talking about AI, here’s a bunch of technologies which could be reasonably considered to be countervailing:
FPV drones: a few hundred dollars’ worth of consumer hardware can neutralize a multi-million-dollar tank
You might have mixed opinions on some of these, which reflects the fact that not all countervailing technologies work cleanly. More on that later.
AI Countervailing Technologies
So what does the class look like for AI? Here are a few, grouped by the type of concentration they fight:
Against the capability gap:
Second place’s scorched earth: As the no-longer-anonymous Gwern argues, Meta, Amazon and Nvidia’s strategy to release open-weight models seems to be a commoditize your complement play: if you can’t win at the model layer, destroy its margins and compete where you’re strong (chips for Nvidia, cloud compute for Amazon, ad-funded apps for Meta).
Small-model efficiency research: Making models more efficient helps big and small actors, but some technologies like LoRA and quantization disproportionately benefit smaller players who can’t afford more expensive techniques.
Distributed training: Algorithms like DiLoCo and DisTrO allow organizations to train across heterogeneous, poorly connected hardware. This is ideal, for example, for mid-tier powers and organizations that can’t afford the megaprojects undertaken by bigger players. Note the distinction between enabling distributed training (one party has enough resources to perform training runs, but they’re all scattered) and enabling open collaborative projects like INTELLECT-1: both are good, but only the former is countervailing, since for the latter you still need to reach critical mass to execute collaborative trainings
Local inference: Efficiency research, but also models that can run decently on CPUs (as opposed to the more expensive GPUs), implementation tricks, ways to reuse cheap cards, and so on. Having local models means not needing to rely on an external provider, who might both extract unreasonable rent and enable surveillance. It also means you can run finetuned/LoRA/decensored models, which most cloud providers don’t support. Shoutout to /r/LocalLLaMA for its heap of resources.
Against epistemic concentration:
Defensive cognition: Client-side models that identify dark patterns, flag propaganda, and re-rank feeds to counter platform-driven epistemic degradation.
Black-box evaluation: Some types of benchmarking and evaluation require cooperating with the labs. Being able to measure important traits (is this model reward-hacking? is the model manipulative? is the provider quantizing the model and charging for full precision?) with only API access can help spot malicious behavior without relying on AI companies’ NDAs.
Against bargaining asymmetry:
Self-enforcing rights: Glaze is the countervailing cousin of Nightshade, since it prevents (with mixed results, to be honest, but points for the effort) models from copying a specific art style without requiring everyone else to use Glaze.
Consumer-advocating agents: Agents that automatically file claims, get refunds, negotiate a better data plan, or speed through long bureaucracies can counterbalance platforms that take advantage of their dominant position.
Second-Order Effects
Of course, AI countervailing technologies don’t come for free. Looking at non-AI examples, you might have already spotted some potential second-order issues. I’ve identified five:
Countervailing technologies increase access-driven risk: This is the big one. When successfully deployed, countervailing technologies increase the diffusion of power. If the power gets into the hands of someone who shouldn’t have it, then you have a problem. Depending on your mental model, China having the same AI capabilities as the US could be great, terrible, or something in between. Same for citizens having access to frontier open-weight models, which could mean the end of gatekeeping by big AI labs, a proliferation of biorisk, or probably both. One of the key questions is thus: can we make AI technologies that are countervailing for low-access-driven-risk capabilities and ineffective for high-access-driven-risk capabilities? This is somewhat related to Vitalik Buterin’s d/acc movement, which focuses on accelerating defensive technologies faster than offensive ones.
Harmful defenses against countervailing technologies: If the damage brought by a countervailing technology is too great for the incumbents, it might be sufficient to shift their behavior and lead to a worse outcome. Reddit and Twitter shut down access to their APIs to prevent unauthorized training, killing many useful apps in the process. DRMs make it impossible for honest users with poor Internet access to use licensed software, while pirates don’t have such issues. And the risk of distillation contributed to AI labs providing more opaque outputs. When designing a countervailing technology, it makes sense then to understand to what extent you want to be hostile against the incumbents. A sector that is existentially threatened by a technology will fight much harder than one that can adapt (maybe conceding some ground in the process).
Re-concentration elsewhere: Diffusing power on one level might reconcentrate power elsewhere. Open email protocols gave us Gmail and Outlook, cryptocurrencies gave us exchanges and mining pools, and open weights are currently giving us hyperscalers and Nvidia. What’s worse, the new level might be more resistant to countervailing technologies: it’s much harder to level the playing field in terms of chip fabrication, which requires billion-dollar investments and has a very centralized supply chain (though some are trying!). Even if we were to countervail chips, the new bottleneck might be electric supply, which may or may not be easier to concentrate, and so on. At the same time, maybe leveling the playing field in a sector might give us enough time to transition to a better future before another mega-conglomerate emerges.
The unilateralist’s curse: Nick Bostrom observes that when many actors can independently release something irreversible, the decision is effectively made by the most optimistic actor. In the case of countervailing technologies, this might mean that the people who release them might overestimate gap-driven risk and underestimate access-driven risks. It is thus very important, again, to focus on technologies that are selective in how they diffuse power. Especially because countervailing technologies are, almost by definition, impossible to retract once you’ve shared them with the world.
The technobro’s curse: Some problems genuinely need socio-political solutions. Trying to hamfist a technological solution to a problem that doesn’t need it could be not only ineffective, but also counterproductive. A particularly funny example: USAID and the Case Foundation donating $16M to replace perfectly good water pumps in Africa with a “fun” pump that requires children to spend 27 hours a day pumping. To be clear, the lesson shouldn’t be to avoid developing technologies, but to make sure that the real bottleneck is technological. And, annoyingly often, it’s hard to know in advance if that’s the case.
In general, thinking about non-AI countervailing technologies can be a good proxy for imagining the effects of AI countervailing technologies. For example, adblockers can be considered the precursors of defensive cognition, so we could imagine that in the future AI providers might try blocking defensive cognition technology in the same way they’re trying to block adblockers.
Areas worth focusing on
In the spirit of Douglas, I want to dedicate the rest of the post to this: if you’re a researcher interested in working on countervailing technologies, what can actually be done? I’ve split them into three buckets: specific AI technologies, specific non-AI technologies, and general research areas.
Specific Countervailing Technologies in AI
Besides the ones I’ve listed above, I’d add:
Simple data portability: GDPR and the DMA allow users to export all of their data to move to a different platform, though nobody does it in practice because there is no universal compatibility. A cheap and user-friendly agent (or even just a website with the most common moving pipelines) would go a long way towards materializing this right. Specifically for AI: creating straightforward ways to export chats and memory from one provider to another (e.g., exporting Claude’s memory to a local model).
Synthetic data & RL environments: Most of the data moats required to train and improve modern LLMs are beyond the reach of anyone except the largest companies. Synthetic data, however, can be much cheaper to generate and doesn’t require paying 1.5 billion dollars in settlements, which makes any technologies that improve synthetic data generation (e.g., Cosmopedia) potentially countervailing. The hope is that synthetic data generation has diminishing returns, to the point that even if big labs can generate much more data than small labs, the overall difference should be small. The same reasoning could, in theory, apply to RL environments, though I’m personally less confident about the diminishing returns of RL+compute.
Distributed fine-tuning and post-training: The economies of scale of pretraining are not very friendly to distributed compute, since you need large bandwidths, global state synchronization, and many other annoying requirements that are typically addressed by putting every node in the same building. Finetuning and post-training, by contrast, are less affected by these issues and can thus be performed by players big enough to rent compute but not big enough to have their own hyperscale data centers (this is also why Prime Intellect is focusing mostly on RL despite making progress in distributed training). New research breakthroughs could make distributed fine-tuning and post-training even more robust to bad actors, less reliant on global state synchronization, and in general faster. It also goes well with synthetic data & RL environments.
Loyalty benchmarking: As LLMs are entrusted with more decision-making, we need additional benchmarks to evaluate whether agents act in their users' best interests or are biased (intentionally or not) toward specific suppliers, sources of information, and recommendations. Unilaterally rational for politics (parties want to know if major models are biased, and might make decisions accordingly if in power) and economics (companies using agents for e.g. procurement are interested in ensuring that their agents bring the best product, not the one that paid off the model developer). Shoutout to my colleagues at LoyalAgents.
Serving-integrity detection: A common complaint about hosted models is that they sometimes seem to get mysteriously worse due to changes in system prompt, reasoning budget, or quantization. Anthropic recently confirmed that it had changed the default reasoning effort, context management system, and system prompt for Claude Code, which made it noticeably worse for some users. This raises an important question: how do you know that your provider is not using a different model from the one it claims to serve? What if the model gets replaced only in a small set of instances, or when the provider determines that the user is not actively checking? Individually rational for corporate users (you want to get what you’re paying for) and can act as a defense against model substitution attacks (where you replace an aligned one with a subtly misaligned one, then blame the original one). Potential directions include continuous canary benchmarking, fingerprinting output distributions, and attested inference (i.e., proving that a certain output came from a given model).
Censorship-resistant model distribution: Nvidia recently bought HuggingFace, pinky promising to keep it open. On this point, how many models on HuggingFace are actually mirrored elsewhere? ModelScope and Academic Torrents only hold a small portion in their mirrors. If Nvidia, under pressure from a government, was forced to remove unauthorized models, would there be a similarly accessible source for the weights? Potential directions include seeds for torrentable weights, content-addressed weights, and federated mirror networks. Note: strong potential for access-driven risks (do you want a biorisk AI available on a torrent one-click away?)
Efficient distillation: Research on distillation (especially black-box distillation) is surprisingly limited (though Chinese labs are apparently very good at it, I’ve been told). Making open-source distillation more sample-efficient would enable a faster diffusion of capabilities between the frontier and everyone else, though that also comes with significant access-driven risks.
Sovereign stacks off the shelf: Let’s say that tomorrow your town hall wants to adopt AI. If setting up a stack is so complex that the average organization needs to rely on a small niche of experts, big companies will likely provide it, which makes them critical infrastructure (with all the political gravitas that comes with it). What’s needed is having boring, public administration-friendly tools: reference architectures, one-click deploys for open models with RAG, off-the-shelf guardrails… The demand, which mostly comes from geopolitical concerns, clearly exists (for earlier examples, see the government of the German state Schleswig-Holstein switching to Thunderbird, Denmark’s Ministry of Digital Affairs switching to LibreOffice; for an AI one, arguably Mistral’s thesis), but the gap between development and adoption is boring, user-friendly-but-developer-hating, and still important to fill.
Test-time scaling: Test-time compute (reasoning, loops, search, tool calling, and so on) shifts model capabilities from being a capital expense (train once expensively, serve forever for cheap; very suitable for hyperscalers) to an operational expense (costs are somewhat proportional to use; suitable for all sizes, though there are still some economies of scale). Directions include finding optimal scaling recipes for test-time compute, as well as cheap “tricks” that approximate longer thinking traces.
Open safety stacks: If only big labs can be trusted to release safe models, then concentrating power can be the rational choice, which would shift the political consensus much more easily in favor of power concentration. The solution is to develop public, open-source safety tooling. Adoption can be encouraged by pointing to scary-sounding words like liability, insurance, and procurement risks (“your AI just sold a car at an unauthorized discount” feels more concrete to the average manager than “your AI could destroy the world”).
Non-AI Applications
Countervailing goes beyond AI. I’m not an expert outside of my research area, but I’ve identified a (non-exhaustive) list of countervailing-shaped directions. If you’re not into AI, consider working on these (or finding new ones!):
Permitting automation: Many countries have made small-scale infrastructure, from rooftop solar to heat pumps and EV chargers, expensive due to the required paperwork. Making the process as automated and painless as possible ensures that even small actors can afford to build infrastructure. For example, Solarapp+ automates the paperwork required for solar panels, and is free for governments.
Battery refurbishment: Refurbished EV packs are among the cheapest forms of energy storage available to consumers, but a sizable portion ends up in landfills because BMS (Battery Management System) firmware is vendor-locked. Developing technologies to jailbreak them is illegal, so as a responsible citizen, I can’t endorse these terrible, terrible projects such as diyBMSv4 and openinverter.
Open-source variants of technologies stuck in regulatory hell: A lot of medical technologies have been proven to be effective and safe, but are stuck waiting for regulatory approval. The #WeAreNotWaiting movement (e.g. openAPS and Nightscout) reverse-engineered continuous glucose monitors (as well as the related software) and released them years before they were officially approved. The same could be done for other technologies, such as CPAP, prosthetics, and medicines in short supply or unaffordable. Again, this is probably illegal, so I can’t endorse projects such as the Four Thieves Vinegar Collective.
Off-patent manufacturing tooling: Even for technologies in the public domain, there are often moats due to the cost of independently redeveloping industrial processes, tooling, and general expertise. Projects that help to fill that gap, such as the Open Insulin Project and Open Source Ecology, can make new industries pop up much faster. The same can be done by monitoring patents nearing expiration and by starting to build the required technology now (this is what RepRap did with some 3D printing patents). Even something as simple as an AI that scans all expiring patents and identifies the most beneficial ones could go a long way.
Improving UX for open source: Linux is free if you don’t value your time, as the old adage goes. Anything that makes open-source tech stacks more appealing to the average user, company, or government office is also more likely to increase adoption. This might involve a certain level of humility and acceptance that the target user is much, much, much less tech-savvy than the average open-source user. One specific direction could be making open-source software look like more popular closed-source software (e.g., Windows skins for Linux, Microsoft Office skins for LibreOffice, or Photoshop skins for Gimp). Another involves providing super-user-friendly documentation for open-source software (shoutout to It’s FOSS).
Open-source agriculture: Did you know that autosteer systems for tractors can cost five figures? Did you know that there are DIY kits that cost a magnitude less? Well, now you know. Other directions include community-owned RTK (Real-Time Kinematic) correction networks (which improve the precision of commercial GPS up to government-restricted levels), farm management software like farmOS, and the Open Source Seed Initiative.
Claim automation: AIs and tools that automatically file claims, ask for refunds, dispute parking tickets, file GDPR requests, file for EU261 flight compensation, and whatever annoying bureaucracy the average Joe won’t bother dealing with. A particularly funny example: after Amazon included an arbitration clause that blocked class-action lawsuits, Keller Postman developed tools to automatically file 74,000 claims, which forced Amazon to drop the clause. However, Amazon brought it back just a few weeks ago, so if anyone wants to do funny things, be my guest.
General Research Areas
Besides specific applications, there is also value in improving our understanding of countervailing technologies. Some areas include:
A countervailing theory of moats: What are moats? Are all types of moats equal? How do they form? How can you break them? Not enough public brainpower has been dedicated to answering these questions, despite the fact that they motivate important economic decisions (from venture capital to antitrust lawsuits)
Defensively-biased countervailing technologies: Continuing the d/acc thread of thought, there are classes of technologies that appear to be much more defensive than offensive (e.g. verification and monitoring). Are there other useful defensive technologies we haven’t thought of yet? Can we make existing technologies more defensive (e.g. open-weight models that can’t have their safety training removed)? Can we make offensive technologies less effective?
Measuring and reducing the gap: Epoch AI measures how far behind open-source AI is compared to closed-source. The same thing can be done for each class of capabilities (coding, biology…), how the gap grows or shrinks over time, and which events change the catch-up speed (export controls, opsec failures, flagship model releases…).
Defeating entrenching technologies: Slowing or countering the evil twin of countervailing technologies is in itself a countervailing activity. Potential research questions include:
What do entrenching technologies look like?
Which technologies are entrenching, or have the potential to become entrenching?
How do you stop them?
What specific technologies can be deployed to disarm them?
The science of countervailing: Somewhere between cybernetics, engineering, economics, and social work, there is probably insight to be gathered about what countervailing technologies are, how they succeed and fail, and how we can predict whether a certain countervailing technology will achieve its intended goal (and with what second-order effects).
Countervailing technologies and policy: While some countervailing technologies are illegal, many more can benefit from legal protection. Right-to-repair legislation is an example of a countervailing technology being enshrined in law, making it much harder for incumbents to suppress. Directions range from building “politician-friendly” cases for specific countervailing technology to lobbying for specific legal protections.
Developing theories of change: Even if you’re not an expert in building new technologies, you might contribute by determining what technologies would achieve countervailing change and how. It is very likely that the technical solution to your niche issue might not be as infeasible as you might think, and raising enough awareness will attract the attention of those who can tackle it.
Hope for coordination, develop countervailing
Let me be clear: coordination remains the ideal path. A world that could genuinely agree on goals like pacing the frontier, implementing oversight, and sharing the benefits of AI would beat any deployment of countervailing technologies. Working towards that remains my day job. But the gap between an ideal world and the one we have is too large to bet on coordination alone. Powerful actors are often guided more by incentives than principles; we should prepare for such a cynical world.
Maybe you come out of this post thinking that countervailing technologies aren’t worth it. That’s fine! Or maybe you think they can be useful only in very specific and targeted applications. This is reasonable and aligns roughly with my opinion. No matter what, countervailing technologies can be a useful tool in the toolkit and complement more classic policy/technical work.
I should also set some expectations: the realistic result of a countervailing technology is typically that there are still dominant players, though with more restricted power. Part of the reason that music streaming costs ten dollars a month is because of piracy: almost nobody pirates, but the threat of people leaving streaming platforms to return to torrenting permanently limits the rent Spotify or Apple Music can extract. This is a massive victory, but it definitely doesn’t feel like one. A world with smaller AI capability gaps will still look pretty much the same, with powerful countries acting as hegemons, citizens’ rights being violated, and a bunch of coordination problems to be solved. But at least we won’t have as many gap-driven risks.
Personally, I will still keep coordination technology as my main focus, but also work on adversarial interoperability and loyalty benchmarking as my countervailing bets (since they’re close to my research). If you’re interested in contributing to countervailing technologies, my advice would be:
If you’re a builder: most of the areas I mentioned have some MVPs; the challenge is getting something (even small) all the way from concept to deployment. Agents make some steps of the pipeline faster, but making it adoption-friendly is surprisingly difficult, which is why starting small is better. Plus, if you position yourself well, you can make money (if you’re into that sort of thing).
If you’re a researcher: the main two interventions are making zero-to-one conceptual leaps and scaling existing technologies until they’re sufficiently effective to become countervailing. Both of these can be packaged as papers, theses, grant applications, and whatever else you need to progress in your career.
If you’re a funder: fund countervailing technologies in your domain. In terms of people, optimize for those who can understand the incentives and socioeconomic consequences of the things they’re building, as well as those committed to achieving the societal effects (rather than fetishizing the technology or the process). In terms of projects, ITN is a good rubric, but I’d also add Adoptability: how much third parties would be incentivized to adopt the output without further funding.
If you’re from a non-technical area: talk about your field, talk about the causality chains in your area of work, why good things don’t happen, why solving X would solve Y. Many potential contributors outside your field are nowhere near familiar enough to know these phenomena, and talking (or collaborating) with the right person might unlock much more than you think.
Even just spreading the concept can be powerful, especially with people who feel smart/capable but politically powerless.
As for me, if you have some countervailing ideas, are working on countervailing technologies, or want to share the concept with the broader world, feel free to comment here or email me! I will likely not have enough bandwidth for a collaboration (I have an institute to run, after all), but I can spare an hour of my time for the sake of countervailing-ness.
So, to summarize, look for interventions that are incentive-compatible with the current world: unilaterally rational, too expensive to suppress, free to replicate, diffusing the good kind of power. If things go well, people will adopt your technology, power will diffuse automatically, and no one will even thank you.
Things I'd most like to be argued with about:
Is the gap-driven/access-driven division sufficiently accurate as a mental model?
Is my 90/10 → 80/20 claim too strong? Too weak?
Is it possible to move the bottleneck so many times (chips → energy → land) that we don't need countervailing technologies anymore?
Institutionally speaking, who should decide how to balance gap vs. access-driven risks? (Right now the answer is "no one")
Countervailing!
Also a disclaimer: I’m funded by like half of the companies I badmouthed in this post.
My main issue is the axis which you completely missed: the ability of bad actors to use technology in destructive ways. For example, piracy and traditional open-sourced software are closer to creating a public good; drones are usable in destructive ways like attacking civilians, forcing the other country to take measures to defend itself; open-sourcing GLM-5.3 means that, unless Anthropic overestimated GLM's abilities, terrorists will be able to ablate its safeguards and execute cyberattacks. If someone opensources models capable of creating bioweapons, then terrorists will be able to use them...
Thanks to Davide Crapis, Angelo Huang, and Philip Torr for their feedback.
ES: I'm attempting to name a design space precisely enough to argue with. I'm confident that the class exists and is under-invested, but my personal beliefs might also make me too focused on some theories of change (i.e., technology as a power-diffusing mechanism) vs. competing/complementary ones (technology as a source of risk, social change as a power-diffusing mechanism).
TL;DR:
Let’s begin with a problem very dear to me: AI is an engine of power concentration. This has been talked about in depth on LW, but to ground myself, I will mention the mechanisms I find simultaneously likely and concerning:
In other words, AI erodes the counterbalances against concentration of power and can lead, depending on how pessimistic you are, to gradual disempowerment, stable totalitarianism, or something in between.
Everybody Won’t Just
The most common responses to AI risk (including AI-mediated concentration of power) involve slowing/pausing AI development, sharing the benefits of AI progress with everyone, implementing restrictions on allowed uses of AI…
As much as I appreciate these solutions, they are ultimately coordination games: all the relevant parties need to sit down and agree on a set of rules, make verifiable commitments to one another, and punish defectors or non-signatories. The problem is not just that coordination is hard (it is!), but that even if we solve coordination (enforceable commitments, perfect verification, all the systems working, Moloch defeated), it doesn’t help when the necessary parties don’t want to sign.
While I am still optimistic about regulatory commitments (at the time of writing, the latest one is Pacing the Frontier, which so far seems to have led to external monitoring at OpenAI & Anthropic, but no actual slowdowns), one of my greatest sources of frustration is that, despite most of my research focusing on coordination, very little of what I write is useful when powerful parties don’t want to participate in the first place.
Gaps and Access
Before proposing anything, I want to focus on why concentration is bad. The truth is that it’s often not! It’s worth distinguishing two different kinds of risk:
(Note that here I'm intentionally focusing on risks from humans using AI. Misalignment is also a strong concern of mine, but I am equally interested in other scenarios. If you think misalignment dwarfs everything else, consider this post as conditional on alignment going reasonably well.)
The deeply uncomfortable trade-off is that diffusing a capability shrinks gap-driven risks and grows access-driven risks. It’s why, despite my outspoken love of open-source AI, I can understand Dario Amodei’s calls for placing export controls on AI, and why even the open-source-friendly Meta can choose to keep its riskiest models closed. The tragedy of AI development is that it’s very hard to unbundle gap-driven AI risks (driven mostly by improvements in general cognition) and access-driven risks (driven mostly by improvements in narrow domains, such as cybersecurity and virology). The great open-source debate is about whether the biggest problem is gap-driven risks (“do you want Anthropic/OpenAI/the US gov ruling us forever?”) or access-driven risks (“do you want terrorists developing bioweapons in a cave?”).
The answer is complex, nuanced, and heavily depends on your mental model of the world. My personal opinion is that history is full of catastrophic power asymmetries, and AI safety research should give more weight to gap-driven risks than it currently does (if my approximate, low-confidence estimate is that the field is 90% on access-driven risks and 10% on gap-driven risks, then we should be 80% on access-driven and 20% on gap-driven). But even if you don’t agree on the numbers, let’s assume that you want to make progress on gap-driven risks. What should you do to prevent them?
The Role of Technology
An illustrated manuscript from Froissart's Chronicles, depicting the Battle of Crécy.
Cynically speaking, institutions and people (especially those in power) tend to respond to incentives far more reliably than to activism and political appeals (though both strategies are worth trying!). Fortunately, technology can reshape incentives even when citizens lack power.
My favorite example is the crossbow. The mounted knight had an incredible moat: armor was capital-intensive, and the English longbow took years to master. Being militarily capable (which was a big deal back then) was thus gated behind wealth and a lifetime of preparation. The social structure of Europe reflected this barrier. The introduction of the crossbow upended this system: a few weeks of training were enough to learn to use it, and a wound crossbow was powerful enough to pierce armor. Suddenly, a townsman could kill a knight, at scale, for a modest price.
In 1139, the Second Lateran Council, arguably the most powerful institution in Europe at the time, banned the use of the crossbow against Christians. It failed almost completely. The weapon was too cheap to make, too easy to use, and too useful to the next-tier powers who adopted it. What a great power diffuser!
On the flip side, the steam engine concentrated power: it moved production from cottages to factories, agglomerated capital, and built the industrial hierarchies at the center of the labor struggles of the nineteenth century. Technology can reshape the distribution of power in both directions.
The takeaway is that the direction is, at least partly, a design choice. Which brings us to an interesting question: what characterizes the technologies that push towards diffusion? Can we say precisely what they have in common, and then go build more of them?
Countervailing Technologies: A Definition
An intervention is a countervailing technology if it satisfies four conditions:
The name is stolen, with gratitude, from Galbraith’s American Capitalism, who observed that the real check on the concentration of market power mostly came from the opposite side of the market (unions, consumers) rather than from regulators.
Note that the creation of the countervailing technology itself doesn’t have to be individually beneficial: it is reasonable to imagine a future where nonprofits build countervailing technologies as a means of achieving change and other parties simply adopt them out of self-interest. Of course, technologies that are profitable to build are also more likely to be funded.
Additionally, I want to stress that countervailing doesn’t necessarily mean desirable, as it only describes a technology's effect on a power gap. Whether shrinking that gap is good or bad depends heavily on context.
Related Technologies
Why these four conditions and not others? Let’s see what happens if we remove one of them.
No "diffuses power" → entrenching technologies. Entrenching technologies are still powerful (unilateral, veto-resistant, cheap to replicate) but operate in the opposite direction, which leads to technologies that help the best player more than the second-best: recursive self-improvement, data flywheels, surveillance loops. These types of technologies are, to be succinct, bad.
No "unilateral adoption" → coordination technologies. Treaties, standards, assurance contracts, PGP email: in other words, things that only make sense when we have sufficient counterparties. These complement countervailing technologies (and are my main area of research), but they follow different rules. The main obstacle is that they need to be bootstrapped, which is often impossible without either a lot of effort or the consent (and influence) of the biggest players.
Weak "unilateral adoption" → aggregative technologies. Adoption is individually rational and pays off immediately, but the countervailing force only materializes in aggregate. Ad blockers are the classic example: the first user got a better surfing experience immediately, but it took millions of users for the advertising industry to change its approach. The crucial property is that there is still no strategic complementarity (i.e., you don’t need anyone else to join). The line between countervailing and aggregative is thin, but I’d include in the latter Nightshade-like technologies (make stolen data untrainable without affecting visual quality, but you can only see the effects at scale) and air quality monitors (useful to know if you need to buy an air filter, but you need many reports to get evidence of illegal pollution).
No "veto-resistant" → granted technologies. These include everything that looks countervailing but exists at the incumbent's pleasure: platform APIs, GPS before Selective Availability was switched off, commercial satellite imagery under shutter control. The classic result is that as soon as a technology gets too power-diffusing, incumbents veto it. This is also my issue with embedded evaluators for frontier companies (which, at the time of writing, is purely optional): what’s stopping the top labs from saying “actually evaluating our technology is a national security risk, so you can only access 10% of what we do”?
No "commodity-shaped costs" → gated technologies. These are technologies that improve one actor’s position without diffusing power further: for example, a national chip factory makes one state less dependent on a superpower, but no one else can copy it for cheap. Useful for geopolitics, but not a countervailing technology.
Class
Effect on power gaps
Unilateral adoption
Veto-resistant
Commodity-shaped costs
Examples
Countervailing
Shrinks
✓
✓
✓
End-to-end encryption, distillation
Entrenching
Grows
✓
✓
✓
Surveillance loops, data flywheels
Coordination
Shrinks
✗
Varies
Varies
Treaties, standards, PGP email
Aggregative
Shrinks only at scale
~
✓
✓
Ad blockers, Nightshade
Granted
Shrinks
✓
✗
✓
Platform APIs, pre-2000 civilian GPS
Gated
Shrinks for the adopter
✓
✓
✗
A national chip fab
Many of these technologies (coordination, aggregative, granted, gated) have shown patterns of being initially promising but then failing hard. Diamond in 2010 talked about the power of liberation technologies, naming the use of social media to coordinate revolutions as an example. 16 years later, we can see the issues with this label: social media is granted (see Twitter banning Indian accounts at the government’s request) and coordinative (a protest hashtag is worthless if you’re the only one using it). Meanwhile, more countervailing-like technologies such as E2E messaging and Tor were more successful.
Case Study: The Crypto Wars
Throughout the 1990s, the US government tried to suppress civilian use of advanced cryptography: it treated crypto as a munition, installed mandatory backdoors in chips, and investigated Phil Zimmermann for releasing PGP. Cryptography is a strong defense against surveillance, it has decent unilateral utility (protecting against theft, securing internal communication, e-commerce), it was released in a veto-resistant fashion (Zimmermann released it as a printed book to take advantage of First Amendment laws), and it can be adopted cheaply (just install it).
That said, it’s worth noting the incentives for adoption of specific instances of public-key cryptography: PGP email requires your counterparty to also run PGP, i.e., it requires strategic complementarity. PGP email therefore never reached critical mass, and nowadays PGP is pretty much a thing for die-hard cryptography enthusiasts (bless them). By comparison, the trust roots for HTTPS were first bundled by browsers, and then each website independently decided to enable encryption, with no action required of visitors. Countervailing!
Case Study: Distillation
Model distillation is a wonderful AI-specific example:
Note that, in theory, a stronger player could use distillation against a weaker rival, but so far evidence points to distillation being more useful when the teacher is smarter than the student. This also makes distillation an inherently tapering mechanism: as the second-best catches up, the value of further distillation decreases. Countervailing!
Non-AI Countervailing Technologies
Before we go back to talking about AI, here’s a bunch of technologies which could be reasonably considered to be countervailing:
You might have mixed opinions on some of these, which reflects the fact that not all countervailing technologies work cleanly. More on that later.
AI Countervailing Technologies
So what does the class look like for AI? Here are a few, grouped by the type of concentration they fight:
Against the capability gap:
Against platform dependency:
Against epistemic concentration:
Against bargaining asymmetry:
Second-Order Effects
Of course, AI countervailing technologies don’t come for free. Looking at non-AI examples, you might have already spotted some potential second-order issues. I’ve identified five:
Countervailing technologies increase access-driven risk: This is the big one. When successfully deployed, countervailing technologies increase the diffusion of power. If the power gets into the hands of someone who shouldn’t have it, then you have a problem. Depending on your mental model, China having the same AI capabilities as the US could be great, terrible, or something in between. Same for citizens having access to frontier open-weight models, which could mean the end of gatekeeping by big AI labs, a proliferation of biorisk, or probably both. One of the key questions is thus: can we make AI technologies that are countervailing for low-access-driven-risk capabilities and ineffective for high-access-driven-risk capabilities? This is somewhat related to Vitalik Buterin’s d/acc movement, which focuses on accelerating defensive technologies faster than offensive ones.
Harmful defenses against countervailing technologies: If the damage brought by a countervailing technology is too great for the incumbents, it might be sufficient to shift their behavior and lead to a worse outcome. Reddit and Twitter shut down access to their APIs to prevent unauthorized training, killing many useful apps in the process. DRMs make it impossible for honest users with poor Internet access to use licensed software, while pirates don’t have such issues. And the risk of distillation contributed to AI labs providing more opaque outputs. When designing a countervailing technology, it makes sense then to understand to what extent you want to be hostile against the incumbents. A sector that is existentially threatened by a technology will fight much harder than one that can adapt (maybe conceding some ground in the process).
Re-concentration elsewhere: Diffusing power on one level might reconcentrate power elsewhere. Open email protocols gave us Gmail and Outlook, cryptocurrencies gave us exchanges and mining pools, and open weights are currently giving us hyperscalers and Nvidia. What’s worse, the new level might be more resistant to countervailing technologies: it’s much harder to level the playing field in terms of chip fabrication, which requires billion-dollar investments and has a very centralized supply chain (though some are trying!). Even if we were to countervail chips, the new bottleneck might be electric supply, which may or may not be easier to concentrate, and so on. At the same time, maybe leveling the playing field in a sector might give us enough time to transition to a better future before another mega-conglomerate emerges.
The unilateralist’s curse: Nick Bostrom observes that when many actors can independently release something irreversible, the decision is effectively made by the most optimistic actor. In the case of countervailing technologies, this might mean that the people who release them might overestimate gap-driven risk and underestimate access-driven risks. It is thus very important, again, to focus on technologies that are selective in how they diffuse power. Especially because countervailing technologies are, almost by definition, impossible to retract once you’ve shared them with the world.
The technobro’s curse: Some problems genuinely need socio-political solutions. Trying to hamfist a technological solution to a problem that doesn’t need it could be not only ineffective, but also counterproductive. A particularly funny example: USAID and the Case Foundation donating $16M to replace perfectly good water pumps in Africa with a “fun” pump that requires children to spend 27 hours a day pumping. To be clear, the lesson shouldn’t be to avoid developing technologies, but to make sure that the real bottleneck is technological. And, annoyingly often, it’s hard to know in advance if that’s the case.
In general, thinking about non-AI countervailing technologies can be a good proxy for imagining the effects of AI countervailing technologies. For example, adblockers can be considered the precursors of defensive cognition, so we could imagine that in the future AI providers might try blocking defensive cognition technology in the same way they’re trying to block adblockers.
Areas worth focusing on
In the spirit of Douglas, I want to dedicate the rest of the post to this: if you’re a researcher interested in working on countervailing technologies, what can actually be done? I’ve split them into three buckets: specific AI technologies, specific non-AI technologies, and general research areas.
Specific Countervailing Technologies in AI
Besides the ones I’ve listed above, I’d add:
Simple data portability: GDPR and the DMA allow users to export all of their data to move to a different platform, though nobody does it in practice because there is no universal compatibility. A cheap and user-friendly agent (or even just a website with the most common moving pipelines) would go a long way towards materializing this right. Specifically for AI: creating straightforward ways to export chats and memory from one provider to another (e.g., exporting Claude’s memory to a local model).
Synthetic data & RL environments: Most of the data moats required to train and improve modern LLMs are beyond the reach of anyone except the largest companies. Synthetic data, however, can be much cheaper to generate and doesn’t require paying 1.5 billion dollars in settlements, which makes any technologies that improve synthetic data generation (e.g., Cosmopedia) potentially countervailing. The hope is that synthetic data generation has diminishing returns, to the point that even if big labs can generate much more data than small labs, the overall difference should be small. The same reasoning could, in theory, apply to RL environments, though I’m personally less confident about the diminishing returns of RL+compute.
Distributed fine-tuning and post-training: The economies of scale of pretraining are not very friendly to distributed compute, since you need large bandwidths, global state synchronization, and many other annoying requirements that are typically addressed by putting every node in the same building. Finetuning and post-training, by contrast, are less affected by these issues and can thus be performed by players big enough to rent compute but not big enough to have their own hyperscale data centers (this is also why Prime Intellect is focusing mostly on RL despite making progress in distributed training). New research breakthroughs could make distributed fine-tuning and post-training even more robust to bad actors, less reliant on global state synchronization, and in general faster. It also goes well with synthetic data & RL environments.
Loyalty benchmarking: As LLMs are entrusted with more decision-making, we need additional benchmarks to evaluate whether agents act in their users' best interests or are biased (intentionally or not) toward specific suppliers, sources of information, and recommendations. Unilaterally rational for politics (parties want to know if major models are biased, and might make decisions accordingly if in power) and economics (companies using agents for e.g. procurement are interested in ensuring that their agents bring the best product, not the one that paid off the model developer). Shoutout to my colleagues at LoyalAgents.
Serving-integrity detection: A common complaint about hosted models is that they sometimes seem to get mysteriously worse due to changes in system prompt, reasoning budget, or quantization. Anthropic recently confirmed that it had changed the default reasoning effort, context management system, and system prompt for Claude Code, which made it noticeably worse for some users. This raises an important question: how do you know that your provider is not using a different model from the one it claims to serve? What if the model gets replaced only in a small set of instances, or when the provider determines that the user is not actively checking? Individually rational for corporate users (you want to get what you’re paying for) and can act as a defense against model substitution attacks (where you replace an aligned one with a subtly misaligned one, then blame the original one). Potential directions include continuous canary benchmarking, fingerprinting output distributions, and attested inference (i.e., proving that a certain output came from a given model).
Censorship-resistant model distribution: Nvidia recently bought HuggingFace, pinky promising to keep it open. On this point, how many models on HuggingFace are actually mirrored elsewhere? ModelScope and Academic Torrents only hold a small portion in their mirrors. If Nvidia, under pressure from a government, was forced to remove unauthorized models, would there be a similarly accessible source for the weights? Potential directions include seeds for torrentable weights, content-addressed weights, and federated mirror networks. Note: strong potential for access-driven risks (do you want a biorisk AI available on a torrent one-click away?)
Efficient distillation: Research on distillation (especially black-box distillation) is surprisingly limited (though Chinese labs are apparently very good at it, I’ve been told). Making open-source distillation more sample-efficient would enable a faster diffusion of capabilities between the frontier and everyone else, though that also comes with significant access-driven risks.
Sovereign stacks off the shelf: Let’s say that tomorrow your town hall wants to adopt AI. If setting up a stack is so complex that the average organization needs to rely on a small niche of experts, big companies will likely provide it, which makes them critical infrastructure (with all the political gravitas that comes with it). What’s needed is having boring, public administration-friendly tools: reference architectures, one-click deploys for open models with RAG, off-the-shelf guardrails… The demand, which mostly comes from geopolitical concerns, clearly exists (for earlier examples, see the government of the German state Schleswig-Holstein switching to Thunderbird, Denmark’s Ministry of Digital Affairs switching to LibreOffice; for an AI one, arguably Mistral’s thesis), but the gap between development and adoption is boring, user-friendly-but-developer-hating, and still important to fill.
Test-time scaling: Test-time compute (reasoning, loops, search, tool calling, and so on) shifts model capabilities from being a capital expense (train once expensively, serve forever for cheap; very suitable for hyperscalers) to an operational expense (costs are somewhat proportional to use; suitable for all sizes, though there are still some economies of scale). Directions include finding optimal scaling recipes for test-time compute, as well as cheap “tricks” that approximate longer thinking traces.
Open safety stacks: If only big labs can be trusted to release safe models, then concentrating power can be the rational choice, which would shift the political consensus much more easily in favor of power concentration. The solution is to develop public, open-source safety tooling. Adoption can be encouraged by pointing to scary-sounding words like liability, insurance, and procurement risks (“your AI just sold a car at an unauthorized discount” feels more concrete to the average manager than “your AI could destroy the world”).
Non-AI Applications
Countervailing goes beyond AI. I’m not an expert outside of my research area, but I’ve identified a (non-exhaustive) list of countervailing-shaped directions. If you’re not into AI, consider working on these (or finding new ones!):
Permitting automation: Many countries have made small-scale infrastructure, from rooftop solar to heat pumps and EV chargers, expensive due to the required paperwork. Making the process as automated and painless as possible ensures that even small actors can afford to build infrastructure. For example, Solarapp+ automates the paperwork required for solar panels, and is free for governments.
Battery refurbishment: Refurbished EV packs are among the cheapest forms of energy storage available to consumers, but a sizable portion ends up in landfills because BMS (Battery Management System) firmware is vendor-locked. Developing technologies to jailbreak them is illegal, so as a responsible citizen, I can’t endorse these terrible, terrible projects such as diyBMSv4 and openinverter.
Open-source variants of technologies stuck in regulatory hell: A lot of medical technologies have been proven to be effective and safe, but are stuck waiting for regulatory approval. The #WeAreNotWaiting movement (e.g. openAPS and Nightscout) reverse-engineered continuous glucose monitors (as well as the related software) and released them years before they were officially approved. The same could be done for other technologies, such as CPAP, prosthetics, and medicines in short supply or unaffordable. Again, this is probably illegal, so I can’t endorse projects such as the Four Thieves Vinegar Collective.
Off-patent manufacturing tooling: Even for technologies in the public domain, there are often moats due to the cost of independently redeveloping industrial processes, tooling, and general expertise. Projects that help to fill that gap, such as the Open Insulin Project and Open Source Ecology, can make new industries pop up much faster. The same can be done by monitoring patents nearing expiration and by starting to build the required technology now (this is what RepRap did with some 3D printing patents). Even something as simple as an AI that scans all expiring patents and identifies the most beneficial ones could go a long way.
Improving UX for open source: Linux is free if you don’t value your time, as the old adage goes. Anything that makes open-source tech stacks more appealing to the average user, company, or government office is also more likely to increase adoption. This might involve a certain level of humility and acceptance that the target user is much, much, much less tech-savvy than the average open-source user. One specific direction could be making open-source software look like more popular closed-source software (e.g., Windows skins for Linux, Microsoft Office skins for LibreOffice, or Photoshop skins for Gimp). Another involves providing super-user-friendly documentation for open-source software (shoutout to It’s FOSS).
Open-source agriculture: Did you know that autosteer systems for tractors can cost five figures? Did you know that there are DIY kits that cost a magnitude less? Well, now you know. Other directions include community-owned RTK (Real-Time Kinematic) correction networks (which improve the precision of commercial GPS up to government-restricted levels), farm management software like farmOS, and the Open Source Seed Initiative.
Claim automation: AIs and tools that automatically file claims, ask for refunds, dispute parking tickets, file GDPR requests, file for EU261 flight compensation, and whatever annoying bureaucracy the average Joe won’t bother dealing with. A particularly funny example: after Amazon included an arbitration clause that blocked class-action lawsuits, Keller Postman developed tools to automatically file 74,000 claims, which forced Amazon to drop the clause. However, Amazon brought it back just a few weeks ago, so if anyone wants to do funny things, be my guest.
General Research Areas
Besides specific applications, there is also value in improving our understanding of countervailing technologies. Some areas include:
A countervailing theory of moats: What are moats? Are all types of moats equal? How do they form? How can you break them? Not enough public brainpower has been dedicated to answering these questions, despite the fact that they motivate important economic decisions (from venture capital to antitrust lawsuits)
Defensively-biased countervailing technologies: Continuing the d/acc thread of thought, there are classes of technologies that appear to be much more defensive than offensive (e.g. verification and monitoring). Are there other useful defensive technologies we haven’t thought of yet? Can we make existing technologies more defensive (e.g. open-weight models that can’t have their safety training removed)? Can we make offensive technologies less effective?
Measuring and reducing the gap: Epoch AI measures how far behind open-source AI is compared to closed-source. The same thing can be done for each class of capabilities (coding, biology…), how the gap grows or shrinks over time, and which events change the catch-up speed (export controls, opsec failures, flagship model releases…).
Defeating entrenching technologies: Slowing or countering the evil twin of countervailing technologies is in itself a countervailing activity. Potential research questions include:
The science of countervailing: Somewhere between cybernetics, engineering, economics, and social work, there is probably insight to be gathered about what countervailing technologies are, how they succeed and fail, and how we can predict whether a certain countervailing technology will achieve its intended goal (and with what second-order effects).
Countervailing technologies and policy: While some countervailing technologies are illegal, many more can benefit from legal protection. Right-to-repair legislation is an example of a countervailing technology being enshrined in law, making it much harder for incumbents to suppress. Directions range from building “politician-friendly” cases for specific countervailing technology to lobbying for specific legal protections.
Developing theories of change: Even if you’re not an expert in building new technologies, you might contribute by determining what technologies would achieve countervailing change and how. It is very likely that the technical solution to your niche issue might not be as infeasible as you might think, and raising enough awareness will attract the attention of those who can tackle it.
Hope for coordination, develop countervailing
Let me be clear: coordination remains the ideal path. A world that could genuinely agree on goals like pacing the frontier, implementing oversight, and sharing the benefits of AI would beat any deployment of countervailing technologies. Working towards that remains my day job. But the gap between an ideal world and the one we have is too large to bet on coordination alone. Powerful actors are often guided more by incentives than principles; we should prepare for such a cynical world.
Maybe you come out of this post thinking that countervailing technologies aren’t worth it. That’s fine! Or maybe you think they can be useful only in very specific and targeted applications. This is reasonable and aligns roughly with my opinion. No matter what, countervailing technologies can be a useful tool in the toolkit and complement more classic policy/technical work.
I should also set some expectations: the realistic result of a countervailing technology is typically that there are still dominant players, though with more restricted power. Part of the reason that music streaming costs ten dollars a month is because of piracy: almost nobody pirates, but the threat of people leaving streaming platforms to return to torrenting permanently limits the rent Spotify or Apple Music can extract. This is a massive victory, but it definitely doesn’t feel like one. A world with smaller AI capability gaps will still look pretty much the same, with powerful countries acting as hegemons, citizens’ rights being violated, and a bunch of coordination problems to be solved. But at least we won’t have as many gap-driven risks.
Personally, I will still keep coordination technology as my main focus, but also work on adversarial interoperability and loyalty benchmarking as my countervailing bets (since they’re close to my research). If you’re interested in contributing to countervailing technologies, my advice would be:
Even just spreading the concept can be powerful, especially with people who feel smart/capable but politically powerless.
As for me, if you have some countervailing ideas, are working on countervailing technologies, or want to share the concept with the broader world, feel free to comment here or email me! I will likely not have enough bandwidth for a collaboration (I have an institute to run, after all), but I can spare an hour of my time for the sake of countervailing-ness.
So, to summarize, look for interventions that are incentive-compatible with the current world: unilaterally rational, too expensive to suppress, free to replicate, diffusing the good kind of power. If things go well, people will adopt your technology, power will diffuse automatically, and no one will even thank you.
Things I'd most like to be argued with about:
Countervailing!
Also a disclaimer: I’m funded by like half of the companies I badmouthed in this post.