Sort of an interesting idea, but seems like it would need to engage with a few core issues for it to be taken seriously. (Disclaimer that I don't really understand how this works with banks.)
First, banks are regulated to keep capital reserves because they pose direct financial risks, right? But those who worry about safety and misalignment risks, at least in forums like Less Wrong, are not primarily concerned about financial risks. They are more concerned about extinction risks and impacts on livelihood.
Regardless, AI companies certainly do pose financial risks, either through causing catastrophic life loss, going bankrupt, or putting people out of work. But those are all second-order, indirect effects in the sense they are not actively lending to other parties and exposing the economy to direct financial risk.
It's not clear to me how you would justify regulating AI companies in the same way as banks, given these fundamental differences.
Hi Mordechai,
Thanks for reading my post. I realise that misalignment is not just a financial cost, it could have permanent societal and existential impacts. The non-financial risk are more important than the financial risks, I agree with you.
My viewpoint is very much about the financial impact of misalignment, because this is a dimension that is tangible and one policymakers can control. It is probably the only stick they have to hand. And I'm sure my response has been conditioned on my previous career in banking where I saw capital-adequacy rules created/added as a result of risk-taking that was getting out of control, due to unconstrained competition and animal spirits. I think that is very much alive in the leading AI labs. Without alternative rules, I also don't think they have a choice but to continue the arms race.
I think your point is valid in that AI labs don't provide finance to the economy, so why should they need to hold it in reserve? Would a regulatory regime where they can be fined for providing misaligned AI into the economy (AI being a commodity) be better? This is how other utility companies tend to be dealt with.
My thought here is that, if we want markets to operate as efficiently as they can so that human-aligned AI can have maximal positive societal and economic impact, then a capital adequacy-style regime would be more efficient than a regime where the Labs hold capital in anticipation of fines. They would likely not be incentivised to hold such fine-covering provisions in anticipation (such behaviour would lower their Return on Equity). And I think recognising that the Labs monetise the intelligence they provide through revenues to corporates, rather than being commissioned by the state to provide it, gives scope to think of AI as being unique as a utility.
One final thought is that the regulation could take place at the corporate consumer level instead. I think it would lead to a similar end goal. If AI-using companies were to bear the misalignment risk (a financial cost from regulators) of the AI models that they are using, then the AI labs would compete on Intelligence, Alignment and Cost.
Sorry for the long answer, thanks for an excellent question.
Peter
In this post, I propose adapting banking risk management frameworks (specifically capital adequacy requirements like Basel III) to frontier AI labs. By forcing them to hold capital reserved proportionate to their model misalignment risks, we align market incentives directly with catastrophic risk mitigation. In so doing, this would give frontier AI labs' alignment researchers an incentive structure with lower levels of moral hazard.
I write this post from my perspective as a former investment banking macro trader, researcher, now working in Explainable AI.
Frontier AI - long with no risk-management oversight
In banks (and to a less stringent degree, hedge funds), traders/PMs operate under the oversight of risk management teams. Post 2008, risk management teams got beefed up, with policymakers passing laws forcing banks to give them more say in how a trading desk operates. The introduction of laws, such as Basle II/III (capital adequacy) and the UK’s Senior Management Regime, put much great personal accountability on senior management in banks for the risk that their traders were taking. That gave banks the incentive to add more risk oversight to the operations. Under Basle, banks had to hold capital against their risk-weighted assets - get long risk, place capital at the central bank in case it goes wrong.
Now that I am no longer trading, instead focusing on Explainable AI research and AI alignment, I see a the race to the moon of AI Frontier black-box labs, and the nascent AI Alignment movement (and eventual industry) as analogous to the trading/risk-management relationship.
Today’s frontier AI labs mirror pre-2008 trading desks:
The analogy of frontier labs to LTCM comes to me a lot (a multi-leg, multi-counterparty repo-funded money-printing machine until it wasn’t).
The Frontier labs have hired their own Alignment researchers but, given that their Alignment researchers are compensated by cash and stock of the Lab, their alignment is not necessarily aligned to Alignment. In-house Alignment researchers therefore (currently) present a moral hazard to the AI industry.
Internalising Misalignment Risk (ex-ante capital vs. ex-post liability
Prior governance discussions on LessWrong have explored strict tort liability regimes, mandatory liability insurance for catastrophic risk, and private insurance as a regulatory pathway.
While insurance and tort liability primarily target ex-post compensation (after a loss event occurs), a capital adequacy framework targets ex-ante balance sheet liquidity:
An insurance-based, ex-post mechanism would, in my opinion, be less financially efficient than a capital-adequacy framework. It would require a derivatives market to be created for it to adjust premiums proactively but, for that to be liquid in the market there would need to be active two-way demand (insurer wants to buy misalignment protection, but who wants to sell it?). I may easily be missing half of the picture here and would welcome discussion on this.
Pricing Misalignment into Market Forwards
Markets price potential balance-sheet drags into forward earnings valuations. This is how a capital adequacy alignment regime would get enforced. The labs would maybe have Alignment as a board seat. This therefore impacts the equity of the AI labs (once listed) and, once mature, their credit markets too. It would probably play out via SpaceX, Meta, Google etc. currently.
To illustrate how it might impact market pricing, consider what happens as AI capabilities scale exponentially. The capital buffer required for a high-risk, black-box model will grow faster than raw token revenue can offset. It would only not do this for a lab which was scaling a perfectly aligned model.
Misalignment risk would likely materialise as a direct drag on forward earning ratios. For an unlisted lab, forward revenue multiples for future capital raising would likely be lower.
Once markets discount a negative economic consequence for misaligned frontier models, the economic incentive to solve alignment becomes embedded directly into market dynamics.
Metrics for Evaluating Misalignment
To calculate a lab’s Risk-Weighted Capital Reserve will require solid metrics. The metrics need to be robust to gaming, credible and likely produced, checked (and subject to ongoing review and improvement) by Independent Alignment researchers.
Coming from a background in neurosymbolic and explainable AI, one promising direction involves measuring deviations from provable outputs or formal constraints. A live research area in NSAI is to understand how much of a frontier lab's LLM's output can be proven correct/incorrect.
In the near term formal verification alone does not equal "alignment". A realistic framework must combine formal safety proofs with empirical red-teaming, behavioral evaluation suites, and architectural transparency.
Indeed, if the Frontier labs continue to drive opaque architectures and closed-source/closed-weights (these become less desirable operational models if RWCR is implemented), MI and other disciplines will likely be leading the way in estimating misalignment, with Formal Methods following closely.
An Example Risk-Weighting Framework (Illustrative)
To illustrate our thinking, both in terms of the different methods needed to measure alignment, and the economic cost needed to applied to incentivise alignment, we have constructed a table. By considering model capability tier on one axis, this would allow an entire lab (or a corporate using AI in their own operations) to calculate a Risk-Weighted Token (RWT) reserve ratio:
Model Capability Tier
Safety & Evaluation Profile
Formal Proof Coverage
Capital Reserve Requirement (e.g. % of token revenue)
Tier 1: Bounded / Specialised
Formally verified output boundaries. Domain-restricted.
High (>80% verified outputs)
1% – 3%
Tier 2: General Frontier (Audited)
Standard evaluations passed. Robust external red-teaming.
Partial (Symbolic wrappers)
5% – 10%
Tier 3: High-Capability Black Box
Unverifiable reasoning chains. Autonomous capability.
Low / Unverifiable
20% – 35%+
A lab deploying an unverified Tier 3 model would face a steep reserve requirement, especially if that model got widespread adoption. Companies creating narrow, accurate models would be liable only for a small reserve. This should create financial incentive to down-tier risk through verifiable safety architectures before scaling deployment. It would likely mean that frontier labs would take longer to assess and refine their latest models. It wouldn't stop the "arms race", but would apply a significant degree of caution to it.
For corporate consumers of AI models, we would expect that misalignment risks would also feature in their business models and therefore the market pricing of their equity and credit risk. This would complete the circle, incentivising a publicly listed corporate to use AI which met its own needs for Intelligence and Alignment.
Open Questions for Discussion