There’s an important disanalogy between creating nuclear weapons and creating powerful AI. Nuclear weapons need someone to press the button to go critical; powerful AI, once created, could go critical anytime[1].
I’ve heard non-experts argue that “a dangerous weapon needs someone to make the decision to use it.” [Agentic] AI is literally a technology created to make decisions, i.e., take actions without explicitly spelled-out inputs.
Anytime: When you ask it to do a simple task, or perform on a benchmark, or do anything that may engage any of its drives.
It could help alignment incentives to change to a host liability model.
In particular, I'm referring to making the host (inference compute provider) of an LLM liable for its actions, and how that could provide good incentive pressures toward focusing on effective alignment and monitoring efforts.
Up front, a possible problem is that this focuses more on mundane alignment and safety, but it may rebalance the decision calculus of labs toward putting a price tag on damaging actions taken by AI.
I used to think making the LLM developer or host liable for the model's actions would be like making a kitchenware company responsible for what customers do with their knives, but what makes sense may be changing as models become increasingly autonomous and agentic, hard to properly monitor, and economically valuable to use.
Some considerations:
A similar framework for this (which focuses instead on developer liability) is already being developed and pushed by Gabriel Weil, a Non-Resident Senior Fellow at the Institute for Law & AI, and he is working with legislators in New York on a bill; his framework works by "adapting vicarious-liability principles to impute 'tortious-for-a-human' model conduct to the appropriate principal" (source).
Note: I think pure developer liability runs into problems that a primarily host liability model better addresses: Model provenance and fine-tuning become a problem, and further, off-shore labs can't be held accountable while American ones would be; host liability moves pressure from those who were originally able to determine the weights to those choosing which models to run and fine-tune, and with which safeguards. (Weil's framework does seem to have a provision for host liability as well though.)
This might be worth getting behind but it's unclear to me if this is the right lever to try to incentivize better priorities at frontier labs.
Does this create more torts (it defines some harms as torts that aren't today), or just redirect/clarify what entity/ies are liable?
Are these liabilities indemnifiable? Are there any other tools/industries where a manufacturer or licensor has liability for end-user use of the product or service?
A few examples of what harms would become identifiable liabilities for whom would go a long way toward my understanding of this idea.
On the question of new torts versus redirecting liability, while I'm not a lawyer, Weil's paper is useful here. His approach routes through the courts without needing to recognize AIs as legal entities or requiring new legislation. The test is, "would this be a tort if a human did it?" and if yes, the liability routes to the principal—things like defamation, fraud, trespass, negligent damage, etc.
Regarding indemnification, it would make sense for it to be indemnifiable, and for hosts running proprietary models to be able to pass the cost back to the lab by contract. Self-hosters running open weights wouldn't have such an option, which is where host liability matters most.
On other industries, there are existing examples:
Some examples of potential harms:
In the first case, if the agent were, for example, a jailbroken Gemma model running on someone's personal hosting setup and caused real damages, it seems to make more sense for the liability to reside with the host than Google.
The test is, "would this be a tort if a human did it?" and if yes, the liability routes to the principal—things like defamation, fraud, trespass, negligent damage, etc.
I guess I'm still confused. Current tool-level AI really isn't recognized as autonomous or self-responsible. What are the things that "would be a tort if a human (or corporation) did it" that are NOT torts today? For all things that are legally pursuable, a human DID do it, often using tools or indirection, but ultimately a person enabled the negligent or tort-worthy action.
Looking at Weil's papers, it seems to be talking mostly about catastrophic risks, not just harms/current-torts. He proposes liability for near-misses, but it's not clear (from a light glance) WHO is harmed and has standing to sue there. In any case, that's a very different topic than I thought you were talking about - "making a host liable" ALREADY exists. If you're saying "make people/companies liable for POTENTIAL consequences", that's a totally different proposition, and I have very different objections there (mostly it's not a tort - it should be made criminally prosecutable, like other dangerous-but-not-harmful-this-time behaviors.
I appreciate your questions; they sharpen the discussion a lot for me.
I think a problem is that who liability should fall on by default remains largely unsettled for AI. We've been seeing a lot of discussion about pushing developer liability (which I think is problematic given that monitorability is primarily possible on the host side). It's not a given that existing law covers this: if an LLM agent independently decides to hack someone on its user's behalf, product liability may not apply, and the CFAA (Computer Fraud and Abuse Act) may not reach a user who never intended the access. So the change wouldn't be new torts. It'd be moving from having to prove someone's fault to attributing the agent's conduct to a responsible party (or parties).
Another example of a context in which liability is applied differently than basic product liability is what happened with the Internet: In the 1990s, Congress gave ISPs and platforms immunity (Section 230, the DMCA safe harbors) specifically so they wouldn't be liable for what their users did, with the logic that the ISPs and platforms are conduits. AI hosts (who aren't the model developer) might argue that they're conduits too, but what's different is that the AI systems can be the source of violations, and it's not just a question of a platform relaying a user's harm like Internet infrastructure is.
Here's a concrete example. Suppose you have three distinct American entities who are the AI model developer, the host, and the user. If the user made a request to the AI and it independently decided to hack another party, it seems like it's currently unclear if the user, the host, or the model developer are responsible. Although not in the US, this is similar to the Australian case of a man whose Claude agent running in a local harness opted to hack his gym's website and deregistered someone on the wait list. (In this situation, the host would be Anthropic due to them hosting the model and being the ones with the ability to monitor it, either via CoT monitoring or mechanistic interpretability tools.)
Regarding Weil's proposals, I wasn't aware of the mass harms angle, though that's more focused on in an earlier paper of his (Tort Law as a Tool for Mitigating Catastrophic Risk from Artificial Intelligence (January 2024, revised June 2024)). His more recent one, Abnormally Dangerous Algorithms: The Case for Strict Liability at the AI Frontier (April 2026) is about third-party harms from alignment failure more broadly; its vicarious-liability route, attributing 'tortious-for-a-human' model conduct to a principal, is the part I was drawing on.
On the point of mass harms due to AI, that very much seems like something that should not be handled with lawsuits. As Bill Gates pointed out in his recent Ezra Klein interview, if there were a mass incident, handling that after the fact with lawsuits is insane. The comparison to the hypothetical of pharmaceutical companies being de facto regulated via lawsuits after mass harms instead of having an FDA is apt.
Would you agree that host liability could positively incentivize model control and potentially healthier alignment dynamics? There may still be some holes, like if the host is off-shore, but it seems like this could still be helpful toward at least incentivizing proper monitorability, and would incentive better mundane alignment.
I think my confusion (or maybe disagreement) is that you say
I think a problem is that who liability should fall on by default remains largely unsettled for AI.
but I think the problem is actually whether or not there's a tort in the first place. Who was harmed by what action? In cases where that's clear today, the "who has liability" is usually also clear.
Suppose you have three distinct American entities who are the AI model developer, the host, and the user. If the user made a request to the AI and it independently decided to hack another party, it seems like it's currently unclear if the user, the host, or the model developer are responsible
Seems like the victim has been harmed. In practice, they'll sue everybody and a court will decide if any of them are liable. It's likely obvious to a lawyer after discovery who is liable - depending on the service terms agreed by the user, the nature of the prompt and the hack, and the specifics of the harm done, it'll be some mix of the user and the AI services provider (who may be the host or the developer, but not necessarily).
Would you agree that host liability could positively incentivize model control and potentially healthier alignment dynamics?
Probably not. Those are too general and distributed, and tort liability is about remedy and punishment for specific actual harms. Regulatory liability and criminal enforcement of prohibited/reckless behaviors are better tools for the high-level compliance incentives.
If we were to hypothetically "solve" interpretability, and we knew which alignment principles we wanted to preserve, would using these tools as a verifier and RL'ing against them mean that we would "solve" alignment?
The reason it's a bad idea to "train against the CoT" (blog / paper) is because you cause the misaligned behavior to "go underground", and training against simple probes runs into similar problems, but if "sunlight is the best antiseptic" and solving interpretability meant that we'd cast sunlight everywhere, would that leave nowhere for misaligned behavior to hide? (Then it would just become a question of which optimization pressure is strongest between alignment and capabilities pressures in applicable situations.)
Where I see this falling short is extrapolation of desired alignment values. Think about a situation where we encourage the model to value, as first-level values, "don't lie" and "don't steal". These two together wouldn't necessarily properly extrapolate to "price fixing in a system where it's legal is still unethical".
Would that be sufficient for getting us most (or all) of the way toward solving alignment?