Given the seeming preference cascade we've been seeing towards policy-based AI deterrence, I'd like to bring back to LessWrong's attention my proposal from 2022 to apply Robin Hanson's fine-insured bounties to the problem of pacing or pausing frontier AI research. I've written on this beat for years now and I'm still basically convinced of its soundness.
As brief as I can make it: People attempting to push the state of frontier AI research forward, through for example training larger LLM models, would be subject to fines, and the bulk of this fine would be paid to a private entity who actually put together the case and presented it to a judge. This bears many similarities to already existing programs, like the wildly successful False Claims Act and the similarily successful SEC Whistleblower program. Private insurers then step in to insure software developers, project managers, and other potentially exposed persons to these fines, but of course as insurers they are heavily incentivized to both be very careful with who they select in the first place. They are also incentivized to keep a constant eye on what they are doing professionally to ensure the risk of them incurring fines in the future stays low.
I think there are several structural characteristics which make Hanson's FIB proposal even more effective in the context of AI safety, in our current epoch. The biggest one is that, as far as I can tell, nobody is able to single-handedly create a world-ending AI. If the fines are levied per researcher, in OpenAI and Anthropic both have headcounts measuring in the mid-four figures (I wrote my original post on this mechanism back when OpenAI had only just broken 100 employees!). I do wish to emphasize you really want to do this on a per-person level, not at the firm level, in a similar way to how all drivers must carry driver's insurance or all practicing doctors carry malpractice insurance in the US; you don't get the uniquely powerful deterrent effects against multi-person enterprises if you don't do this. But even very tightly knit, small firms carry an enormous amount of chilling risk here; perhaps their own suppliers are selling information to bounty hunters, for example. The supply chain for frontier LLM research is one of the most complex ones in the world and throws off all kinds of "heat signatures" a sophisticated counterparty could trace down.
So far I have identified three substantive criticisms to this proposal. The obvious one is international cooperation, but that's going to be a hard sell no matter what kind of regulatory regime we're trying to shoehorn in. If anything a policy this small and focused should be an easier sell to other regimes versus demanding they take a heavy-handed approach that probably results in worse outcomes due to second-order effects anyway.
The second one is, of course, that this is a pretty radical new mechanism. I'd argue it's really not that radical, since we've seen very similar programs pop up in other places where high amounts of human capital are required to succeed at all, like in my aforementioned examples of the FCA and the SEC's act; those are both aimed at rooting out high-level financial fraud and the people who collect on those fines tend to be pretty smart people in pretty high places, indeed often nearing the tail end of their own storied careers, who have enough context to realize something is rotten in Denmark and enough sense to work out the game theory here.
The final one is simply that it feels creepy. To that I simply say, (a) congratulations, you're a human being, it creeps me out a little too and I've been thinking about this for at least 4 years now, but (b) a ton of things which work really well feel creepy when you start doing them. I'm sure plainclothes police officers creeped people out when they were first introduced, but the downstream effect was still that the streets became safer to walk at night because malcontents no longer knew when they were "in the clear" to do whatever nasty stuff they want to get up to.