TLDR: I'm managing a new fund, housed at Lightcone Infrastructure, that will award at least $200,000 in grants and prizes for corrigibility research in 2026. Roughly half will go to traditional grants (first application deadline August 23rd) and half for prizes recognizing excellent work done this year. If you have...
Look, I'm as much of a Rationalist with a special interest in AI x-risk as anyone. But oh my god do I hate talking about "P(doom)". When it gained widespread usage in the wake of ChatGPT, I assumed that it was floating around variously adjacent circles of faux-intellectuals, but surely...
(...but also gets the most important part right.) Bentham’s Bulldog (BB), a prominent EA/philosophy blogger, recently reviewed If Anyone Builds It, Everyone Dies. In my eyes a review is good if it uses sound reasoning and encourages deep thinking on important topics, regardless of whether I agree with the bottom...
Last year I wrote the CAST agenda, arguing that aiming for Corrigibility As Singular Target was the least-doomed way to make an AGI. (Though it is almost certainly wiser to hold off on building it until we have more skill at alignment, as a species.) I still basically believe that...
Is focusing on corrigibility our best shot at getting to ASI alignment? Max Harms and Jeremy Gillen are current and former MIRI alignment researchers who both see superintelligent AI as an imminent extinction threat, but disagree about Max's proposal of Corrigibility as Singular Target (CAST). Max thinks focusing on corrigibility...
This post is a (somewhat rambling and unsatisfying) meditation on whether it's possible, given a somewhat powerful AI that is more or less under control and trained in a way that it behaves reasonably corrigible in environments that resemble the training data, whether one could carefully iterate towards a machine...
> It's plausible that humanity could make a corrigible ASI by 2035 if the planet was united around that goal and being very careful. Are there any knowledgeable people outside MIRI who might disagree with me on this statement and be interested in arguing with me about it? I'm more...