AI is humanity's first through fifth largest problem, but one stands head and shoulders above the rest. Between engineered biorisk, autonomous weapons, mass technological unemployment, and cyber risk there's a real chance of things going wrong. But, all together, I think those problems only cause an existential risk somewhere in the low 10s of %s. Unaligned ruthless superintelligence, on the other hand seems like it would near-certainly cause an existential catastrophe[1] at anything like current levels of alignment theory, and that kind of unaligned superintelligence seems the default outcome of the transition away from AIs trained mostly to mimic patterns in human text towards lots of RL and continuous learning and the capabilities growth from massive investment.
As a result, my grantmaking strategy is focused narrowly on interventions which seem like they might help delay or avert unaligned superintelligence, especially those that increase the odds of aligned superintelligence coming first.
Classes of project I'm interested in
Technical AI safety work that is sufficiently ambitious that it might apply even to strongly superintelligent systems.
Work to improve the chances of automated research systems being used to develop theory which is aiming towards proven alignment properties which aim to hold for strongly superintelligent systems, not just automating incremental empirical work.
Resolution? Strategic red teaming of Anthropic's plan?
Attempts to prepare for worlds where timelines end up longer and there's more time to do ambitious theory work.
Projects which improve the onboarding funnel, including courses, resources, and other materials that help people understand the risks of superintelligent AI (not just AI risk broadly).[3]
Highly competent people who are confused about what to do and will spend time figuring that out and upskilling and developing inside views rather than having a plan in advance.[4]
People who have built useful things before seeking funding
'Science of ML'-flavour theory without a clear backchained path to helping a lot with superintelligence alignment[8]
Academics who didn't evaluate (or disagreed with) the arguments for misaligned superintelligence risk and therefore are not addressing what I see as the core bottlenecks to survival
Undirected / unopinionated / deferring to consensus or Uniformitarian field-building,[9] especially the kinds that end up being largely a pipeline to the labs or fails to teach the big picture strategic situation
Projects which would look weird to low context onlookers.[12]
Mild CoIs, on the level of being friends or having collaborated with the grantee.[13]
Context & me as a grantmaker
A private donor has bought me into Lightcone Commons. I welcome more people boosting my regranting pool,[14] and expect to find good use for significantly more funds than I have been allocated so far.
In my personal grantmaking[15] I have historically preferred to fund people I have seen doing good things in the wild rather than people who have good applications, tended to have at least a short call with most people I fund, positioned myself as a helpful mentor rather than overbearing assessor, and often provided extensive incubation, and networking for projects I have funded.
Please don't pitch me via DMs, instead apply to Lightcone Commons and grantmaking.ai[16]. I'll be searching for the terms 'superintelligence risk' and 'superintelligence alignment' in applications and expect to read any application which uses those words. Also, unless I am allocated further funds, I will mostly be giving out grants in the sub $50k range, so individuals and very small/frugal orgs only.
Evals can give you warnings, but kind of fundamentally can't solve the problems of misaligned superintelligence, have major issues with situational awareness, and are imo massively overinvested in plus have notable capabilities externalities.
Interpretability has an ordering problem afaict, where it opens the black box enough to allow it to be optimised into world-ending capabilities before it is able to open the black box enough to align them, and therefore is mostly harmful to fund. Connor has a great talk about the challenges of using MI for AGI safety. There are rare exceptions, like I think Lucius's work is plausibly the kind of thing that generates critical insights, but I mostly don't expect to fund interp.
Openly acknowledged by the founders of the field of control as not sufficient for superintelligence. Also, I think somewhat likely to backfire horribly in ways I keep meaning to write up.
Given that my previous categories excluded the vast majority of current work, it should not be a surprise that I am not excited about field builders who are going along with the current rather than building strategy formed by strong inside views.
They seem to be not just splitting attention between many AI risks, but specifically not including superintelligence misalignment risk and assuming business as usual on at least some courses.
What I have observed in personal grantmaking doesn't look like an efficient market where opportunities passed up by major grantmakers are often lemons, it looks like a lot of the best low cost opportunities going unfunded while huge amounts of money are poured into high legibility prestigious projects (with an often very questionable sign of impact on superintelligence risk). I think people focusing on adverse selection probably mostly in effect are managing PR in a way I expect to be net negative.
I am not nervously concerned about which stories might look bad to a very low context person, and side with e.g. the people who funded giving out lots of copies of HPMOR to math olympiads as an outreach program, not the people who thought that looked weird and bad.
I do however have a strong filter for people who read as high dark triad or low integrity, because those are correlated with harmful mistakes and negative impact
My life kind of revolves around AI x-risk reduction and people I think are doing good work are often people I want to collaborate with, and therefore end up friends with, and also get to observe for skill and good character over a longer time. However, I will not grantmake to serious CoIs like partners or people who are plausible romantic interests.
I am happy to forward the normal 2% fee on to grantees rather than taking compensation, at least for the first round, as I still have adequate personal runway and have typically used spare money for donations anyway.
I am also happy to share my identity with people who are seriously interested in grantmaking through me, please reach out via LW DMs.
I used to be a crypto millionaire before spending down and donating the vast majority of my wealth over the past 7 years of working full time uncompensated on reducing AI x-risk.
AI is humanity's first through fifth largest problem, but one stands head and shoulders above the rest. Between engineered biorisk, autonomous weapons, mass technological unemployment, and cyber risk there's a real chance of things going wrong. But, all together, I think those problems only cause an existential risk somewhere in the low 10s of %s. Unaligned ruthless superintelligence, on the other hand seems like it would near-certainly cause an existential catastrophe[1] at anything like current levels of alignment theory, and that kind of unaligned superintelligence seems the default outcome of the transition away from AIs trained mostly to mimic patterns in human text towards lots of RL and continuous learning and the capabilities growth from massive investment.
As a result, my grantmaking strategy is focused narrowly on interventions which seem like they might help delay or avert unaligned superintelligence, especially those that increase the odds of aligned superintelligence coming first.
Classes of project I'm interested in
Classes of grantee I'm excited by
Things I am mostly not excited by
Classes of thing I don't consider particularly important
Context & me as a grantmaker
A private donor has bought me into Lightcone Commons. I welcome more people boosting my regranting pool,[14] and expect to find good use for significantly more funds than I have been allocated so far.
In my personal grantmaking[15] I have historically preferred to fund people I have seen doing good things in the wild rather than people who have good applications, tended to have at least a short call with most people I fund, positioned myself as a helpful mentor rather than overbearing assessor, and often provided extensive incubation, and networking for projects I have funded.
Please don't pitch me via DMs, instead apply to Lightcone Commons and grantmaking.ai[16]. I'll be searching for the terms 'superintelligence risk' and 'superintelligence alignment' in applications and expect to read any application which uses those words. Also, unless I am allocated further funds, I will mostly be giving out grants in the sub $50k range, so individuals and very small/frugal orgs only.
Likely on the level of '... there will be no Earth and no biological life, but only a rapidly expanding sphere of darkness eating through the Milky Way as the AI reaches and extinguishes or envelops nearby stars.'
Italic indicates organisations I have donated to with personal funds and/or spent significant time volunteering for.
Specifically, I think the Catastropist-end of funnel building has not been developed anything like as well as the more prosaic side.
Hard to identify with personal connection or having seen them excel in other domains.
Evals can give you warnings, but kind of fundamentally can't solve the problems of misaligned superintelligence, have major issues with situational awareness, and are imo massively overinvested in plus have notable capabilities externalities.
Interpretability has an ordering problem afaict, where it opens the black box enough to allow it to be optimised into world-ending capabilities before it is able to open the black box enough to align them, and therefore is mostly harmful to fund. Connor has a great talk about the challenges of using MI for AGI safety. There are rare exceptions, like I think Lucius's work is plausibly the kind of thing that generates critical insights, but I mostly don't expect to fund interp.
Openly acknowledged by the founders of the field of control as not sufficient for superintelligence. Also, I think somewhat likely to backfire horribly in ways I keep meaning to write up.
This field risks being a potent capabilities enhancer / timelienes shrinker without sufficient gains to alignment to compensate.
Given that my previous categories excluded the vast majority of current work, it should not be a surprise that I am not excited about field builders who are going along with the current rather than building strategy formed by strong inside views.
They seem to be not just splitting attention between many AI risks, but specifically not including superintelligence misalignment risk and assuming business as usual on at least some courses.
What I have observed in personal grantmaking doesn't look like an efficient market where opportunities passed up by major grantmakers are often lemons, it looks like a lot of the best low cost opportunities going unfunded while huge amounts of money are poured into high legibility prestigious projects (with an often very questionable sign of impact on superintelligence risk). I think people focusing on adverse selection probably mostly in effect are managing PR in a way I expect to be net negative.
I am not nervously concerned about which stories might look bad to a very low context person, and side with e.g. the people who funded giving out lots of copies of HPMOR to math olympiads as an outreach program, not the people who thought that looked weird and bad.
I do however have a strong filter for people who read as high dark triad or low integrity, because those are correlated with harmful mistakes and negative impact
My life kind of revolves around AI x-risk reduction and people I think are doing good work are often people I want to collaborate with, and therefore end up friends with, and also get to observe for skill and good character over a longer time. However, I will not grantmake to serious CoIs like partners or people who are plausible romantic interests.
Habryka is happy to onboard people planning to donate at least $50k this coming round, smaller donations can go via ARMF.
I am happy to forward the normal 2% fee on to grantees rather than taking compensation, at least for the first round, as I still have adequate personal runway and have typically used spare money for donations anyway.
I am also happy to share my identity with people who are seriously interested in grantmaking through me, please reach out via LW DMs.
I used to be a crypto millionaire before spending down and donating the vast majority of my wealth over the past 7 years of working full time uncompensated on reducing AI x-risk.
My sponsor has a preference for publicly viewable applications.