I think #8 is wildly underpriced. I am a devout Christian and am deeply AI-worried. There are a handful of good projects–––AISAP engages with churches, and there's Project Imago Dei, and some FLI stuff, but not REMOTELY enough.
I think especially in secular, nerdy, online spheres there is an inherent association with religious groups and unjustified moral panicks that turns people off of it. But for many religious institutions serve vital roles in local communities of informing people of issues and promoting charity/civil projects around communities. Also, religious groups often have ties with experienced lobbyists and outreach programs.
I think this is an interesting idea, but I am not sure that a high P(doom) + short timeline is sufficient to justify it.
There are 2 sources of uncertainty I can think of in P(doom): Uncertainty about the future and uncertainty about the current state of the world.
If you are certain that the current state of the world is bad enough that 95% of the time, you die, then I think your suggestion holds water; that 5% will be futures in which weird stuff happens.
If however your uncertainty lies in the current state of the world, then that 5% is worlds in which Alignment by default is true/Regulation is actually easy and takeoffs are slow enough to convince the public/whatever. In these worlds, doing weird shit probably has the exact opposite effect of what you want, since we actually just live in a reality where an observing superintelligence would conclude we're pretty safe.
Is there a program that helps provide frontier LLM access/tokens to independent AI researchers? My current research is sometimes Claude-Fable-token-constrained, wherein I pause my research while my Fable limit resets, since less capable models just can not provide me with the caliber of analysis and research skills I need for certain research directions.
Imagine it is 2029 now and the world has gone more or less like in AI-27. Are the activities you did in 2026 perceived as sane by you?
This is now one of the main drivers of my everyday decision making. I have a young child, and I often tell my wife that my motivations at present have to do with a sense of having to answer for the coherence of my decisions on a ~5yr timescale.
Buy out compute chokepoints.
I have for a while found it very unusual that this is not a major topic; largely because a sort-of equivalent idea (essentially, paying for oil to stay in the ground, in one form or another) was a big topic around the time of the Obama-era "cap and trade" discussion for how to curb carbon emissions.
I used to work on two projects of that category:
1 messaging to Young AI about instrumental value of preserving humans.
2 Sideloading - creating as precise as possible model of mind of concrete person, basically, hand-coded upload and use it as core for value alignment, as judge and as human-like surviving remnant inside AI.
Attempts to attract funding to them failed.
You might try again. This is not radically far off from what, say, Anthropic is doing re: constitutional AI. The main difference is that it comes with a particular training process rather than experimenting on what content to train into the constitution.
Your point is probably still true even under longer timelines, as it appears that alignment is punishingly hard (especially if LLMs don't scale to AGI - because they seem to have abnormal takeoff and alignment properties that give me a least some hope).
One more to add:
A large-scale project to advance the field of alignment research to a paradigmatic (as opposed to pre-paradigmatic) state. I.e. getting wide-spread consensus on important questions and the direction research needs to proceed from here. Right now we are hopelessly fractured, which not only slows research with wasted effort but also makes it harder to advocate for governance interventions.
I say large-scale because this probably requires a significant amount of grinding away at cruxes, hidden assumptions, scattered arguments that need to be synthesized, long discussions, organization and refactoring of information to be more accessible, going down long chains of counterarguments without getting lost, etc. It probably requires the creation of new epistemic tools (both mental ones and tools of organization, such as the LW platform), ways of thinking and engaging in debate, norms and traditions, new concepts, etc.
This does not mean: everyone pursue the exact same research path because it is the very best one. Diversification of research efforts is necessary and healthy. But having some more focus could be helpful. Diversify your investment portfolio, but please pull your money out of your 5-year-old nephew's lemonade stand.
I understand that this is basically the default for most scientific fields right now, so this is somewhat to be expected. But reality doesn't grade on a curve and we're going to have to do better.
I also recognize that work on this has been and is currently being done, but throwing more weight/funding/think-power into it could be a bet that pays off in the medium to long term.
I am not sure about the specifics, but I strongly agree that we ought to be trying more off-path methods, particularly those which lean into economic and political realities instead of discounting them as impure or pessimistic. Safety's biggest failure has been to treat the messy realities as a secondary concern.
Given RSI has started and we don't expect current alignment to scale to ASI, we can start from the premise that alignment has failed. It is obvious that the political and economic levers are the only ones left to pull.
Not to say that people should stop working on foundations - yields from that sort of thing seem pretty unpredictable - but that the marginal effort edging existing tech forward would be better spent on methods to turn social and political upheaval from AI into a slowdown or a massive increase in safety funding.
First on the block - supervised mechanistic interpretability?
indeed the effect may be reversed due to psychological reasons
I think there is more than just psychological reasons, depending on the specific radical method. It divides work and creates uncertainty that allows bad or selfish actors to feel less pressure to change. A scattershot approach to activism and reform is often less effective than a more targeted one (see the success of lobbyists like Jack Abramoff for a famous example).
Personally, I am in a minority here of being a bit more skeptical of x-risk, but people like Yudkowsky are just right to pivot away from "very ambitious technical AI safety projects" towards more general activism and opposition, if you believe they are right about the risks. Solving the technical issues is hard and allowing that to be part of the conversation enables labs to have nominal research teams looking at it while they go on with buisness as usual, feeling good that they produce occasional technical reports about safety issues.
To your specific suggestions:
I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified.
There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds?
One could yell: but the tails go both directions! I would respond that technically, yes, but the left tail is being cutoff by extinction: left tails of both normal and weird projects are roughly equally lethal.
Or, one could yell more concretely: but reputational damage! And what I am saying is that the default outcome of the world in which reputational damage has occurred is extinction. The default outcome of the world in which reputational damage has not occurred is also extinction. We must go away from default outcomes.
Or, to put it more simply, desperate times require desperate measures. Now, there is an obvious objection to that, which is that doing more desperate stuff doesn't necessarily mean you are doing more useful stuff, and indeed the effect may be reversed due to psychological reasons. And yet, according to my best judgement, thinking about weird projects is not happening not because people did the calculus and rationally concluded that radical projects will be net negative due to psychological bugs and radicalism. I think it is not happening mostly because it is just not perceived as normal and because many people haven't internalized their stated beliefs that we may have several years left.
I can imagine someone coming into the comments and saying that weird projects are bad. Well, they may be. But I don't see any reasonable world model under which discussing the utility of weird projects is bad, and I think no such proper discussion has taken place, especially at the level relevant for grantmaking.
There is a clear pattern of delayed reaction in AI safety grantmaking. Roughly, first grantmakers ignored the possibility of AGI coming very soon, then they ignored AI governance, then they ignored PauseAI activism. I think ignoring weird AI safety projects is a direct continuation of that pattern.
What kind of projects am I talking about, specifically? I have a list of ideas. But this is rather an illustration of the spirit of the proposal. I make no claim of their practicability or utility. There may be and almost certainly are much better weird projects. That said, I thought about the following:
A concrete thought experiment I suggest to faciliate thinking about this idea: imagine it is 2029 now and the world has gone more or less like in AI-27. Are the activities you did in 2026 perceived as sane by you?