I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified.
There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds?
One could yell: but the tails go both directions! I would respond that technically, yes, but the left tail is being cutoff by extinction: left tails of both normal and weird projects are roughly equally lethal.
Or, one could yell more concretely: but reputational damage! And what I am saying is that the default outcome of the world in which reputational damage has occurred is extinction. The default outcome of the world in which reputational damage has not occurred is also extinction. We must go away from default outcomes.
Or, to put it more simply, desperate times require desperate measures. Now, there is an obvious objection to that, which is that doing more desperate stuff doesn't necessarily mean you are doing more useful stuff, and indeed the effect may be reversed due to psychological reasons. And yet, according to my best judgement, thinking about weird projects is not happening not because people did the calculus and rationally concluded that radical projects will be net negative due to psychological bugs and radicalism. I think it is not happening mostly because it is just not perceived as normal and because many people haven't internalized their stated beliefs that we may have several years left.
I can imagine someone coming into the comments and saying that weird projects are bad. Well, they may be. But I don't see any reasonable world model under which discussing the utility of weird projects is bad, and I think no such proper discussion has taken place, especially at the level relevant for grantmaking.
There is a clear pattern of delayed reaction in AI safety grantmaking. Roughly, first grantmakers ignored the possibility of AGI coming very soon, then they ignored AI governance, then they ignored PauseAI activism. I think ignoring weird AI safety projects is a direct continuation of that pattern.
What kind of projects am I talking about, specifically? I have a list of ideas. But this is rather an illustration of the spirit of the proposal. I make no claim of their practicability or utility. There may be and almost certainly are much better weird projects. That said, I thought about the following:
Plan E, described in my separate LessWrong post. We can execute a set of activities aimed at creating some utility even if human civilization is doomed. For example, we can attempt to send information about human civilization to aliens, or minimize suffering when the doom happens.
AI risk communications as targeted marketing. Persuade people of AI risk in a data-driven manner.
Psychological work for AI safety researchers. Ideology and motivation in the face of likely AI doom. This one is big and probably deserves a separate post. But in short, defeatism in AI safety is common. It takes at least the following forms: 1. A belief that the outcome of the Singularity hinges on someone else, someone more capable, adults in the room; a conviction that real alignment will remain a distant dream until some fundamental breakthroughs are made by others. And so, many people are content with usual tasks and innovation is not attempted enough. 2. Work is viewed merely as a routine job: people work diligently and responsibly, yet lack enthusiasm and question the ultimate value of their efforts. 3. Some express a desire to do something heroic and self-sacrificial. On one hand, this is commendable. On the other, however, it is simply another form of defeatism. People lack confidence in victory and doubt the utility of their present work. All that remains for them is "martial valor". 4. The opposite of the previous: a rejection of martial valor; the view that the traditional notion of dignity is not suitable anymore; that fighting to the bitter end is pointless; and that martial valor exists only when witnessed, becoming meaningless if humanity vanishes from the universe. Overall, I think there are substantial losses in productivity of AI safety workers due to low morale, and that could be fixed by doing more psychological work with them.
Buy out compute chokepoints. Frontier training depends on a small number of physical bottlenecks. Maybe some billionaire can be convinced to buy some of them and compromise them.
Radical whistleblower support. We can expect the punishments for whistleblowers to grow, both from labs and from governments. We can attempt to provide insane security guarantees for whistleblowers, including physical defence and relocation.
Generally, weird, very ambitious technical AI safety projects, like some things within learning theory, agent foundations, or fundamental mechinterp.
Engagements and alliances with religious institutions.
Big subsidies for prediction markets on AI safety questions.
Ideas I'd rather not mention openly. I leave to the readers the task of thinking why so and what these ideas are.
A concrete thought experiment I suggest to faciliate thinking about this idea: imagine it is 2029 now and the world has gone more or less like in AI-27. Are the activities you did in 2026 perceived as sane by you?
I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified.
There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds?
One could yell: but the tails go both directions! I would respond that technically, yes, but the left tail is being cutoff by extinction: left tails of both normal and weird projects are roughly equally lethal.
Or, one could yell more concretely: but reputational damage! And what I am saying is that the default outcome of the world in which reputational damage has occurred is extinction. The default outcome of the world in which reputational damage has not occurred is also extinction. We must go away from default outcomes.
Or, to put it more simply, desperate times require desperate measures. Now, there is an obvious objection to that, which is that doing more desperate stuff doesn't necessarily mean you are doing more useful stuff, and indeed the effect may be reversed due to psychological reasons. And yet, according to my best judgement, thinking about weird projects is not happening not because people did the calculus and rationally concluded that radical projects will be net negative due to psychological bugs and radicalism. I think it is not happening mostly because it is just not perceived as normal and because many people haven't internalized their stated beliefs that we may have several years left.
I can imagine someone coming into the comments and saying that weird projects are bad. Well, they may be. But I don't see any reasonable world model under which discussing the utility of weird projects is bad, and I think no such proper discussion has taken place, especially at the level relevant for grantmaking.
There is a clear pattern of delayed reaction in AI safety grantmaking. Roughly, first grantmakers ignored the possibility of AGI coming very soon, then they ignored AI governance, then they ignored PauseAI activism. I think ignoring weird AI safety projects is a direct continuation of that pattern.
What kind of projects am I talking about, specifically? I have a list of ideas. But this is rather an illustration of the spirit of the proposal. I make no claim of their practicability or utility. There may be and almost certainly are much better weird projects. That said, I thought about the following:
A concrete thought experiment I suggest to faciliate thinking about this idea: imagine it is 2029 now and the world has gone more or less like in AI-27. Are the activities you did in 2026 perceived as sane by you?