I think this makes good points. Definitely we shouldn't talk about p(doom) without an assumption about how we approach making AGI. I think the assumption that we'll just race flat out is breaking down, making more approaches more likely.
But your core point doesn't apply to most people, beyond a little slowdown.
If the the drive takes so long you'll die on the way if you don't drive 200 mph, it's not a mistake to risk it.
Humanity isn't remotely longtermist on average, nor are they utilitarian. They primarily care about themselves, and a few loved ones and friends. So the pitch "we can get there safely in 40 years" means asking most decision-makers (who tend to be over 40) to die and let their parents die. Sure they could save their kids and maybe some of them would if they understood the stakes. But there's a powerful incentive to convince themselves it's not too dangerous to try to save themselves.
This is less common in the rationalist community, but rationality does not demand utilitarianism (at least not clearly).
This tradeoff is also less common in the AI developer community just because they tend to be young.
I realize that's not why most people who oppose slowdown do. I hope we get to a situation where this distinction does become relevant, because that would mean everyone is seeing the situation more clearly. Which would improve our odds of navigating it.
Perhaps this is partly what you get if you promote reasoning in terms of expected values too hard, too generally (instead of "saner", less extremistan-conducive decision rules, or even heuriatics whatever "normal people use").
I don't know how much this has to do being overly neoclassical-econ-brained, but it's a live option for me that it does have non-trivial amount, whether it's in terms of actually skewing people's "honest thinking" or by giving them more concepts to produce palatable justifications for why they are riding the cancerous wave.
I think this post boils down to the fact that you need to compare outcome likelihoods of a plan with outcome likelihoods of other plans (or the counterfactual of not pursuing the plan), not against a hypothetical where "nothing ever happens".
As the post argues, if following a plan implies 30% of utopia and 10% of extinction, whether you actually prefer it or not really depends on what you think the counterfactual outcome distribution is if you don't take the plan (or if you take the best alternative plan instead of this one).
The natural implication is that pessimists should be a lot more likely to take such a plan. To these people, gambling on AI is worth it because it will save everything, otherwise something else bad will happen that will doom humanity with much greater than 10% likelihood! (e.g. getting outraced by less responsible people, WW3, covert AI projects, etc)
In contrast, optimists should probably reject such a plan because they broadly think that better plans will become available the longer they wait, e.g. humanity will eventually figure out better alignment techniques that reduce the extinction risk closer and closer to 0%.
Advanced AI is generally expected to have some very high variance outcomes—it might herald everything good, it might destroy humanity. For instance, here are 800 random AI researchers’ expectations about how good the future is, lined up:
From my 2023 survey
As you can see, most AI researchers put a serious chunk of probability on very different overall outcomes: maybe doom, maybe utopia. This is common. Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things might be insanely good.
I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there’s a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’.
That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there’s only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job will radically improve your life.
The things you should be comparing are driving at 200mph and driving at a normal speed! The things you should be comparing are attempting to attain advanced AI by the current route, and by other routes!
We can debate whether all the other routes are bad or impossible somehow, for instance if constraining projects that risk loss of human control risks sending humanity into an irrecoverable ruin. But I don’t think having ruled out such things is why people are usually thinking in trade-off terms.
Rather I think this error comes from a few things:
If you are bullish on some kind of advanced AI utopia, you should generally be lesskeen to try to achieve it via a careless route that leaves you at high risk of dying and losing it on the way there.