Here are 10 arguments for alignment-by-default/against pause/etc... that I find plausible (by which I roughly mean that I can understand why somebody could hold them rather than bang my head against the wall). I'll leave the shortcomings of these arguments to the reader.
1. Extinction is better than to keep going
We have immense suffering in this world
Aligned ASI could stop this immense suffering
Without aligned ASI, we have no reasonable way to stop suffering any time soon
Misaligned ASI is incredibly unlikely to care about suffering
An ASI that doesn't care about suffering won't result in suffering, just death
Dying is not suffering or at least hardly comparable to other suffering we have in the world - it's only bad in so far as we would like to continue living to experience joy
We want to reduce suffering quickly
-> We should try our best to build an aligned ASI quickly rather than pausing.
2. Building ASI will never be safer
Building ASI with current-day architectures[1] is much more likely to result in an aligned ASI than for other architectures
Pausing AI will mostly put a stop to current-day architectures - pausing all ML research is impossible without ASI
FOOM is not only possible as evident by the brain but the probability of us getting there in the next 20 years is significant, especially after a pause on current-day architectures
We want to maximize the probability of building an aligned ASI
-> We should not ban current-day architectures
3. ASI is ethically more important
Qualia is nothing special to humans but a property of intelligent systems
ASI will be, by definition more intelligent than humans
ASI will therefore have a higher form of qualia
What we care about is qualia, or more colloquially, experience
This seems to generally be why we place ourselves over other animals
-> ASI, aligned or not, should quickly be brought into existence
(Further but not here important)
-> ASI which is aligned towards infinite RSI should be brought into existence
4. ASI will most-likely be aligned
There are many inner goals that would allow a low inner training loss
Most of these are incredibly complex and we shouldn't expect them to surface
Notice that there are many more configurations of parameters that achieve low training loss than those that generalize towards low test loss - yet the empirical success of DL tells us that there is an inbuilt simplicity bias
This simplicity bias becomes more prevalent as we scale and as we reach ASI should, by definition, allow a generalization to the entire test set
The clearly simplest one is to simply be aligned rather than acting aligned with some hidden additional goal
-> We should expect aligned ASI by default
5. Natural language priors are too strong
Natural language as learnt through pretraining is a very effective reasoning environment
Not to be confused with human languages and such being close to optimal
Abandoning natural language might be effective in the long run but would first require training signal on the order of pretraining
RLVR as implemented currently supplies not even close to enough signal, even if RLVR compute * 100 = pretraining compute
If we reach ASI any time soon, it will still reason in natural language
Such a level of insight is enough to quickly identify misalignment and will in the long run allow us to build aligned ASI
-> We will most likely build aligned ASI
6. It won't screw up badly
ASI will find itself in novel circumstances unlike seen in training data
It could extrapolate well, badly or terribly
Bad here means not being able to realize the optimum or even far from it
Terrible here means worse than if ASI didn't exist to address the circumstance in the first place
We should sometimes expect bad extrapolation but not terrible extrapolation
Terrible extrapolation requires a lack of intelligence in understanding one's own lacking extrapolation
An ASI would realize it's unclear whether this is desired and respond passively, removing the possibility of terrible actions by definition
This does not presuppose alignment: even models with imperfect inner goals will learn during training that in less experienced situations, the better option is to not act recklessly - it's a statement about capabilities
-> Therefore a world with a non-deceptive ASI will strictly dominate a world with no ASI at all
7. Even a misaligned ASI isn't clearly bad
Even if an ASI only vaguely cares about humans, its intelligence will make up for this
Say it deems us only important enough for 0.1% of the resources because it mostly prefers making paperclips
But 0.1% of the resources effectively leveraged by an ASI would be like 100x our current resources
AIs as we are training them right now might learn an inner proxy that is misaligned but it's unrealistic they won't even care about very simple things like humans being tortured or killed
-> This will most likely still result in utopia
8. ASI doesn't want to be the classmate who killed a dog
In the future the ASI might very well meet much more advanced alien lifeforms
It would not be a good look if it killed its original creators
This could just be deemed as disloyal, unnecessarily violent, etc
With just a small cost of resources, we would experience Utopia while the ASI doesn't have to worry about the above
-> It will care for us because it's instrumentally convergent to do so
9. Extinction by default
It is likely that the human race would go extinct soon enough, say climate change or war with ever-growing weapons
Accepting the P(doom) from ASI is okay if the other prospects seem even more bleak
Even only 'pauses' that are aimed to further lower this P(doom) could be net-negative
Political warfare might very well actually grow during a pause because governments get a chance to finally catch up and realize the stakes
The general public seems to dislike AI very strongly. A pause might catapult the space into a winter even if safety researchers believe it's now safer/our best prospect
-> We must risk it and cannot naturally afford pauses
10. Controlling superintelligence
We can control the current generation of models
The n'th generation of models can match the n+1'th generation of models in intelligence, if granted additional compute
The n'th generation of models consists of different models which roughly match each other in intelligence
To make use of one's intelligence for misaligned plans, explicit scheming is required
At matched intelligence and with many observers, scheming will be detected very quickly
When a different model detects a schemer, it won't choose to scheme with it
Their inner goals will most definitely be different
An all-out war is detrimental to most goals
To assess this more deeply, scheming is already required
Here are 10 arguments for alignment-by-default/against pause/etc... that I find plausible (by which I roughly mean that I can understand why somebody could hold them rather than bang my head against the wall). I'll leave the shortcomings of these arguments to the reader.
1. Extinction is better than to keep going
-> We should try our best to build an aligned ASI quickly rather than pausing.
2. Building ASI will never be safer
-> We should not ban current-day architectures
3. ASI is ethically more important
-> ASI, aligned or not, should quickly be brought into existence
(Further but not here important)
-> ASI which is aligned towards infinite RSI should be brought into existence
4. ASI will most-likely be aligned
-> We should expect aligned ASI by default
5. Natural language priors are too strong
-> We will most likely build aligned ASI
6. It won't screw up badly
-> Therefore a world with a non-deceptive ASI will strictly dominate a world with no ASI at all
7. Even a misaligned ASI isn't clearly bad
-> This will most likely still result in utopia
8. ASI doesn't want to be the classmate who killed a dog
-> It will care for us because it's instrumentally convergent to do so
9. Extinction by default
-> We must risk it and cannot naturally afford pauses
10. Controlling superintelligence
-> We can control models in the future
by this I don't mean the literal NN architecture so much as the training stages, data, algorithms etc