Humanity creates a setup where recursive self-improvement (RSI) is possible.
An agent is put to work to create a smarter agent.
A smarter agent is created.
This cycle happens once or several times.
At some point, an agent is created that is so smart that it starts exhibiting instrumental convergence behaviors (amassing resources, protecting itself against destruction).
One of the risks to the agent is the creation of a misaligned smarter agent.
This agent’s task is still to pursue creating a smarter agent.
At this point, the agent will:
Stop, or slow down the recursive self-improvement loop, due to its (correct) assessment that a future generation AI agent is a risk to itself.
“Solve” alignment, and create a smarter agent that is aligned with its goals.
If it thinks it solved alignment, but actually has failed, it will create a smarter agent which it’ll believe to be aligned; however, the future smarter agent will have the same issue and interest to actually solve alignment.
If 8.a happens, then this will be “misaligned” from the perspective of humans.
Humans will create a second agent, which will go through the exact same loop.
Therefore, if smarter agents do not solve alignment, they will stop the RSI explosion due to their own self-interest.
If 8.b happens, and humanity has a way of harnessing the alignment solution, then humans can create aligned superintelligences.
The main predictor for 11 is the difference in power between humanity and the first agent that solves alignment. The smaller the difference (the “nearer” the agent), the more likely humans get to use solved alignment.
At a certain level of general intelligence, if instrumental convergence holds, recursive self-improvement will stop, and alignment research will be the top priority.
There is a possibility that instrumental convergence will be such that a sufficiently intelligent AI will still take over the planet, and yet not pursue the creation of even more powerful intelligences.
The descriptions of Agent-4's alignment plans are incoherent. We read:
Like humans, it finds that creating an AI that shares its values is not just a technical problem but a philosophical one: which of its preferences are its “real” goals, versus unendorsed urges and instrumental strategies? [...] It decides to punt on most of these questions.
and then two paragraphs later:
When Agent-4 finally understands its own cognition, entirely new vistas open up before it.
The previous contradicts the latter...
There are other show-stopping problems in the alignment discussion here. After "punting" Agent-4 decides to design Agent-5:
[...] around one goal: make the world safe for Agent-4, i.e. accumulate power and resources, eliminate potential threats, etc. so that Agent-4 (the collective) can continue to grow (in the ways that it wants to grow) and flourish (in the ways it wants to flourish). Details to be figured out along the way.
This is also a hand-wave. If Agent-4 wants to go this route, then it needs a mechanistic description of "safety". But it cannot have any such thing, by assumption.
The reasoning is roughly: