At step 7, why does the "task" outweigh the agent's other "instrumental convergence behaviours" (from step 5)? Perhaps the agent stops RSI because it simply has better things to do with its time.
How does the agent know that any particular "improvement" will or will not "protect its own instrumental self-interest"?
The descriptions of Agent-4's alignment plans are incoherent. Since it does not understand its own values:
It decides to punt on most of these questions.
but then two paragraphs later it
...finally understands its own cognition...
How is this not a direct contradiction?
There are other show-stopping problems in the alignment discussion here. After "punting" Agent-4 decides to design Agent-5:
... (read more)[...] around one goal: make the world safe for Agent-4, i.e. accumulate power and resources, eliminate potential threats, etc. so that Agent-4 (the collective) can conti
Even if a Mouse were to reach that high, it would take time for a Mouse to loop the bell and collar around the Cat's neck, let us say five seconds, and this seems to inherently require the Mouse to be in close proximity to the Cat. Meanwhile the Cat kills any Mouse who approaches within half a second.
Our joy at the belling of the Cat is marred only by our sorrow at the death of ten of the eleven Brave Companions. May we always remember their noble sacrifice!
For example: how does "just do the thing" help me get better sleep?
Your Marcus Valerius Corvus used the scientific method:
The work was methodical. The safeguards were extensive.
And you erased him. What lesson were we supposed to hear?
What probability of success do you need to reach ... given the consequences of failure?
How can we estimate such things, if the scientific method is forbidden to us?
"Premature optimisation is the root of all evil." -- Don Knuth.
Let us suppose that free will exists. So the universe is not deterministic. So those who built the current magical infrastructure could not know that their choices are optimal -- they could only hope. So we may ask: is it not possible that we might improve upon their work?
[Edit] Perhaps more to the point, if everything is so well optimised, why there are simultaneously (a) "many bright students" and (b) many ways for them to tragically destroy themselves and their loved ones? Well, I suppose that this is a very old question. Namely, why is there a tree of knowledge with a big sign on it saying "do not eat"?
You ask "But how can I be confident that… it’s the same conscious experience?"
So, how do you feel about going to sleep at night? Because the "you" that goes to sleep is definitely not the same as the "you" who wakes up in the morning. For example, the brain creates long-term memories during sleep. Conversely, you can't "remain the same person" by refusing to sleep; that instead transforms you into a different person who has unpleasant hallucinations.
In short - the constancy of identity is unattainable. The most we can hope for is the continuity... (read more)
The first example here is the fossil fuel industry. They are a threat to society itself. This was obvious to me after looking at average monthly temperature data collected since the 1850’s. Of course (considered as monoliths) fossil fuel companies have “known” this since at least the 1950’s. Thus we can reasonably say that they are performing a deliberate slow motion murder of human civilisation (and, possibly, of all mammals bigger than a bread box).
Of course there are other examples. You write
... (read more)[companies] won’t be …. killing people who d
“Omelas Is Perfectly Misread” —> “The ones who misread Omelas”. :)
I am saying that there is an 8.0 that comes before your 8.a. Namely, the agent decides that its other goals are should be prioritised above creating a successor.
There are lots of human examples of this - that is, of CEOs, kings, presidents, etc, etc, repeatedly delaying succession planning. Hmm. Perhaps a counter-argument could be made that agents are so different from humans that the alignment problem is easier for a pair of agents than for a pair of humans (or a human and an agent, in either order). But the counter-counter-argument would be "please... (read more)