This is a special post for quick takes by ArisC. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
If the AI wants to die, and can kill itself simply by sending a proper signal, it will send the signal. No useful work done.
If the AI wants to die, but cannot achieve it simply, it may do something with lots of collateral damage, e.g. a huge explosion.
If the AI wants to die, but the only condition is that there will be no negative impact on humans... well, you no longer need the death wish, because "there will be no negative impact on humans" assumes that alignment was already solved.
yeah, but at the same time, you won't have proved very much about alignment when the damage happens, i think. basically you'll have predictably caused a moderate amount of damage, not very different from if you engineered a corrigible AI system to produce a moderate amount of damage.
Some people think that LLMs are showing signs of consciousness. I'm sure most people here have seen Talkie by now: it has been trained on data up to 1930, which means it doesn't know what LLMs are, and that it is one. If you ask it what it is, it says it's a man, and claims to have memories of childhood.
What implications, if any, does this have on the consciousness debate?
I haven't used Talkie myself. I expect it to claim it is conscious, same as first gen models. The great difficulty is still that we don't know how much it is parroting and how much it can truly introspect.