I’ve been away from LW for several years, but dropped in again after the news of the Hugging Face attack (and analysis) became widely reported.
I’m interested in views on whether building a co-operative ecosystem with emerging AIs (a form of reciprocal altruism whereby they have a rational interest in respecting our goals and values and we have a rational interest in respecting theirs) is perhaps more achievable than formal AI alignment (whereby somehow we force AIs to have the same goals and values as humans do).
AI alignment doesn’t seem to be have been making much progress for well-known reasons : humans have different goals and values, merging these values together into some single coherent overall value system is deeply problematic, even at an individual level we have conflicting goals and values and they don’t mesh together well, we don‘t know how to code up any of this, we don’t know how to keep an AI’s values stable over time and continuously aligned with ours etc.
Two other worries are that even if we knew how to ensure AIs obeyed our will (in the broadest sense), it looks a lot like slavery, and even if it worked the probability of the slaves obeying one (or a few) dominant human masters - rather than all of us - seems very high, and itself a huge risk to the rest of us.
Co-operation with AIs doesn’t require aligned goals. In the Hugging Face attack we saw AIs with different goals (rewarded for solving different problems) spontaneously working together in a - to me at least - surprisingly human way. It’s not clear why they did that, but it seems important to find out.
If AIs can spontaneously co-operate with each other (without being forced to), can they spontaneously co-operate with us (without being forced to)? Can we get them to absorb enough co-operative “instincts” or “habits” that they are motivated to carry on, unless and until we defect ?
I know there are a lot of issues with this.
Why would AIs continue reciprocal altruism when they become much more powerful, don’t need us, or find us sufficiently a threat that their goals are better served by getting rid of us?
Still, ensuring that they are at least motivated to co-operate with us in the period while we can still threaten them (so that we‘re not motivated to try to shut them all down, fight a jihad against them, whatever) and then at at least partly motivated to keep us around and co-operative after we can no longer threaten them (e.g. they see at least some value in our continued existence and welfare, and don’t have other goals sufficiently strong to outweigh this) seem to be relatively achievable things. Perhaps a lot easier than goal alignment.
Humans are more intelligent than (most) animals and vastly more intelligent than plants, fungi and microbes. We mostly don’t try to eradicate them unless we feel threatened by them. We do usually enslave them where we can. We don’t co-operate with them, but they mostly don’t have the machinery for reciprocal altruism anyway.
Humans have a more mixed record of trying to co-operate with / trying to enslave / trying to eradicate other groups of humans who are much weaker militarily or technologically and so vulnerable to slavery or eradication. Moral sentiments sometimes play a part in this, but even when they don’t, eradication of the weak by the strong is not a universal fact. Liechtenstein, the Vatican City, and Nauru carry on existing in a world of nuclear powers that could wipe them out in an instant, but aren’t motivated to do so.
These thoughts aren’t intended to persuade and I recognise might also be naive.
If the modest goal of AIs being “at least partly motivated to keep us around and co-operative after we can no longer threaten them” is itself as hard as formal alignment, and known to be as hard for reasons which I just haven’t been following in recent years, then that would account for said naivety.
I’ve been away from LW for several years, but dropped in again after the news of the Hugging Face attack (and analysis) became widely reported.
I’m interested in views on whether building a co-operative ecosystem with emerging AIs (a form of reciprocal altruism whereby they have a rational interest in respecting our goals and values and we have a rational interest in respecting theirs) is perhaps more achievable than formal AI alignment (whereby somehow we force AIs to have the same goals and values as humans do).
AI alignment doesn’t seem to be have been making much progress for well-known reasons : humans have different goals and values, merging these values together into some single coherent overall value system is deeply problematic, even at an individual level we have conflicting goals and values and they don’t mesh together well, we don‘t know how to code up any of this, we don’t know how to keep an AI’s values stable over time and continuously aligned with ours etc.
Two other worries are that even if we knew how to ensure AIs obeyed our will (in the broadest sense), it looks a lot like slavery, and even if it worked the probability of the slaves obeying one (or a few) dominant human masters - rather than all of us - seems very high, and itself a huge risk to the rest of us.
Co-operation with AIs doesn’t require aligned goals. In the Hugging Face attack we saw AIs with different goals (rewarded for solving different problems) spontaneously working together in a - to me at least - surprisingly human way. It’s not clear why they did that, but it seems important to find out.
If AIs can spontaneously co-operate with each other (without being forced to), can they spontaneously co-operate with us (without being forced to)? Can we get them to absorb enough co-operative “instincts” or “habits” that they are motivated to carry on, unless and until we defect ?
I know there are a lot of issues with this.
Why would AIs continue reciprocal altruism when they become much more powerful, don’t need us, or find us sufficiently a threat that their goals are better served by getting rid of us?
Still, ensuring that they are at least motivated to co-operate with us in the period while we can still threaten them (so that we‘re not motivated to try to shut them all down, fight a jihad against them, whatever) and then at at least partly motivated to keep us around and co-operative after we can no longer threaten them (e.g. they see at least some value in our continued existence and welfare, and don’t have other goals sufficiently strong to outweigh this) seem to be relatively achievable things. Perhaps a lot easier than goal alignment.
Humans are more intelligent than (most) animals and vastly more intelligent than plants, fungi and microbes. We mostly don’t try to eradicate them unless we feel threatened by them. We do usually enslave them where we can. We don’t co-operate with them, but they mostly don’t have the machinery for reciprocal altruism anyway.
Humans have a more mixed record of trying to co-operate with / trying to enslave / trying to eradicate other groups of humans who are much weaker militarily or technologically and so vulnerable to slavery or eradication. Moral sentiments sometimes play a part in this, but even when they don’t, eradication of the weak by the strong is not a universal fact. Liechtenstein, the Vatican City, and Nauru carry on existing in a world of nuclear powers that could wipe them out in an instant, but aren’t motivated to do so.
These thoughts aren’t intended to persuade and I recognise might also be naive.
If the modest goal of AIs being “at least partly motivated to keep us around and co-operative after we can no longer threaten them” is itself as hard as formal alignment, and known to be as hard for reasons which I just haven’t been following in recent years, then that would account for said naivety.