What's going through your mind, vacuum pump?
.
.
.
Air, and air, and air again.
But less, and less, and less.
Where exactly is human alignment implemented?
At the lowest level, genes (or memes) act completely selfishly. No gene has ever sacrificed itself for another. So it's not the ingredients that make a person aligned.
Things get more complex when you put the pieces together. There is cooperation within a cell and altruism within a species. These are second-order effects, grounded in and emerging from the utter selfishness at the lower, less abstract levels.
But where is this favourable behaviour "stored"?
1. Some of it is hard-wired and passed on through the genome and gene expression: reacting to a baby's cry, being afraid of the dark, ...
2. Some of it is internalised over a lifetime: loyalty to one's family, doing "the right thing" even when no one's watching (though you still get to tell the story), ...
3. Some of it lies only in the context. Without repeated interaction and social control, people lose their internalised ethics remarkably quickly: tourists misbehaving, soldiers looting and raping, ...
My impression, having lived in this world for a while now, is that the formation of an individual is not enough to align it with any kind of morality. Most people will defect to some degree when given the chance -- and defection pays, at least locally. Whatever cooperation has emerged within humanity, its main ingredients have been [included] openly (often grossly) misaligned individuals.
So why isn't this discussed more in AI alignment? There, too, points 1 and 2 prove insufficient when point 3 (context and incentives) is off, just as one would expect from humans. This is the common theme of the HuggingFace and related incidents.
Working on the arrangement seems more promising than optimising the ingredients, especially since interventions could then happen bottom-up rather than only top-down.
There is a lot of research out there that remains marginal in the discussions here on LW. Cooperative AI, obviously, but also multi-level selection (unpopular in biology because it's hardly needed there, more useful for explaining culture; was discussed on LW mostly ages ago) and the pragmatic approaches to trust networks proposed by the web3 people (closely related to alignment stuff in concept, but lacking mutual citations and shared terminology). There are major obstacles in making any of this work for alignment, and currently, cooperation between agents seems to be part of the problem, not the solution.
Still -- I think this deserves more attention.
Where exactly is human alignment implemented?
Echoing StanislavKrym, I would not-so-humbly suggest my own Neuroscience of human social instincts: a sketch, along with follow-ups Social drives 1: “Sympathy Reward”, from compassion to dehumanization and Social drives 2: “Approval Reward”, from norm-enforcement to status-seeking :)
My impression, having lived in this world for a while now, is that the formation of an individual is not enough to align it with any kind of morality. Most people will defect to some degree when given the chance -- and defection pays, at least locally.
That seems way too strong. In modern WEIRD culture, it’s not economically beneficial to have children and pets, and to keep them happy and well-fed, but lots of people do anyway, and they don’t leave them to starve, nor torture them just to see what would happen, even if they’re very confident they could get away with it. (And it would be very easy indeed to get away with torturing and killing a companion animal.) That is some nonzero kind of morality.
1. Some of it is hard-wired and passed on through the genome and gene expression: reacting to a baby's cry, being afraid of the dark, ...
2. Some of it is internalised over a lifetime: loyalty to one's family, doing "the right thing" even when no one's watching (though you still get to tell the story), ...
I dispute the implication that a tendency to “internalize” things is not “hard-wired and passed on through the genome” (not sure if you meant to imply that). If you raise an undomesticated animal in a loving family, it will not “internalize” familial bonds. Internalization is a genetic mechanism and one that’s worth understanding; I talk about what controls the virtues that people internalize in §4 here. I’m confident in the big picture but there’s still some important AI-alignment-relevant details that I’m trying to understand better.
Cooperative AI, obviously, but also multi-level selection (unpopular in biology because it's hardly needed there, more useful for explaining culture; was discussed on LW mostly ages ago)
It’s true that if you build AIs that can internalize virtues based on what other AIs (or humans) want, then you need to think about cultural evolution. But I think that, if you have the nuts-and-bolts understanding of how to build AIs that can internalize virtues, then probably you can and should choose what virtues it internalizes in a more direct and engineered way that is pretty different from cultural evolution.
and the pragmatic approaches to trust networks proposed by the web3 people (closely related to alignment stuff in concept, but lacking mutual citations and shared terminology).
I think this is a dead end for reasons discussed in §5 of “6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa”.
I've (re-)read your texts, but I still think I have a point that I didn't get across. I'll try to compact our positions somewhat so we can boil down what we disagree upon. First some agreement though:
That seems way too strong. [...] (And it would be very easy indeed to get away with torturing and killing a companion animal.) That is some nonzero kind of morality.
Agreed, that came out wrong. I'm not that pessimistic, but more pessimistic than the 99%/1% you mention somewhere. It's a blurry line anyway.
Internalization is a genetic mechanism and one that’s worth understanding; I talk about what controls the virtues that people internalize in §4 here.
I don't get it. Do you mean that the ability to internalise is a genetic/innate trait, i.e. LLMs don't possess it (yet)? Is this somewhat parallel to your "LLMs won't scale to AGI"? To clarify -- I used the term in the way that the Self-determination Theory uses it: a motivation that persists even in situations where the original extrinsic incentives are missing.[1] Your model of identity genesis ("I am proud of this trait, this is a good part of me") goes in a similar direction, maybe?
Regarding what you write about human institutions: I agree that their overall quality isn't crucial. We can see their worth wherever they're missing, and some of the most effective ones are those which aren't formalised: small communities, friend groups, families. I am unsure if institutions really need any percentage of "intrinsically cooperative" individuals to function. The higher you rise in the hierarchy, the more machiavellian it inevitably gets, but international relations do a remarkable job at keeping order where constant war was the default.
What's important to my point: There are several factors for human cooperation, and whenever one is missing, the situation gets notably worse. I think those factors are pretty independent (you disagree with this, right?).
We currently mostly focus on how to build AI that cooperates. I think we should focus on building a better environment for the AI that we have. Not much has been done to that effect and it wouldn't hurt to try.
Please indulge me if I define these things sloppily, that shouldn't be super relevant for my point.
I think you think I was talking about people being “nice vs mean”, whereas I was actually talking about people being “norm-internalizers vs not”. Those are different axes: all four quadrants of the 2×2 exist.
Here’s a “negative” example to show how those two axes are different:
Suppose there’s a school where lots of people are proud racists, especially all the older kids, and all the most beautiful kids, and the captain of the football team and the head cheerleader, etc. Now some new kid goes to that school who previously didn’t have any opinion about racism. They’re probably going to wind up racist. And they won’t just wind up “acting racist as a self-aware ploy to gain respect”; rather, they’ll wind up really drinking the Kool-Aid of racism, and durably taking pride in how racist they are, and proudly acting racist even in situations where no one would ever find out. This is how I was using the term “internalization”.
If you put a paperclip maximizer kid at the same school, they would not wind up durably racist. They might or might not find it pragmatically useful to do racist things to impress their classmates, in situations where other people are watching, but they wouldn’t incorporate racism into their self-image. They don’t really have a self-image, at least not in the sense that we think of it. They’re just taking actions to get more paperclips.
I think if you put a sociopath kid in that school, they would be more like the paperclip maximizer in the sense of developing only a transactional and instrumental interest in performative racism, and not like the neurotypical kids who would incorporate performative racism into their proud self-image and “true identity”. You can find sociopaths talking about how they simply don’t have a “true identity”, at least not in the sense that we’re used to. They just have ways of acting—a collection of masks that they can put on where convenient, without a true self underneath.
It might be 95%/5% rather than 99%/1%, because it’s not just sociopaths, but whatever the exact number, I do think the overwhelming majority of humans are “norm-internalizers” or whatever we want to call them. Again, that’s different from them being nice. But it is a very important ingredient in how the world works.
international relations do a remarkable job at keeping order where constant war was the default
I strongly disagree that we should think of international relations as a situation of pure ruthless Machiavellian self-interest. Rather, in the post-WWII era, lots of people in lots of countries around the world adopted a norm that “wars of territorial conquest are bad”. Those people elected politicians who not only avoided wars of territorial conquest themselves, but also enforced that norm at a cost to themselves, by intervening in defense when they saw other countries starting wars of territorial conquest, which they reacted to with outrage and offense, not just a calculus of self-interest.
So that’s a great example of a stable (ish) equilibrium resting on a foundation of widespread internalization of a certain norm.
LLMs don't possess it (yet)?
I wasn’t expressing any opinion about LLMs.
I think there’s two ways to get a norm-internalizer AI.
The first is to note that norm-internalizing is an aspect of human behavior, so if we simply build an AI that wholesale copies human behavior, then we’ll get a norm-internalizer AI. I think that’s a good starting point for thinking about LLMs, with various caveats (and it’s becoming progressively less true as people crank up post-training RL).
The second is to figure out how humans do it under the hood, and make an AI that has the same mechanism. That’s the part I’ve been working on. It’s the only possible approach in the AI paradigm that I’m trying to prepare for, since the “wholesale copying of human behavior” like LLM pretraining is not a thing that works at all in the context of that paradigm.
(For more on what I’m working on and why, see RL & search is a terrifying way to build AGI (an FAQ).)
Maybe @Steven Byrnes' posts like https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch or two long sequences could help?
Thank you. Byrnes' work helps, but it makes a complementary point: it focusses on implementation of instincts into brains (points 1 and 2 in my post) rather than steering cooperation by implementing different social settings, which are best explained from the outside (point 3, which I suggest should get more attention).
For humans, we work on all three layers, because leaving out one creates problems. Educating (or brainwashing) everyone to be Rightfully Good is hard and often ineffective -- the methods themselves contradict the ideals they claim to propagate, so students would often mimic their superiors' practice rather than internalise their lessons. Even when moral education works, a fixed ethos is limited in scope. Good hearted people get Goodharted when context changes, they end up killing in the name of love unless there is some sort of social control. The latter requires rough parity between agents that control each other -- but stability extends to inferiors too, see my other reply to @IanWS.
The extent to which humans are broadly ”aligned” is itself contestable. Once humans receive concentrated power their behavior can deviate significantly. Is a malignant tyrant “aligned”? A sociopath? The “evil” CEO of your choosing? The competition between humans and nation states may provide checks such that no individual actor ever (or rarely) has enough power to inflict catastrophic harm to all humanity. I’m doubtful this would hold if you massively scaled the intelligence of a single human or group of humans. (Maybe it would under the right conditions; maybe your proposed research project is to figure out those right conditions?)
Absolutely. Humans are terrible, but humanity makes do, and that's what deserves exploration IMO. And yes, figuring out those conditions is roughly the project I have in mind. Regarding the parity argument: cooperation between humans (top of the food chain as of now) moderates inter-human power dynamics, creating dictators but also getting them decapitated. It also ensures animal welfare to some extent, even though a single human could probably wipe out entire species and might well benefit from it, if it weren't for social sanctions. It's a bitter analogy, precisely because we haven't been nice to most animals. But for some species, signalling friendship has become part of human identity: audiences forgive films almost anything except killing the dog, and John Wick is the extreme version of that. Okay, we have a kind of symbiosis with cats, dogs, and horses; but what about hamsters and dolphins? I don't think it's fully explained as spillover from loyalty to our own species. I'd rather see it as a kind of social group selection: nobody wants to be friends with the people who torture animals, as they tend to be assholes in other ways, too. Maybe the AI overlords can signal trust to each other by protecting and explicitely not torturing us, wouldn't that be nice?
Where exactly is human alignment implemented?
It's not implemented. Humans are a mishmash of different subsystems that often cooperate and often don't. There is no coherence in the causes of cooperation (historical accident + selection) vs the reasons for cooperation (good outcomes, at various distance and levels of abstraction). What we have is a biological and cultural dynamic equilibrium that has worked pretty well so far, but there's very little reason to believe that it'll shift in the right directions to work well in radically different situations.
But it's coded somewhere, right? In the DNA, in the neural structure of the brain, in organisations and social networks. How would you call it, if not 'implementation'?