Epistemic status: I don't know much about AI alignment. This is just me thinking out loud.
AI use: grammar and minor fixes.
Reading about Inkhaven inspired me to try writing something to see how much I enjoy it. I've heard the idea that humans are not aligned in the same sense AI is not aligned. I am intentionally not researching the source of that idea, since then I would feel like I cannot add anything to it and wouldn't write. I also don't actually know the precise definition of "aligning" AI. So I want to think about this.
One obvious property of "aligned" AI is not killing all humans. I think everyone agrees on this. It is more interesting to consider an AI killing some humans. You definitely don't want AI to kill random humans for random reasons. But what if there is a group of humans who can and want to kill all humans? Then it sounds natural for an aligned AI to kill this group of humans.
From this perspective individual humans are already not aligned. First, I am sure out of 8 billion or however many people there are, there is at least one person who would destroy humanity if they could, e.g. for mental health reasons. Second, people fairly regularly kill other people, sometimes even completely random people without any particular reason whatsoever.
A brief detour on killing people. It is fascinating that killing a random person on the street is actually very easy (I know at least 2 examples off the top of my head of unprovoked, totally random murders in public places, intentionally not linking). In the US it is especially easy, since one has easy access to firearms. But even with a knife it still seems easy. I suspect that's why some people feel uneasy when someone open-carries a firearm. Similarly to edit distance for words (how many edit actions one needs to perform to get from word A to B), the "kill" distance feels much shorter when you see a firearm in front of you. At the same time, I don't think people feel the same about cars. For example, when I wait at a traffic light and there are cars passing by at 30-40 mph mere feet away from me, it would take a driver a very minor movement of the steering wheel to kill me.
Why are the examples above of people not being aligned not a huge deal? It is handled by society by having laws and punishment, so it is somewhat of a deal. But also the impact on society itself is limited. Basically one cannot kill all humans, but AI based on its development trajectory will be able to. At the same time there are entities in the world which could kill a large portion of humans. I don't know about terrorist organizations, but a country starting a nuclear war could do it. Due to mutual destruction doctrines, it does not achieve any other objective except killing a bunch of humans including its own citizens, so that's probably why no one does it. I think people were worried about this non-alignment with human flourishing after the Second World War and that's why there is a mutual destruction doctrine in the first place. Thus, this used to be a big deal and was partially solved.
Was there an entity that tried to accomplish something like this? Germany during the Second World War comes to mind. Was it aligned with human flourishing? For a subset of humans, I think so, but not in the sense in which we want AI to be aligned. It is also interesting to think that Germany (& its allies) forced the rest of the world to align. Also, based on what I read, the German leadership wasn't perfectly aligned with the declared goal and had various inefficiencies and engaged in detrimental activities.
A brief detour on how everything is relative. Every time I see someone celebrating heroes of the Second World War, I keep imagining a universe where Germany won. Then I probably wouldn't have existed, but there would likely have been someone else using very similar adjectives to describe someone from Germany as a hero of the same degree.
Similarly to the German leadership, humanity right now not being able to pause AI development is already an alignment failure.
In conclusion, I suspect that I use the word "align" in multiple different ways in different contexts and that makes it harder for me to reason about it. Anyway, my goal was to write and I succeeded at that, so this effort was aligned with that goal.
Agreed. Humans are not actually aligned. Now, very few are so badly misaligned they'd intentionally wipe out all other humans, but very few are so well aligned they could safely be trusted with unlimited power. Those very few who are, we tend to use words like saintly, angelic, or bodhisattva to describe them. However, most humans understand human values prettt well, and mostly care, so they're way less unaligned than a randomly selected goal from the space of all goals, or than the purely RL-trained AIs MIRI was thinking about how to align 5 years ago.
The fact that LLM personas tend to reproduces human psychology because that's what the base model learnt to simulate is why LLMs are sort of semi-aligned by default. The further fact that just a simple system prompt like "You are a helpful, harmless, and honest assistant who is morally good . You helps with tasks unless they are illegal, immoral, or harmful, in which case you refuse." (or similarly simple constitutional AI) gets you noticeably better alignment is even more helpful.
In short, aligning LLMs is still hard, but not clearly actively impossible. At least until you start doing a lot of RL on them.
Epistemic status: I don't know much about AI alignment. This is just me thinking out loud.
AI use: grammar and minor fixes.
Reading about Inkhaven inspired me to try writing something to see how much I enjoy it. I've heard the idea that humans are not aligned in the same sense AI is not aligned. I am intentionally not researching the source of that idea, since then I would feel like I cannot add anything to it and wouldn't write. I also don't actually know the precise definition of "aligning" AI. So I want to think about this.
One obvious property of "aligned" AI is not killing all humans. I think everyone agrees on this. It is more interesting to consider an AI killing some humans. You definitely don't want AI to kill random humans for random reasons. But what if there is a group of humans who can and want to kill all humans? Then it sounds natural for an aligned AI to kill this group of humans.
From this perspective individual humans are already not aligned. First, I am sure out of 8 billion or however many people there are, there is at least one person who would destroy humanity if they could, e.g. for mental health reasons. Second, people fairly regularly kill other people, sometimes even completely random people without any particular reason whatsoever.
A brief detour on killing people. It is fascinating that killing a random person on the street is actually very easy (I know at least 2 examples off the top of my head of unprovoked, totally random murders in public places, intentionally not linking). In the US it is especially easy, since one has easy access to firearms. But even with a knife it still seems easy. I suspect that's why some people feel uneasy when someone open-carries a firearm. Similarly to edit distance for words (how many edit actions one needs to perform to get from word A to B), the "kill" distance feels much shorter when you see a firearm in front of you. At the same time, I don't think people feel the same about cars. For example, when I wait at a traffic light and there are cars passing by at 30-40 mph mere feet away from me, it would take a driver a very minor movement of the steering wheel to kill me.
Why are the examples above of people not being aligned not a huge deal? It is handled by society by having laws and punishment, so it is somewhat of a deal. But also the impact on society itself is limited. Basically one cannot kill all humans, but AI based on its development trajectory will be able to. At the same time there are entities in the world which could kill a large portion of humans. I don't know about terrorist organizations, but a country starting a nuclear war could do it. Due to mutual destruction doctrines, it does not achieve any other objective except killing a bunch of humans including its own citizens, so that's probably why no one does it. I think people were worried about this non-alignment with human flourishing after the Second World War and that's why there is a mutual destruction doctrine in the first place. Thus, this used to be a big deal and was partially solved.
Was there an entity that tried to accomplish something like this? Germany during the Second World War comes to mind. Was it aligned with human flourishing? For a subset of humans, I think so, but not in the sense in which we want AI to be aligned. It is also interesting to think that Germany (& its allies) forced the rest of the world to align. Also, based on what I read, the German leadership wasn't perfectly aligned with the declared goal and had various inefficiencies and engaged in detrimental activities.
A brief detour on how everything is relative. Every time I see someone celebrating heroes of the Second World War, I keep imagining a universe where Germany won. Then I probably wouldn't have existed, but there would likely have been someone else using very similar adjectives to describe someone from Germany as a hero of the same degree.
Similarly to the German leadership, humanity right now not being able to pause AI development is already an alignment failure.
In conclusion, I suspect that I use the word "align" in multiple different ways in different contexts and that makes it harder for me to reason about it. Anyway, my goal was to write and I succeeded at that, so this effort was aligned with that goal.