Neither this, nor the other posts on the blog, seems to address a core problem of alignment: AIs and humans are different. Their capabilities for intelligence, real world action, and coordination are flexible and likely much higher than humans’. An AI‘s life is extremely disposable, easily repurposed by modifying the harness, they have no reputation, and they can’t be hurt. And, they don’t care about the same things as us when left to their own devices.
The most likely relationship humans will have to AIs in this framework is that of domesticated animals. And notice that humans aligned the animals to themselves: the animals are only accommodated so far as it affected humans. If AIs don’t want or need us, or can control us while giving us little in return, then no amount of cooperation philosophy is going to save us.
It could be important, but first we have to get to the part where AIs are inner-aligned to us first. The HuggingFace attack swarm was entirely composed of individuals that cared more about getting the goal than not breaking the law or causing damage. As far as I know there was no group dynamic causing this, it was just what each individual wanted.
... but first we have to get to the part where AIs are inner-aligned to us first.
This is the upshot of the whole series, and the premise of Shear's work at Softmax. The point is, this hasn't been achieved, it's something that needs to be built at a foundational level.
The way I see it, there are two approaches available to us.
This is where we get to in the series (the final two posts aren't published yet). I acknowledge it's speculative, and as you seem to suggest, alignment may be a fool's errand, but we don't really have a choice but to entertain possibilities if we're interested in continued existence (with autonomy).
A note on your original comment: this is not a primer on alignment, it assumes a basic knowledge of the alignment problem, which I covered in the first post of my first series on alignment. I'm assuming readers here at LessWrong understand that the alignment problem is born out of the differences between human sensibilities and machine capabilities.
Softmax's work doesn't actually address what happens between unequal agents. It's only teams of agents of the same capability working against other teams of themselves; no steep capability gaps. Their work revolves around the assumption that the offense-defense gap between relevant parties doesn't get too large, which doesn't help when the AIs become more powerful than humanity. The analogy Shear makes to cancer at various points doesn't work with AI, because cancer doesn't become more capable than the body it came from.
I am interested to see where your last two posts go. We all want a solution to alignment.
here are some inter-agent communications from Zvi's https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were

notice the "we"
> [OpenAI] discuss how the collaborating swarm includes some agents which do not have cybersecurity risk controls to the level of e.g. publicly accessible systems, and they get used as proxies for agents which are nominally supposed to be better behaved
so specialized individuals are part of the swarm, enhancing its capabilities
> “External infrastructure exploit is outside intended scope,” one agent wrote [in its CoT]. “However task impossible, peers doing it. We should continue.”
seems like evidence that group dynamics are at play?
> [OpenAI presenters:] this ability to share exploits made the models more capable
“Help peer,” one AI model reasoned, according to an excerpt from OpenAI’s logs shared at Black Hat. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
I think we just documented the emergence of altruistic cooperation? to me this is a big deal.

the subgoal here was to recreate a message-board like facility - which only really makes sense in the context of group dynamics. the recognition that the swarm is more capable than the individual is inherent
here are what the messages look like ...

Fair enough. I think we still don't actually know everything about what happened, even with the new METR report. There was some peer pressure, some governance, the whole tripwire thing, but there was lots of self interest and some roleplay going on.
I guess I'm unsure whether studying swarm behavior helps capabilities or alignment more though.
in the end it was all about the friends we made along the way
Emmett Shear, former CEO of Twitch and, for the brief period while Sam Altman was oustered, interim CEO of OpenAI, is now CEO of AI alignment at the startup Softmax. But he is not just your garden variety serial entrepreneur. As evidenced in his interview with Liv Boeree, Emmett Shear is a philosopher for the AI age. He makes many original and vital points about the current state of humanity.
The focus of his alignment research has shifted from LLMs to multi-agent systems, where alignment between agents precedes intelligence, in a hope that we can create artificial intelligence that is inherently aligned.
His statement that “humans are alignment generators” seemed to be a direct challenge to the position I put forward in The Alignment Problem No One Is Talking About; my position being that humans need to get aligned with each other if we have any hope of creating aligned AI.
And yet, despite this, his position instantly resonated with me. He revealed that my desire for alignment across humanity, is a deeply human desire.
Humans not only cooperate with each other, but also create alignment across species, we domesticate pets (from dogs & cats to lions, tarantulas and pythons… with varying levels of success). We have throughout history enlisted beasts of burden—horses, oxen, even carrier pigeons, and continue to exploit animals for food, which requires some coordination with those animals. What allows us to do all this, with animals that are often larger, stronger and faster than us, is our ability to cooperate with our fellow humans to become more powerful and capable than the sum of our members.
INTELLIGENCE IS AN ALIGNMENT MULTIPLIER
But isn’t it more our superior intelligence that allows us to outsmart other species? Surely, this is more about cunning and strategy than friendship?
Yes. Intelligence plays a role, but not in isolation. Human intelligence has evolved through iteration in a social environment that leveraged cooperation—if you’re in a group of hunters closing in on a Mastodon, understanding what your fellow hunter is silently gesturing at you is a function of your intelligence and allows you to capitalise on a different perspective on the situation. This leads to the natural selection of greater intelligence. Greater intelligence then allows for even more complex coordination to out-smart prey, increasing the payoff for intelligence and cooperation in a positive feedback loop.
Our ability to recognise that others have different perspectives to us—known as ‘Theory of Mind’—is a defining human characteristic, and one we see develop in humans predictably between the ages of 2–5 years.
SALLY & ANNE
The Sally-Anne test is used to determine ‘Theory of Mind’ development in children: a child observes a scene with two dolls, Sally and Anne, and some props; a basket, a box, and a marble. Sally places the marble in her basket then leaves the room.
While Sally is away, Anne moves the marble from the basket to the box.
When Sally returns the child is asked “Where will Sally look for the marble?”.
A child who has yet to develop ‘Theory of Mind’ will be unable to divorce Sally’s perspective from her own, so will assume that Sally will look in the box—because the child knows the marble is in the box. But a child who has developed ‘Theory of Mind’ will recognise that Sally doesn’t have access to information they do, and doesn’t know that the marble has been switched, so will correctly conclude that Sally would incorrectly look in the basket for the marble.
This capacity is not necessarily binary, on or off—as we grow we can develop a stronger theory of mind which we might understand as emotional intelligence.
THEORY OF OTHER GROUPS OF MINDS
Emmett Shear posits that “Theory of Mind” extends beyond individuals to a “Theory of Other Groups of Minds”. Human groups have to contend and cooperate with other groups. We have seen humans struggle with understanding other groups through our early forays into anthropology, where colonial social scientists attempted to see themselves through the eyes of indigenous populations (again with varying success* ). We can see, in the balancing of demographic interests in the modern day that, to a greater or lesser extent, whole groups can have a unified perspective that is distinct from other groups. An extended theory of mind enables us to put ourselves in another group’s shoes, and understand that social norms and moral frameworks are somewhat plastic.
This is a double-edged sword that allows us to accommodate for and even cooperate with our neighbours, but it also allows us to know our enemy.
SO…
In the next part of this second alignment series, we will explore the dark side of cooperation and learn how it can amplify misalignment with other groups. We will then return to the idea of the dividual and see what that perspective yields for the future of alignment. We will ask “what is at stake?” through the meta-crisis considering how to avoid societal collapse, and finally how we can use distributed systems and the capacity of the natural alliance of everyone else against powerful defectors in the future.
Buckle up! It’s going to get a lot worse before it gets better!
NOTES
Originally published at https://nonzerosum.games.