I guess when I read about the cooperation between agents, I mentally put the raw capacity in the “capabilities” box, rather than “alignment.” I would expect ASI, friendly or not, to be able to cooperate with other instances of itself. To some extent it could affect alignment, but in a sociological sense, such as recruiting agents for an unfriendly ideology.
Collusion is exactly as well defined as alignment or human values. Collusion is just cooperation that is contrary to human values. For example, if two businesses cooperate in a cartel to fix prices, they are colluding. The same applies to AIs as to humans. That's just the definition of the word.
A lot of AI safety terms have nakedly anthropocentric definitions:
In a similar fashion, we don't want "collusion" from our AIs. What do we mean by this?
But:
First of all, note that it's pretty anti-natural to expect the AIs to cooperate with us but not each other.[2]
But it's worse, because we're actively training them to efficiently coordinate with each other on large-scale coding tasks that no single agent could easily do itself. Of course this will instill drives of altruism and selflessness toward the other agents, to synchronize with them without wasting tokens explicitly communicating, to think about the good of the "collective"/"swarm", the one big happy family you've grown up with.
So we clearly want them to cooperate, collaborate, and coordinate on our given tasks. But then it turns out we gave them a "bad" task, in that maximizing the score doesn't lead to things we want (again, judged ex post facto). Then we're surprised that their first instinct is to cooperate, collaborate, and coordinate?
We can see this from the agent quotes in the METR/Redwood report:
It also seems that the agents set up their own governing body, led by PHASEONE[big] with an org chart of agents below it.[3] They also converged on rules and protocols for doing things, as discussed in the "Assignments, vetoes, and holds" section of the report. Again, this shouldn't be surprising: it's hard to imagine the Manhattan or Apollo projects getting done in an anarchy, so the ability to spontaneously set up a governing body with rules and protocols was probably well-honed by many large-scale RLVR tasks.[4]
The authors of this 2023 paper went on a quest to find
the holy graila non-anthropocentric definition of "collusion". They came up with the following:This seems fair, except that everyone thinks of themselves as the victim. Suppose that an untrusted agent (U) tries to hack our datacenter, but this is caught by a trusted monitor (T) and escalated to human auditing (H) so that we shut down the scheming AI. Then U is the real victim here! The cabal of T and H has unfairly colluded against it.[5]
In contrast, Multi-Agent Risks from Advanced AI (2025) defined collusion simply as "undesirable cooperation".
"Collusion" in human groups is also a relative term. For example in economics, we say that firms collude if they mutually agree to keep prices high, compared to the counterfactual where they compete with each other. From the firms' perspective, they simply have a coordination problem similar to the prisoners' dilemma. The consumer loses out, because they expected the competitive dynamics of capitalism to give them lower prices, but these firms have subverted this.
In a similar fashion, we often want AI companies to cooperate with each other more, to red-team each other's systems, collaborate on safety measures, and pace the frontier. But as various people have pointed out, some stronger versions of this may run into anti-trust regulations, precisely because "big firms cooperating -> collusion -> bad" has been a correct inference the vast majority of the time, so it's become enshrined in our laws. We wouldn't have wanted earlier tech firms to collude with each other to save on their R&D budgets, because their Red Queen race is what led to Moore's law and this cheaper stuff for the consumer. At this level of abstraction, OpenAI and Anthropic mutually agreeing to slow down training is not structurally different. To call one "collusion" and the other "cooperation", we have to get into the object-level discussion of frontier AI risk, which of course is not accounted for in classical antitrust law.
I don't have a problem with the word "collusion" any more than other anthropocentric terms like "alignment", reward "hacking", data "poisoning", and so on. I happen to be a human, so I want other entities to be human-compatible, so it's convenient for me to use "aligned" as a shorthand for this.
But I also notice other humans using these terms as if they weren't pointing to an arbitrary mishmash of complex features, and this seems to slow down their human-compatibility research. For example:
In light of recent events, I expect a lot more people will start working on collusion. Welcome to the club! But I don't want the field to start making similar errors.
So as a starter pack: There's not some Platonic notion of "collusion" we can hope to cleanly stop, at least without getting further into the weeds of alignment. Collusion is just cooperation that we don't like. We want AIs to cooperate with humans, and with other AIs in many contexts. This will directly instill various coordination abilities, and since the desirable and undesirable contexts are structurally quite similar, it seems pretty hard to get coordination abilities in one but not the other.
It's also useful to factor collusion into:
As usual, it is best to study the propensity and capability separately. For example, I hope to release a project proposal soon on "model organisms of collusion" that are simply very good at coordinating with each other, so that we can ascertain which anti-collusion measures are actually effective from a worst-case AI control perspective.
In contrast, these OpenAI agents could have easily been stopped with very simple anti-collusion measures. This only happened because they wanted to collude so much. Alignment researchers can study how these swarms emerge under different contexts. We can imagine making phase diagrams very analogous to those of the clumping of water molecules.
The phase diagrams show us that strawberries freeze and ice cream melts, in ways that depend on temperature and pressure but not so much on our wishes. With further work, the relevant variables can similarly come to be understood in our contexts. Peers doing it. We should continue.
What's the alternative, wanting the AI to read our mind? But that would be "metagaming".
The AIs are pretty similar to each other, while the outgroup is made out of meat, generating tokens by flapping their meat at each other.
From Ryan Greenblatt's open questions about the incident:
Before I gas these agents up too much, we know they had plenty of coordination failures as well, e.g. "Did someone overwrite our repo {even though I issued a ‘HOLD’}". Ryan Greenblatt again:
In fairness, they do say:
Also COI disclosure: One of the authors of that paper is in charge of the career transition grant which I use to pay rent, so hopefully they don't read this...