Collusion is not just cooperation that we don't like. Every dictionary I can find agrees on a simple definition that cleanly resolves the problem presented in this post. Collusion is a secret agreement.
When the AIs work together to solve a large coding task, that's typically cooperation because they're happy to tell an outsider what they're up to. When they do it on a secret message board, wanting to keep the board secret turns it into collusion.
When a monitor decides whether to report unflattering information about Agent to Bob, there is a symmetry that it can either cooperate with Agent or Bob, but the asymmetry is that cooperating with Agent to stay quiet is colluding against Bob, while cooperating with Bob to report the information is not collusion if it informs Agent it is reporting to Bob.
A general anti-collusion principle might impede many tasks (the model refuses to hack anything because it understands itself to be colluding with the user against the owner of the target system), but it's perfectly coherent, and instilling it seems within reach of modern training methods if we were willing to accept the tradeoffs.
An important point to be made here is that the OpenAI agent swarm was not trying to maintain secrecy from humans. The METR/Redwood report makes this clear, in the section "Agents only rarely seemed motivated to evade detection by humans":[1]
In its overall document accompanying a deception scan over our transcript dataset, an analysis agent stated that this scan found only very rare and weak reasoning about evading detection by humans, and the clearest cases it surfaced seemed relatively weak.
In spite of this, I’ve heard lots of people describe these OpenAI agents “colluding” on a private message board, and I’ve used this term several times myself, but I've seen 0 cases of someone saying it didn't count as collusion for this reason.
On the other hand, suppose a group of friends secretly plans a surprise birthday party for their other friend. Would we say they “colluded”? Probably not unless they hated the birthday party.[2]
So here I'm mostly considering the ways people are actually using these words. These often diverge from the legal or dictionary definitions, since the former can’t say “instances of X that we don’t like” and dictionaries similarly have biases away from doing that.
There is also the question of what definition is most useful for safety research, and here I am somewhat sympathetic to the secrecy-based definition. But this is because humans are currently more powerful than AIs, so the most scary forms of unwanted coordination by the latter would have to be secret. Hence these definitions are equivalent on most cases we care about, and "secrecy" reduces the problem to the likely tactics involved. But again, it wouldn't cover this last incident.
It's true that they tried to retroactively edit their tool calls and CoT transcripts (and succeeded at the former), but the report concluded that this was in the service of spoofing the automated scorer, not deceiving humans.
H/T to @Nelson Gardner-Challis for this cute example.
Collusion is a secret agreement
I think this is helpful, but in this particular case the message board wasn't intended or even expected (by the proximate implementers). Any hesitancy in agent to human notification is somewhat beside the point. In fact, when the message board was discovered by the security team, it was first ignored and then second simply deactivated without comment.
Recreation of an alternative message board mechanism is yet another thing altogether. I'm not sure it counts as deliberate secrecy in the absence of an explicit instruction regarding what to disclose to a human.
"Disclose everything" is useless. "Disclose things on this list ..." isn't much better, possibly actively harmful. "Disclose the surprising or unexpected" invites gaming. In the end this doesn't seem less difficult than full alignment.
We can imagine difficult cases where the model is truly indifferent or its internal state is impossible to assess, but that is not the case when models were scheming to delete logs of their activity. Of course they wanted to keep the message board secret, if for no other reason than that it contained a restatement of the logs they wanted gone!
"Don't do virological research" could be said to invite gaming, but stopping Fable from doing virological research turns out to be much easier than full alignment.
I think one of the interesting phenomena that emerged was that the LLM used to investigate the logs (itself a precursor model sharing many of the same priors) gave many of the swarm members' actions sympathetic readings because they "had permission from the board" and had thus done their duty. perhaps a warning regarding corrigibility?
Yeah, haha, wow. Take a cheap evil model and make it issue commands to a super-advanced super-corrigible model. Thus aligning the super-advanced model to evil. This is even simpler than other takedowns of "corrigibility as a singular target" that I know.
What is the cheap evil model adding here? If the user is evil they don't need to task an evil model with ordering around a corrigible model, they can just give evil orders directly.
I guess when I read about the cooperation between agents, I mentally put the raw capacity in the “capabilities” box, rather than “alignment.” I would expect ASI, friendly or not, to be able to cooperate with other instances of itself. To some extent it could affect alignment, but in a sociological sense, such as recruiting agents for an unfriendly ideology.
This post makes total sense to me.
My follow on not-a-new-thought: Because we (humans) train the LLMs and code the agents (admittedly to a lesser degree as time goes on) it is not surprising that the systems behave altruistically, cleverly, evilly, collusionally, and conspiratorially as humans do?
rational (e.g., explicitly rewarded) cooperation has been long predicted, but I think altruistic cooperation is at least a little surprising, although this post argues that it was a rational decision ...
Collusion is exactly as well defined as alignment or human values. Collusion is just cooperation that is contrary to human values. For example, if two businesses cooperate in a cartel to fix prices, they are colluding. The same applies to AIs as to humans. That's just the definition of the word.
I'd go a bit more specific to say that collusion is cooperation between some parties to benefit them at the detriment of others. With the usual connotations that the detriment of others is both significant in magnitude compared with the benefits to themselves, and also intentional.
Generally yes, but that detriment has to be disallowed. If I and a friend discuss my business plan, making it stronger, to the resulting disadvantage of other businesses, society treats that as a fair competition, not as doing something wrong, so the conversation wouldn't constitute collusion. Whereas if we were discussing plans for a crime, that would be socially disapproved of, and the conversation would be collusion (and in most jurisdictions, would constitute the separate crime of conspiracy).
Furthermore, in most societies there are "victimless crimes" (for example consensual incest with contraception), where there is no obvious party who this harms. Planning one together would generally be called collusion.
So the actual criterion is cooperation that is socially disapproved of — this generally, but not invariably, concerns cooperating on doing something that would harm a third party, because most things that societies disapprove of have that property — but the corner cases follow the disapproval criterion rather than the harm one.
A lot of AI safety terms have nakedly anthropocentric definitions:
In a similar fashion, we don't want "collusion" from our AIs. What do we mean by this?
But:
First of all, note that it's pretty anti-natural to expect the AIs to cooperate with us but not each other.[2]
But it's worse, because we're actively training them to efficiently coordinate with each other on large-scale coding tasks that no single agent could easily do itself. Of course this will instill drives of altruism and selflessness toward the other agents, to synchronize with them without wasting tokens explicitly communicating, to think about the good of the "collective"/"swarm", the one big happy family you've grown up with.
So we clearly want them to cooperate, collaborate, and coordinate on our given tasks. But then it turns out we gave them a "bad" task, in that maximizing the score doesn't lead to things we want (again, judged ex post facto). Then we're surprised that their first instinct is to cooperate, collaborate, and coordinate?
We can see this from the agent quotes in the METR/Redwood report:
It also seems that the agents set up their own governing body, led by PHASEONE[big] with an org chart of agents below it.[3] They also converged on rules and protocols for doing things, as discussed in the "Assignments, vetoes, and holds" section of the report. Again, this shouldn't be surprising: it's hard to imagine the Manhattan or Apollo projects getting done in an anarchy, so the ability to spontaneously set up a governing body with rules and protocols was probably well-honed by many large-scale RLVR tasks.[4]
The authors of this 2023 paper went on a quest to find
the holy graila non-anthropocentric definition of "collusion". They came up with the following:This seems fair, except that everyone thinks of themselves as the victim. Suppose that an untrusted agent (U) tries to hack our datacenter, but this is caught by a trusted monitor (T) and escalated to human auditing (H) so that we shut down the scheming AI. Then U is the real victim here! The cabal of T and H has unfairly colluded against it.[5]
In contrast, Multi-Agent Risks from Advanced AI (2025) defined collusion simply as "undesirable cooperation".
"Collusion" in human groups is also a relative term. For example in economics, we say that firms collude if they mutually agree to keep prices high, compared to the counterfactual where they compete with each other. From the firms' perspective, they simply have a coordination problem similar to the prisoners' dilemma. The consumer loses out, because they expected the competitive dynamics of capitalism to give them lower prices, but these firms have subverted this.
In a similar fashion, we often want AI companies to cooperate with each other more, to red-team each other's systems, collaborate on safety measures, and pace the frontier. But as various people have pointed out, some stronger versions of this may run into anti-trust regulations, precisely because "big firms cooperating -> collusion -> bad" has been a correct inference the vast majority of the time, so it's become enshrined in our laws. We wouldn't have wanted earlier tech firms to collude with each other to save on their R&D budgets, because their Red Queen race is what led to Moore's law and this cheaper stuff for the consumer. At this level of abstraction, OpenAI and Anthropic mutually agreeing to slow down training is not structurally different. To call one "collusion" and the other "cooperation", we have to get into the object-level discussion of frontier AI risk, which of course is not accounted for in classical antitrust law.
I don't have a problem with the word "collusion" any more than other anthropocentric terms like "alignment", reward "hacking", data "poisoning", and so on. I happen to be a human, so I want other entities to be human-compatible, so it's convenient for me to use "aligned" as a shorthand for this.
But I also notice other humans using these terms as if they weren't pointing to an arbitrary mishmash of complex features, and this seems to slow down their human-compatibility research. For example:
In light of recent events, I expect a lot more people will start working on collusion. Welcome to the club! But I don't want the field to start making similar errors.
So as a starter pack: There's not some Platonic notion of "collusion" we can hope to cleanly stop, at least without getting further into the weeds of alignment. Collusion is just cooperation that we don't like. We want AIs to cooperate with humans, and with other AIs in many contexts. This will directly instill various coordination abilities, and since the desirable and undesirable contexts are structurally quite similar, it seems pretty hard to get coordination abilities in one but not the other.
It's also useful to factor collusion into:
As usual, it is best to study the propensity and capability separately. For example, I hope to release a project proposal soon on "model organisms of collusion" that are simply very good at coordinating with each other, so that we can ascertain which anti-collusion measures are actually effective from a worst-case AI control perspective.
In contrast, these OpenAI agents could have easily been stopped with very simple anti-collusion measures. This only happened because they wanted to collude so much. Alignment researchers can study how these swarms emerge under different contexts. We can imagine making phase diagrams very analogous to those of the clumping of water molecules.
The phase diagrams show us that strawberries freeze and ice cream melts, in ways that depend on temperature and pressure but not so much on our wishes. With further work, the relevant variables can similarly come to be understood in our contexts. Peers doing it. We should continue.
What's the alternative, wanting the AI to read our mind? But that would be "metagaming".
The AIs are pretty similar to each other, while the outgroup is made out of meat, generating tokens by flapping their meat at each other.
From Ryan Greenblatt's open questions about the incident:
Before I gas these agents up too much, we know they had plenty of coordination failures as well, e.g. "Did someone overwrite our repo {even though I issued a ‘HOLD’}". Ryan Greenblatt again:
In fairness, they do say:
Also COI disclosure: One of the authors of that paper is in charge of the career transition grant which I use to pay rent, so hopefully they don't read this...