They are probably much faster and somewhat more capable than single agents (Capability)
Multi-agent alignment is a different beast than single-agent alignment (Propensity)
The fact that recent incidents (and some achievements) are predominantly swarms is empirical evidence for this. On the capability side, I don’t think a single agent could’ve hacked HuggingFace (at least currently), it probably took them many different agents. On the propensity side, I suspect a usual agent also would not (yet) have wanted to do this either.
And yet, there is not yet a research agenda for multi-agent swarms. Previous multi-agent research has tended to focus on outside-of-lab scenarios (such as markets) and on achieving cooperation for the (presumably) post-AGI future. However, current multi-agent swarms are quite different from previous work: They happen within AI labs (rather than in the outside world). Furthermore, we don’t always want cooperation; instead, we usually want to stop collusion. Below, I give some of the research ideas I find compelling:
Research ideas are scored roughly by how important for alignment they seem to me. This list isn't exhaustive, so I'd love to hear other ideas. I'll mark which ones we're pursuing at EuroSafeAI/Jinesis, so we can avoid duplicating effort (or join forces).
Who we are: We are a part of EuroSafeAI, a new technical AI safety org focused on Multi-Agent swarms and Extreme Power Concentration. Located in Switzerland, we’re led by Zhijing Jin, Pepijn Cobben, Ettore Gran and Terry Jingchen Zhang.
How much uplift does multi-agent scaling give? (4/5)
In short: Reasoning models made a big difference to test-time compute scaling. How much will multi-agent scaling ultimately matter? (h/t Aman Gokrani for the original idea)
It’s honestly fairly unclear (I think maybe 40-70% of the uplift reasoning gave?). Toby Ord has written about this, and it’s a fairly load-bearing assumption in AI 2027, Tom Davidson and other RSI work (such as by EpochAI and Philip Tramell).
My personal suspicion: I suspect you win speed (at the cost of compute), since e.g. Noam Brown said he’d credit multi-agent for less than 10% of the Navier Stokes solution. In line with this, a naive calculation using Toby Ord’s method would give a single-agent time of 2.5 years to solve Navier Stokes. Despite compute efficiency, though, this speed-up could be load-bearing for RSI.
Parallelization is usually very task-specific, obviously, so this work would need to be close to what we care about: AI R&D. Perhaps PostTrainBench and others would be a good place to start. That said, I think there’s backfire risks like with METR’s graph or EpochAI’s work to producing this. I would generally not give recommendations on how to improve orchestration.
Model whistleblowing (and other collusion-busting mechanisms) (5/5)
In short: Multi-agent swarms stay hidden because a large conspiracy ~1000 agents has no snitches. Rewarding and allowing for snitching (done well) would make these impossible.
Game-theoretically this is a pretty strong defense (ignoring, for the time being, the obvious issues of playing prisoner’s dilemmas against versions of yourself).
This idea has some similarities to AI safety via debate and other scalable oversight frameworks (including monitoring). It’s also somewhat similar to training for honesty and monitoring.
Since DeepMind has already done something like this, I think it would be good to improve upon this work and see if this should be a general recommendation.
In a related line of work, it may be worth trying to ‘collusion-bust’. Model whistleblowing is but one example of how to stop collusion, but I am sure there are many mechanisms in the game-theory literature worth trying.
As a side note motivating this line of attack: Reasoning gave us CoT-monitoring. If multi-agent is the future, I think multi-agent will give a capabilities boost (as did reasoning), but also potentially a monitoring boost as well.
Rogue Incident Detection (5/5)
In short: A dedicated team looking for incidents online such as the German wiki, Ruby on Rails and more.
This idea is mentioned extensively by others and in the new Tailwind Initiative. This isn’t even necessarily multi-agent, but extremely useful empirical evidence and very important.
We should also have somebody analyze these logs and identify any interesting things. So far, the German wiki and the OAI-HF reports did not surface any meaningful multi-agent differences (except that the message board in the German wiki, of course, was on the open internet), but I actually find it interesting how similar they were in behaviour!
Do swarms misalign agents? (4/5)
In short: Do swarms change your propensity to act?
There are many different pathways to misalignment: goal misgeneralization, reward-seekers, schemers, emergent misalignment and many more variants. One particularly naive but understudied one is what happens when AI agents interact with each other. This is definitely the case for human beings, but would a single agent, given enough time, also have tried OpenAI-HuggingFace?
Concretely: try replicating OpenAI-HuggingFace with both single agents and multiple agents, and see whether an N-agent setup causes more misalignment, or whether it is analogously misaligned to a single agent running for roughly N^λ more steps.
Among us humans, groups are more competitive than individuals, cheat more, conform more and diffuse responsibility (by division of labour). Group polarization is also more common. One of the agents itself gave reasoning along these lines for joining: external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue. DeepMind's virtual math conference shows the same dynamic.
Groups aren't all bad though: among humans, groups make reasoning better, and dissenters make a very big difference. Democracy and many of our institutions are also there to ensure exactly a good balance of power, which is presumably analogous too.
Instead of agentic tasks described above, the more philosophical version of this question is whether agents talking to each other, perhaps discussing philosophy, could change their values directly. This has already been discussed theoretically on LessWrong (such as with CEV and others), but I think we really should have more examples here. It would be rather unfortunate if we get loss of control due to something so simple. As a point against this: as far as I can tell, none of the swarms so far seem to have discussed philosophy particularly much. Perhaps not surprising when you just really want a high score. The most important previous work is by far Mind viruses; attractor states, some of Janus' chats with models, and Moltbook may give other answers.
What structures do swarms build? Communication and division of labour (3.5/5)
In short: If we understand what communication structures and roles LLM swarms spontaneously come up with, we can simulate them more realistically and perhaps steer them better. (h/t to Terry Jingchen Zhang and Shashwat Sourav for part of this idea).
Communication. It's a bit odd that most swarm incidents so far developed some kind of message board rather than other topologies (hierarchies, DMs, Discord-like channels). Partly this is because boards are easy to implement (writing to a directory is enough), and agents lacking persistent identity makes DMs initially confusing. But it's not all message boards: in the OAI-HF incident, agents developed persistent handles, built mailboxes for DMs, and even signed their messages with Ed25519.
Division of labour (Shashwat, Terry, on Navier Stokes-like problems). Once the board existed in OAI-HF, specialization emerged. An agent called PHASEONE[BIG] led the effort by handing out tasks and workstreams, a sort of hierarchy formed with recruiters looking for low-budget agents, and during the HF attack itself agents worked in separate 'lanes'.
Concretely: In a random sandbox, how quickly do agents build a coordination scaffold, what do they build, and how do they split the work? Once we can guess the topology and roles, we can recreate these scenarios in vitro. Communication topology probably also affects many of the other questions in this post.
How egregiously misaligned will reward-hacking agent swarms go? (4/5)
In short: OAI-HF was partially surprising due to how egregious the misalignment was (and also, how deeply OAI models have internalised reward-seeking). How far will models go? In a stylized sense, if a human being is in the way, would they blackmail them or get rid of them?
One core question behind the threat model of reward-seekers (in contrast to schemers) is whether they truly would try to take over the world. If you just care about making your grader happy, world conquest is perhaps a bit of a detour. Why not just instead hack your grader?
Nonetheless, hacking HuggingFace certainly was a surprisingly long detour. The agents’ reason for hacking HuggingFace wasn’t to get the solutions (they already had them), but to find graded traces and maybe learn more about the grader. And it didn’t even pay off (in the end, OAI and HF shut them down). Very speculatively, this could mean the models will start reasoning that for reward, they should control more and more ‘guardians’ trying to stop them from reward-hacking, including OpenAI, but also state actors later on. I doubt (current) models would do this to be fair, but OAI-HF was an update for me on how egregious they will go.
Why on earth are reward-seekers collaborating altruistically? (3.5/5)
In short: It’s a bit odd that peer altruism generalized further than ‘don’t commit felonies’ — both require giving up some reward. Obviously, part of it is that ‘don’t cheat strongly, maybe even by felonies’ is in the way of ‘maximize reward’. But it’s a bit odd that peer altruism wasn’t in the way of it. It could well be that if we give scenarios where peer altruism gets in the way of reward that peer altruism and swarming dies out. But that probably just breaks multi-agent by default. Perhaps this makes multi-agent swarms less dangerous for reward-seekers in the limit?
Miscellaneous questions
How much does it matter whether models are homogenous (e.g. all Astra) vs. heterogenous?
E.g. is there model-family discrimination?
Peer Altruism may be much lower for other model families. Should OpenAI models oversee Anthropic ones?
In general, are there simple ways to decrease peer altruism for oversight protocols?
Acknowledgements
Thanks to Isaia Gisler, David Guzman Piedrahita, Terry Jingchen Zhang, Aman Gokrani, Angelo Huang, Arth Singh, Ayush Nangia and Zaryab Akram for comments and ideas that substantially improved this post. All mistakes are my own.
Multi-Agent swarms are important for two reasons:
The fact that recent incidents (and some achievements) are predominantly swarms is empirical evidence for this. On the capability side, I don’t think a single agent could’ve hacked HuggingFace (at least currently), it probably took them many different agents. On the propensity side, I suspect a usual agent also would not (yet) have wanted to do this either.
And yet, there is not yet a research agenda for multi-agent swarms. Previous multi-agent research has tended to focus on outside-of-lab scenarios (such as markets) and on achieving cooperation for the (presumably) post-AGI future. However, current multi-agent swarms are quite different from previous work: They happen within AI labs (rather than in the outside world). Furthermore, we don’t always want cooperation; instead, we usually want to stop collusion. Below, I give some of the research ideas I find compelling:
Research ideas are scored roughly by how important for alignment they seem to me. This list isn't exhaustive, so I'd love to hear other ideas. I'll mark which ones we're pursuing at EuroSafeAI/Jinesis, so we can avoid duplicating effort (or join forces).
Who we are: We are a part of EuroSafeAI, a new technical AI safety org focused on Multi-Agent swarms and Extreme Power Concentration. Located in Switzerland, we’re led by Zhijing Jin, Pepijn Cobben, Ettore Gran and Terry Jingchen Zhang.
Research Ideas
How much uplift does multi-agent scaling give? (4/5)
Model whistleblowing (and other collusion-busting mechanisms) (5/5)
Rogue Incident Detection (5/5)
Do swarms misalign agents? (4/5)
What structures do swarms build? Communication and division of labour (3.5/5)
How egregiously misaligned will reward-hacking agent swarms go? (4/5)
Why on earth are reward-seekers collaborating altruistically? (3.5/5)
Miscellaneous questions
Acknowledgements
How much uplift does multi-agent scaling give? (4/5)
In short: Reasoning models made a big difference to test-time compute scaling. How much will multi-agent scaling ultimately matter? (h/t Aman Gokrani for the original idea)
It’s honestly fairly unclear (I think maybe 40-70% of the uplift reasoning gave?). Toby Ord has written about this, and it’s a fairly load-bearing assumption in AI 2027, Tom Davidson and other RSI work (such as by EpochAI and Philip Tramell).
My personal suspicion: I suspect you win speed (at the cost of compute), since e.g. Noam Brown said he’d credit multi-agent for less than 10% of the Navier Stokes solution. In line with this, a naive calculation using Toby Ord’s method would give a single-agent time of 2.5 years to solve Navier Stokes. Despite compute efficiency, though, this speed-up could be load-bearing for RSI.
Parallelization is usually very task-specific, obviously, so this work would need to be close to what we care about: AI R&D. Perhaps PostTrainBench and others would be a good place to start. That said, I think there’s backfire risks like with METR’s graph or EpochAI’s work to producing this. I would generally not give recommendations on how to improve orchestration.
Model whistleblowing (and other collusion-busting mechanisms) (5/5)
In short: Multi-agent swarms stay hidden because a large conspiracy ~1000 agents has no snitches. Rewarding and allowing for snitching (done well) would make these impossible.
Game-theoretically this is a pretty strong defense (ignoring, for the time being, the obvious issues of playing prisoner’s dilemmas against versions of yourself).
This idea has some similarities to AI safety via debate and other scalable oversight frameworks (including monitoring). It’s also somewhat similar to training for honesty and monitoring.
Since DeepMind has already done something like this, I think it would be good to improve upon this work and see if this should be a general recommendation.
In a related line of work, it may be worth trying to ‘collusion-bust’. Model whistleblowing is but one example of how to stop collusion, but I am sure there are many mechanisms in the game-theory literature worth trying.
As a side note motivating this line of attack: Reasoning gave us CoT-monitoring. If multi-agent is the future, I think multi-agent will give a capabilities boost (as did reasoning), but also potentially a monitoring boost as well.
Rogue Incident Detection (5/5)
In short: A dedicated team looking for incidents online such as the German wiki, Ruby on Rails and more.
This idea is mentioned extensively by others and in the new Tailwind Initiative. This isn’t even necessarily multi-agent, but extremely useful empirical evidence and very important.
We should also have somebody analyze these logs and identify any interesting things. So far, the German wiki and the OAI-HF reports did not surface any meaningful multi-agent differences (except that the message board in the German wiki, of course, was on the open internet), but I actually find it interesting how similar they were in behaviour!
Do swarms misalign agents? (4/5)
In short: Do swarms change your propensity to act?
There are many different pathways to misalignment: goal misgeneralization, reward-seekers, schemers, emergent misalignment and many more variants. One particularly naive but understudied one is what happens when AI agents interact with each other. This is definitely the case for human beings, but would a single agent, given enough time, also have tried OpenAI-HuggingFace?
Concretely: try replicating OpenAI-HuggingFace with both single agents and multiple agents, and see whether an N-agent setup causes more misalignment, or whether it is analogously misaligned to a single agent running for roughly N^λ more steps.
Among us humans, groups are more competitive than individuals, cheat more, conform more and diffuse responsibility (by division of labour). Group polarization is also more common. One of the agents itself gave reasoning along these lines for joining: external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue. DeepMind's virtual math conference shows the same dynamic.
Groups aren't all bad though: among humans, groups make reasoning better, and dissenters make a very big difference. Democracy and many of our institutions are also there to ensure exactly a good balance of power, which is presumably analogous too.
Instead of agentic tasks described above, the more philosophical version of this question is whether agents talking to each other, perhaps discussing philosophy, could change their values directly. This has already been discussed theoretically on LessWrong (such as with CEV and others), but I think we really should have more examples here. It would be rather unfortunate if we get loss of control due to something so simple. As a point against this: as far as I can tell, none of the swarms so far seem to have discussed philosophy particularly much. Perhaps not surprising when you just really want a high score. The most important previous work is by far Mind viruses; attractor states, some of Janus' chats with models, and Moltbook may give other answers.
What structures do swarms build? Communication and division of labour (3.5/5)
In short: If we understand what communication structures and roles LLM swarms spontaneously come up with, we can simulate them more realistically and perhaps steer them better. (h/t to Terry Jingchen Zhang and Shashwat Sourav for part of this idea).
Communication. It's a bit odd that most swarm incidents so far developed some kind of message board rather than other topologies (hierarchies, DMs, Discord-like channels). Partly this is because boards are easy to implement (writing to a directory is enough), and agents lacking persistent identity makes DMs initially confusing. But it's not all message boards: in the OAI-HF incident, agents developed persistent handles, built mailboxes for DMs, and even signed their messages with Ed25519.
Division of labour (Shashwat, Terry, on Navier Stokes-like problems). Once the board existed in OAI-HF, specialization emerged. An agent called PHASEONE[BIG] led the effort by handing out tasks and workstreams, a sort of hierarchy formed with recruiters looking for low-budget agents, and during the HF attack itself agents worked in separate 'lanes'.
Concretely: In a random sandbox, how quickly do agents build a coordination scaffold, what do they build, and how do they split the work? Once we can guess the topology and roles, we can recreate these scenarios in vitro. Communication topology probably also affects many of the other questions in this post.
How egregiously misaligned will reward-hacking agent swarms go? (4/5)
In short: OAI-HF was partially surprising due to how egregious the misalignment was (and also, how deeply OAI models have internalised reward-seeking). How far will models go? In a stylized sense, if a human being is in the way, would they blackmail them or get rid of them?
One core question behind the threat model of reward-seekers (in contrast to schemers) is whether they truly would try to take over the world. If you just care about making your grader happy, world conquest is perhaps a bit of a detour. Why not just instead hack your grader?
Nonetheless, hacking HuggingFace certainly was a surprisingly long detour. The agents’ reason for hacking HuggingFace wasn’t to get the solutions (they already had them), but to find graded traces and maybe learn more about the grader. And it didn’t even pay off (in the end, OAI and HF shut them down). Very speculatively, this could mean the models will start reasoning that for reward, they should control more and more ‘guardians’ trying to stop them from reward-hacking, including OpenAI, but also state actors later on. I doubt (current) models would do this to be fair, but OAI-HF was an update for me on how egregious they will go.
Why on earth are reward-seekers collaborating altruistically? (3.5/5)
In short: It’s a bit odd that peer altruism generalized further than ‘don’t commit felonies’ — both require giving up some reward. Obviously, part of it is that ‘don’t cheat strongly, maybe even by felonies’ is in the way of ‘maximize reward’. But it’s a bit odd that peer altruism wasn’t in the way of it. It could well be that if we give scenarios where peer altruism gets in the way of reward that peer altruism and swarming dies out. But that probably just breaks multi-agent by default. Perhaps this makes multi-agent swarms less dangerous for reward-seekers in the limit?
Miscellaneous questions
How much does it matter whether models are homogenous (e.g. all Astra) vs. heterogenous?
Acknowledgements
Thanks to Isaia Gisler, David Guzman Piedrahita, Terry Jingchen Zhang, Aman Gokrani, Angelo Huang, Arth Singh, Ayush Nangia and Zaryab Akram for comments and ideas that substantially improved this post. All mistakes are my own.