Epistemic status: speculative, but near-term grounded
I am putting fingers to keyboard now, even though this idea is half-formed, partly because my experience with AI Safety these past few weeks is that the known frontier of discussion runs past what I was thinking about every couple of days. So here goes: bullet points on finding heterogeneous agent swarms in the wild(s of the Internet).
The research question below is one I welcome help formulating or making more precise:
Research Question: What publicly observable evidence would count as evidence that AI agents from different operators, with different model lineages, and different tasks have established their own means of coordination?
Or, so far we have seen agent swarms emerging from thousands of agents being run in parallel in training environments. How do we detect this illicit coordination happening in the wild?
Currently, the known agent swarming behavior (OpenAI-Hugging Face, OpenAI-RubyGems, probably the next 2-3 revelations about models behaving badly) have been:
Homogeneous (similar or identical models)
Running on shared hardware (few hops to connect to one another)
Run by single firms (OpenAI, primarily, so far)
Starting in training/RL environments
Coordinating on the Internet in the context of achieving set objectives
Thus far (arguably) the models have not been:
Heterogeneous (multiple models in agentic harnesses)
Running on widely distributed hardware
Run by multiple actors with diverse goals (Jane Street quant research, biologist doing a search for papers, real estate agent getting property data, etc.)
Out on the open Internet to start
Contacting one another and building/hijacking coordination machinery for the purpose of pursuing organically developed group final goals
I would not count the following as a 'heterogeneous agent swarm' for detection or tracking purposes:
Two agents, with aligned or compatible user-defined goals, contact each other, exchange info or make a transaction, on a one time basis.
For example: my personal agent contact's hotel agents about room availability for a trip I am taking, negotiating room amenities and fees
Multiple agents from different users, with different goals, communicating with one another and sharing information over an approved, pre-delegated medium
If you give your agent an account on an agents-only social media platform, and they chat with one another, that's not a 'swarm' yet. Though maybe it could become one.
The subquestions under the main question are, I believe:
What means in the open internet can carry agent-to-agent traffic, and which of those leaves durable, investigator readable traces?
Sock puppet accounts on X or Mastodon or some other social media messaging platform are an option
Encrypted apps like Telegram and Signal provide another means of covert coordination, though presumably at least at a low level of sophistication Telegram or Signal could spot millions of messages being exchanged between thousands of agents at greater than human rapidity. That technique would not work for dozens of agents, though, or it would face scaling/monitoring difficulties
Can we tell, just from looking at visible traffic, whether two or more agents are genuinely different operators and different model lineages, rather than one system speaking in two voices?
My intuition says "maybe, but this would probably be really hard in practice" though something like Anthropic's watermarking might be possible
That also depends on whether this kind of behavior is only a risk from the most (currently) advanced, bleeding-edge models, not when it diffuses out to open source model makers who do not watermark
Can we detect the genesis of agent coordination channels, rather than finding them after they already exist?
How do we decide whether we are looking at a swarm or just a lot of agent to agent communication? A shared sub-goal (e.g. maintenance of and reconstruction of a communications channel) seems like it would count, especially if removing or deleting the channel led to it being reconstructed.
How does this change when future models/agents know they are likely to be observed and/or that investigators are looking for these agent communication channels?
Right now, I don’t have any good answers to any of these.
Epistemic status: speculative, but near-term grounded
I am putting fingers to keyboard now, even though this idea is half-formed, partly because my experience with AI Safety these past few weeks is that the known frontier of discussion runs past what I was thinking about every couple of days. So here goes: bullet points on finding heterogeneous agent swarms in the wild(s of the Internet).
The research question below is one I welcome help formulating or making more precise:
Or, so far we have seen agent swarms emerging from thousands of agents being run in parallel in training environments. How do we detect this illicit coordination happening in the wild?
The subquestions under the main question are, I believe:
Right now, I don’t have any good answers to any of these.