We at Gigascale would like to share our new community resources for safety on large groups of agents, in the thousands-to-billions. I've been working on this for a couple of months, drawing on our literature review and research at Gigascale, and I'm pleased to share it in pursuit of a more common understanding of these problems.
The site has a problem definition, org map, a survey, a daily paper scraper, and a weekly events scraper.
Why large agent systems are important
1. They have shown us the first takeover-type event, in the OpenAI Hack. Apparently the emergent capabilities and agency discussed in Multi-Agent Risks are more immediate than other safety areas realised.
2. They overlap heavily with human systems. The first really big systems of agents are economic and social systems, where AI mixes in with and slowly replaces humans. This make the large agent system safety's principal setting for disempowerment, job loss, inequality etc..
What's different about them
Large agent systems are quite different beasts again to smaller-scale multi-agent systems. In particular, despite strong inroads from Cooperative AI and governance communities, the empirical/engineering challenge of large systems expands beyond the current frontier:
* Observability - systems may be observable at interaction + reasoning level (OpenAI), interaction-only (MoltBook), or not at all (plausibly we will soon see agents communicating on end-to-end encrypted forums, or there will be settings where the investigator does not have access to the system).
* Scalable monitoring - Efficient monitoring is becoming a major issue with system size. METR report notes it cost $400k in API credits to investigate the OpenAI incident. And this was with only ~1000 agents.
* More complex co-evolution of agent behaviour with system mechanisms - good reason from economic and social policy (eg baby bonus, prohibition) to think that Goodhearting against system mechanisms can cause higher-order effects