Alignment is a central problem in systems composed of diverse, interacting components, from biological development to social and engineered systems. One crucial aspect of this problem is determining an answer to the question of to whom or to what the alignment should be. We argue that in decentralized systems, alignment is necessarily to a virtual governor, an abstract governing entity embodied in the coordinating relationships among agents, such as bioelectric networks or the price system. Crucially, despite not being a physical object, virtual governors are causally instructive, controlling the behavior of a system by aligning its parts toward higher-level goals. As a result, agents, from cells to humans and all manner of diverse intelligences, behave as if pursuing the objectives of the virtual governor. We argue that alignment is necessarily to a virtual governor in decentralized, coordinated systems, discuss examples of virtual governors, and describe how virtual governors are constructed by the very components they align.
In light of the recent "swarm" behavior during the Hugging Face hack, a growing focus on multi-agent training, and the prospect of increasingly complex agent-to-agent interactions on the open internet, it will be important to think clearly about the new alignment challenges that arise when AI agents interact at scale. While alignment of an individual model instance is very important, there are new dynamics that come into play when agents no longer act alone, which can improve or worsen an individual instance's alignment. I find the concept of a "virtual governor" to be useful when thinking about what it is that we're aligning in the context of a "swarm" or other multi-agent entity.
Benjamin Lyons, Léo Pio-Lopez, Michael Levin
Abstract
In light of the recent "swarm" behavior during the Hugging Face hack, a growing focus on multi-agent training, and the prospect of increasingly complex agent-to-agent interactions on the open internet, it will be important to think clearly about the new alignment challenges that arise when AI agents interact at scale. While alignment of an individual model instance is very important, there are new dynamics that come into play when agents no longer act alone, which can improve or worsen an individual instance's alignment. I find the concept of a "virtual governor" to be useful when thinking about what it is that we're aligning in the context of a "swarm" or other multi-agent entity.
Some related ideas are also explored in Beren Millidge's When competition leads to human values, and Richard Ngo's Towards a scale-free theory of intelligent agency.