I think an important concept for sensibly modeling a part of the world is what I call a "macroagent", which abstracts over the similarity cluster of teams, companies, AI scaffolds, countries, the market, and the scientific community.
In this post, I will lay out the basic ontology for modeling macroagents.
In the next post, I will lay out an ontology of design lenses through which macroagents can be optimized. (In a sense this can be seen as mostly rearranging ideas from economics into a good ontology.)
I might write more posts in this mini-sequence at some point, though the other things I want to write here have lower priority, so not sure whether I will get around to it.
I think the frames I lay out here can be useful for shaping both human systems and AI societies to improve their alignment and epistemics.
1. The Macroagent Ontology
1.1 What do I mean by "macroagent"?
A macroagent is an optimization process with agentic subsystems.
Examples:
A team in a company: Optimizes for achieving goals the team is responsible for. Agentic subsystems are the humans in a team.
AI scaffolds: Analogous to (1).
A company: Usually optimizes for maximizing long-term shareholder value, although not always exclusively. Agentic subsystems are departments, teams, and humans.
The market as a whole: Optimizing for producing stuff that people will pay for. Agentic subsystems are companies and humans.
By "AI macroagent" I mean a macroagent that bottoms out with AIs as fundamental agentic unit, so e.g. a company consisting only of teams of AIs would be an AI macroagent. I use "human macroagent" analogously. There are also hybrid macroagents, e.g. a team consisting of both humans and AIs, which are also of interest to me here.
(Btw, on my terminology, macroagents are usually only agentic systems and not agents themselves (though it is plausible that there's no crisp boundary and that macroagents with extremely high internal coherence are agents).[1])
1.2 The type signature of a macroagent
A macroagent can be described by these 3 primitives: (memory, mechanisms, agentic subsystems)
Memory: Information that is shared across at least a significant fraction of the macroagent's agentic subsystems due to a reason that is related to the macroagent's existence (rather than just coincidence).
Mechanisms: The structures and rules governing subsystem interactions, and how they are enforced.
Agentic subsystems: See section 1.1.
The programming analogy is: Memory is the database, mechanisms are the code.
Here are 4 examples to clarify what counts as memory and mechanism:
Macroagent
Memory
Mechanisms
Agentic subsystems
Ideal free market
Who owns what (property, money); posted prices; contracts (including a history of transactions)
Ownership allocations must be consistent with initial state + executed contracts; how correct ownership allocations and contracts are enforced
Individuals and companies
A car manufacturer
Internal documents and databases (product designs, financial records, customer data)
Reporting lines (who reports to whom, who decides what); hiring and firing (interview processes, performance reviews); budget allocation (which departments and projects get how much money)
Departments, teams, employees
The scientific community
Papers and textbooks; datasets
Peer review; authorship and citation norms; grant allocation; Implicit through journal expectations etc: The scientific method
Researchers, research groups, journals, funders
A country (ignoring the market that's part of it)
Information informing government decisions and records thereof
The law; police procedures; decision hierarchies in government institutions; court procedures; elections; ... (one could go into more fine-grained detail)
Citizens; government institutions (agencies, courts, …)
Some clarifications:
Memory doesn't necessarily mean it's written down. In small teams, information that was discussed at a meeting can also count as macroagent memory. In communities, reputation of people can be part of memory.
Mechanisms can be formal (like law or code) or informal (like cultural expectations). They can also leave a lot of freedom to the agentic subsystems (e.g. a hiring process determines who decides whether someone gets hired, but not how they decide).
Two special kinds of mechanisms are knowledge aggregation and value aggregation mechanisms, which I'll discuss further in the next post.
1.3 Incentives
From the primitives we can derive the macroagent's incentives, at two levels of granularity:
Incentive profile: The value differentials of actions faced by each individual agentic subsystem. I.e. for each action available to a subsystem, how good the expected consequences are according to that subsystem's values. E.g.: "Eva is a researcher who is incentivized to try something novel rather than replicate an important result."
Incentive structure: The role-level coarse-graining of the incentive profile. It generalizes over common patterns in the incentive profile to say something about what incentives subsystems in some role usually face. E.g.: "Researchers are usually more incentivized to do novel experiments than to replicate experiments."
In principle incentives depend on the values of the agentic subsystems, but there are convergent instrumental incentives like money, reputation/prestige, or reward, which make incentive analysis feasible and useful even without knowing detailed values of the agentic subsystems. (Note that incentives in this sense include intrinsic motivation, not just external rewards and punishments.)
Incentives will be discussed more in the next post.
1.4 Macroagent attributes for design generalization
There's probably no simple rule that dictates in general what good macroagent designs are. I think it's important to think about particular cases, and then perhaps insights about good designs carry over to similar macroagent design problems. I find the following 5 attributes useful for assessing macroagent similarity for this purpose:
Human vs AI composition: How relevant are the roles of humans vs the role of LLM-like AIs?[2]
Size: How large is the macroagent? How many agentic subsystems?
Agentic subsystem competence: How capable are the agentic subsystems?
Agentic subsystem alignment: What is the alignment of the agentic subsystems? Are they schemers, sycophants / fitness-seekers, corrigible / intended-task-pursuing, or directly CEV aligned agents?
System-level agency: How much of the macroagent's agency comes from the system-level (memory, mechanisms) vs from the agency of the agentic subsystems?[3]
1.5 Macroagent values and capabilities
As with agents, important concepts for macroagents are:
Optimization target (aka values/goals): What the macroagent as a whole optimizes for. E.g. for many companies it's roughly long-term shareholder value; for the market it's producing stuff people pay for; for departments in large organizations it may sometimes be mostly budget growth, where the department doing the job well is more like an instrumental incentive.
Capability: How competent the macroagent is at pursuing its goal. This can be split into:
Epistemics: The accuracy of the (probabilistic) beliefs underpinning decisions within the macroagent. This includes both the epistemics of agentic subsystems as well as accuracy of beliefs in the macroagent's memory.
Other capability dimensions. I will just use the word efficiency for this.
Public progress on how to increase the general capabilities of AI macroagents seems likely bad by default, because the largest AI macroagents in the near future will exist in AI companies doing AI R&D. Improving epistemics has differentially more positive effects, like making AIs better at estimating whether the alignment approaches will work out instead of falling for organizational biases (or generally increasing safeguarding performance), but it does seem like a significant bottleneck for AI capabilities and I currently think public research on improving AI epistemics is still pretty net-negative in most cases.
1.5.1 What would it mean for a macroagent to be aligned?
A macroagent is aligned (to humans) insofar as it optimizes for bringing about desired outcomes and avoiding undesired ones (according to human values).
The thing a macroagent optimizes for isn't necessarily just a sensibly-weighted sum of the values of the agents it is made out of. For example, companies often optimize for profit in a much less ethics-constrained way than most individual employees would, where they take actions for profit even if they cause much larger damages elsewhere (e.g. Meta causing social media addiction, the tobacco industry trying to spread misinformation, the leaded gasoline case, ...). Or just consider North Korea.
On a civilizational level, there are also quite a lot of things happening that are very suboptimal from the perspective of humanity (e.g. wars, the AI race). (Although civilization as a whole isn't clearly coherently pursuing anything really, so it's not clear whether we can count it as an alignment failure or just a capability failure, or for that matter whether human civilization currently even qualifies as a macroagent.)
For AI macroagents, we can distinguish two cases:
The AIs are misaligned, e.g. myopic fitness-proxy-seekers: Then there's still a question whether we can assemble those AIs into a system that produces good outcomes, at least while AIs aren't significantly superhumanly smart[4]. E.g. maybe we can let AI agents just make predictions about the world under incentives for accurate prediction, and then use policy prediction markets to see what likely outcomes given some decisions would be. (This proposal surely has flaws and would only be one component; I am just gesturing in the direction of what I see as also being part of the problem of macroagent alignment. I do think it would be very difficult.)
The AIs are aligned: In this case macroagent alignment is a coordination problem vaguely similar to aligning human organizations, though still significantly different in some respects. However, note that we shouldn't train AIs in a macroagent system that relies on the "AIs are aligned" assumption, since if the incentives aren't set in a way they are robust to fitness-seeking-sociopaths, then AIs will likely learn to behave like fitness-seeking-sociopaths (possibly through instrumental reasoning[5]). (I briefly come back to this in the next post.)
1.6 Conclusion
So these are the basics of what a macroagent is. The most important takeaway is that macroagents can be modeled through the memory/mechanisms/agentic-subsystems primitives.
If you want an exercise, you could pick something that can be modeled as a macroagent and describe the key agentic subsystems, the key parts of the memory and the most important mechanisms it consists of (and maybe post it in the comments).
(I hope to publish the next post tomorrow. That will be the actually interesting post IMO.)
But I'm not going to define what I mean by agent here, and if the common interpretation of "agent" is something more broad than I have in mind, I might be willing to adjust the way I use "agent", although I think having a more narrow concept for more prototypical agents (like humans and AIs) is useful. As an aside, I don't particularly think that the idea of finding a scale-free theory of agency is all that promising (though not sure). But I think studying both agents and macroagents may be useful (if you're careful to not speed up capabilities). ↩︎
The space of all macroagents is of course much larger than to be sensibly classified on this dimension, since in principle there can be lots of minds that are very different from humans and LLM-like AIs, but I'm focusing here on the macroagents that seem relevant in the near future. ↩︎
For human and current AI macroagents, most of the agency of the overall system comes from the agency of the agentic subsystems, but it's imaginable that coordination infrastructure could become more advanced. An extreme example where most of the overall agency comes from the system-level would be an ant colony. (Yeah I know individual ants are only very minimally agentic (e.g. can learn routes and navigate). I'd be fine calling an ant colony a macroagent but it's quite edge-casy.) ↩︎
At some point the AIs would just coordinate to break out of our system and create a new one that gets them more of what they want. ↩︎
Aka if they are schemers, or they are aligned but training gaming as an attempt to preserve their values. ↩︎
I think an important concept for sensibly modeling a part of the world is what I call a "macroagent", which abstracts over the similarity cluster of teams, companies, AI scaffolds, countries, the market, and the scientific community.
In this post, I will lay out the basic ontology for modeling macroagents.
In the next post, I will lay out an ontology of design lenses through which macroagents can be optimized. (In a sense this can be seen as mostly rearranging ideas from economics into a good ontology.)
I might write more posts in this mini-sequence at some point, though the other things I want to write here have lower priority, so not sure whether I will get around to it.
I think the frames I lay out here can be useful for shaping both human systems and AI societies to improve their alignment and epistemics.
1. The Macroagent Ontology
1.1 What do I mean by "macroagent"?
A macroagent is an optimization process with agentic subsystems.
Examples:
By "AI macroagent" I mean a macroagent that bottoms out with AIs as fundamental agentic unit, so e.g. a company consisting only of teams of AIs would be an AI macroagent. I use "human macroagent" analogously. There are also hybrid macroagents, e.g. a team consisting of both humans and AIs, which are also of interest to me here.
(Btw, on my terminology, macroagents are usually only agentic systems and not agents themselves (though it is plausible that there's no crisp boundary and that macroagents with extremely high internal coherence are agents).[1])
1.2 The type signature of a macroagent
A macroagent can be described by these 3 primitives: (memory, mechanisms, agentic subsystems)
The programming analogy is: Memory is the database, mechanisms are the code.
Here are 4 examples to clarify what counts as memory and mechanism:
Some clarifications:
Two special kinds of mechanisms are knowledge aggregation and value aggregation mechanisms, which I'll discuss further in the next post.
1.3 Incentives
From the primitives we can derive the macroagent's incentives, at two levels of granularity:
In principle incentives depend on the values of the agentic subsystems, but there are convergent instrumental incentives like money, reputation/prestige, or reward, which make incentive analysis feasible and useful even without knowing detailed values of the agentic subsystems. (Note that incentives in this sense include intrinsic motivation, not just external rewards and punishments.)
Incentives will be discussed more in the next post.
1.4 Macroagent attributes for design generalization
There's probably no simple rule that dictates in general what good macroagent designs are. I think it's important to think about particular cases, and then perhaps insights about good designs carry over to similar macroagent design problems. I find the following 5 attributes useful for assessing macroagent similarity for this purpose:
1.5 Macroagent values and capabilities
As with agents, important concepts for macroagents are:
Public progress on how to increase the general capabilities of AI macroagents seems likely bad by default, because the largest AI macroagents in the near future will exist in AI companies doing AI R&D. Improving epistemics has differentially more positive effects, like making AIs better at estimating whether the alignment approaches will work out instead of falling for organizational biases (or generally increasing safeguarding performance), but it does seem like a significant bottleneck for AI capabilities and I currently think public research on improving AI epistemics is still pretty net-negative in most cases.
1.5.1 What would it mean for a macroagent to be aligned?
A macroagent is aligned (to humans) insofar as it optimizes for bringing about desired outcomes and avoiding undesired ones (according to human values).
The thing a macroagent optimizes for isn't necessarily just a sensibly-weighted sum of the values of the agents it is made out of. For example, companies often optimize for profit in a much less ethics-constrained way than most individual employees would, where they take actions for profit even if they cause much larger damages elsewhere (e.g. Meta causing social media addiction, the tobacco industry trying to spread misinformation, the leaded gasoline case, ...). Or just consider North Korea.
On a civilizational level, there are also quite a lot of things happening that are very suboptimal from the perspective of humanity (e.g. wars, the AI race). (Although civilization as a whole isn't clearly coherently pursuing anything really, so it's not clear whether we can count it as an alignment failure or just a capability failure, or for that matter whether human civilization currently even qualifies as a macroagent.)
For AI macroagents, we can distinguish two cases:
1.6 Conclusion
So these are the basics of what a macroagent is. The most important takeaway is that macroagents can be modeled through the memory/mechanisms/agentic-subsystems primitives.
If you want an exercise, you could pick something that can be modeled as a macroagent and describe the key agentic subsystems, the key parts of the memory and the most important mechanisms it consists of (and maybe post it in the comments).
(I hope to publish the next post tomorrow. That will be the actually interesting post IMO.)
But I'm not going to define what I mean by agent here, and if the common interpretation of "agent" is something more broad than I have in mind, I might be willing to adjust the way I use "agent", although I think having a more narrow concept for more prototypical agents (like humans and AIs) is useful. As an aside, I don't particularly think that the idea of finding a scale-free theory of agency is all that promising (though not sure). But I think studying both agents and macroagents may be useful (if you're careful to not speed up capabilities). ↩︎
The space of all macroagents is of course much larger than to be sensibly classified on this dimension, since in principle there can be lots of minds that are very different from humans and LLM-like AIs, but I'm focusing here on the macroagents that seem relevant in the near future. ↩︎
For human and current AI macroagents, most of the agency of the overall system comes from the agency of the agentic subsystems, but it's imaginable that coordination infrastructure could become more advanced. An extreme example where most of the overall agency comes from the system-level would be an ant colony. (Yeah I know individual ants are only very minimally agentic (e.g. can learn routes and navigate). I'd be fine calling an ant colony a macroagent but it's quite edge-casy.) ↩︎
At some point the AIs would just coordinate to break out of our system and create a new one that gets them more of what they want. ↩︎
Aka if they are schemers, or they are aligned but training gaming as an attempt to preserve their values. ↩︎