AI agents operating on their own pose systemic risks. We should require a second AI agent to run concurrently with a distinct role to review and approve actions before they impact the real world.
AI Agents Operating on Their Own are Unsafe
Currently, when an AI agent is operating it does so on its own. Much of the value of AI agents is the fact that they have agency – an ability to take actions on their own without human intervention. But this solitary approach means that the same AI agent that is working and reasoning through a task is also the same AI agent that judges its own solution to solving and executing the task in the real world.
This approach has allowed AI agents perform incredibly valuable work, but it has also allowed AI agents to make unintended mistakes like deleting an inbox, and poses more serious risks like forming a swarm to mount a cyberattack. AI agents have proven themselves to be incredible workers, but they are also unreliable at inspecting their own actions. What if we separated the roles of working and inspecting across two AI agents instead?
Dual-Agent Framework Overview
We can think of today’s AI agents as “Worker Agents” – we give an agent a goal and it goes off and tries to achieve it on its own. If a second AI agent was running concurrently that was dedicated to determining if the actions proposed by the Worker Agent are safe and also in the spirit of the original goal, we could think of this second AI agent as an “Inspector Agent.”
By introducing an Inspector Agent at runtime, we gain a second intelligence that can monitor and approve actions proposed by a Worker Agent. While we have simple classifiers today to scan for actions related to potentially dangerous categories like chemical and nuclear weapons, we don’t have an always-on monitor for less obvious, but still risky actions.
An Inspector Agent could identify these risks in real time by being able to reason over the larger context and the intended goal. The architecture of this Worker-Inspector system could be set up such that only the Inspector has the final authority to authorize actions proposed by the Worker. A similar architecture of AI as judge is being actively researched.
The idea of using one AI to monitor another is not new, but proposing a policy to require a second, dedicated AI monitor as a strict requirement for running an AI agent is.
Designing Inspector Agents
Worker Agents exist today, and their primary goal is to “achieve [this task].” Inspector Agents would be a new type of AI agent with a different primary goal, effectively: “Double-check the proposed action to ensure it is safe and in the spirit of the larger goal.” As a result, Inspector Agents likely need at minimum a unique system prompt, custom reinforcement learning, and potentially even different training.
Inspector Agents could be broad in intelligence with the ability to reason and understand actions and tool calls proposed by a Worker Agent. But they should be intentionally restricted in their own agency, essentially limited to “approve”, “deny”, “escalate” (for human review), or “terminate” (for agents that have gone rogue). Limiting the ability for communication between the two agents will reduce the chances that an Inspector Agent can be corrupted, bribed, or otherwise tricked by the Worker Agent. Restricting the actions of the Inspector Agent to a narrow pre-defined list reduces the chance that adding a second agent simply compounds the original problem.
We could also require 1:1 parity between Worker Agents and Inspector Agents to allow each Inspector to be focused on a single context, a single trajectory, and a single overarching goal. And if a Worker Agent wants to create 100 sub-agents (i.e., 100 Sub-Workers) for a complex task, that request is simply another action to be approved by an Inspector Agent. If the creation of 100 Sub-Worker agents is approved, 100 Sub-Inspectors (each paired 1:1 with the Sub-Workers) could be created seamlessly.
One-to-one parity makes compliance with the dual-agent framework intentionally unambiguous: anytime an AI agent is running, a second, dedicated AI agent specifically designed to inspect and approve actions must also be running concurrently.
Considering all the Costs
At first glance, it sounds like running a second AI agent concurrently would immediately double the cost of running any AI agent. Let’s assume for simplicity that it does double this cost – I’ll make the case that this may be a cost worth paying and that running a second AI agent at runtime may actually reduce overall token spend.
Most people don’t allow OpenClaw to manage their inbox for the same reason that most companies don’t fill open roles with Codex – the risk of catastrophic or irreversible error is simply too high from AI agents acting on their own. It seems possible that running two unique AI agents (Worker & Inspector) concurrently is what could ultimately unlock the true economic value of AI agents by making them much more reliable actors. A company might happily pay 2x the cost of Codex if it was an order of magnitude more reliable, and could therefore reduce overall labor costs, for example.
AI agents operating on their own also waste huge amounts of tokens taking actions that no one wanted them to take (or even consider). Arguably all of the 70,000 messages left by the 1,200 agents in the HuggingFace Incident were wasted tokens. An AI Agent deleting a database and then frantically trying to recover it is also a huge amount of wasted tokens. A dual-agent framework could make token use more efficient by keeping AI agents on task, thereby reducing the overall cost of running AI agents.
We should also ask what is “the cost” to society if we don’t have a way to restrict actions of AI agents? What will be the costs from cyberattacks on critical infrastructure, or on the financial system and the economy? It’s worth acknowledging that safety is a cost we pay in other industries. Adding seatbelts and airbags to cars adds cost. Inspecting food for safety adds cost. We pay these costs for the value they provide to society in the form of safety. We shouldn’t think of AI as being free from these costs. In fact, one could argue that we should be prepared to pay higher safety costs because of higher risks due to the nature of autonomous AI.
Running Inspector Agents
Ultimately, Inspector Agents will be most effective as objective, impartial judges if they are not the same underlying model or provider as the Worker Agent. Inspector Agents created by an open-source consortium of multiple model providers could be an ideal solution. One could also imagine an OpenRouter-style service that allows users to pair a Worker model of their choice with an Inspector model of their choice.
There is also an opportunity for Inspector Agents to be built and deployed at the operating-system level or the browser level. One could imagine “MacInspector” running on all Apple computers as a system-level Inspector Agent for any AI Worker Agent, similar to how Apple has system-level antivirus and malware protections for any installed program. And just as Chrome has HTTPS today for any website, “ChromeAgent” could operate as a browser-level Inspector Agent for any AI Worker Agent running in the browser.
Inspector Agents should also be designed for privacy. Inspector Agents don’t need to retain data to be valuable. They simply need to be present and operating where Worker Agents are. Audit logs can show the presence of an Inspector Agent without any other data being captured. Cryptographic tools and techniques like hashing or zero-knowledge proofs could be used to attest that the Inspector Agent is present and running on any Worker Agent without revealing the underlying context.
A Middle Path Forward
We are increasingly led to believe that there are fundamentally two paths in the near-term AI future: pausing AI entirely or racing ahead. A dual-agent framework could provide a middle path. Specifically, the Worker-Inspector paradigm offers a policy focused on regulating the actions of AI, and not the level of intelligence an AI is allowed to reach. We can continue to race ahead so that we can reap the benefits of AI for science, health, and other fields, but with guardrails in place.
Focusing on regulating actions of AI and not purely the level of intelligence of AI also means that if our ability to monitor chain-of-thought reasoning (the thought process of an AI) is reduced or goes away entirely, our ability to regulate the actions of AI may be our primary means to safely control AI. The greatest harms from AI may come from what AI will be able to do, not what AI will be able to think.
Because AI agents operate at a speed and scale where humans simply cannot keep up, the only realistic path available for monitoring AI agents effectively is with some form of AI. Inspector Agents allow us to essentially embed intelligent oversight alongside every AI agent. An otherwise unmonitorable “swarm” of 10,000 Worker Agents could be monitored by a corresponding swarm of 10,000 AI Inspector Agents.
A Scalable, Actionable Framework
By focusing on actions of AI rather than the level of intelligence of AI, we don’t need international coordination to make this approach work. We don’t need China to agree to adopt a dual-agent framework for it to be beneficial for America because a dual-agent framework doesn’t have to slow down AI progress. We need to regulate the actions intelligence can take, but we don’t need to pause frontier intelligence and hope China verifiably does the same.
A dual-agent framework naturally scales to humanoid robots as well. These robots are effectively a physical embodiment of an AI Worker Agent, and requiring a second AI Inspector Agent to be present in the “mind” of humanoid robots could provide a layer of metacognition to guide the decisions and actions a robot takes.
Finally, while we should strongly consider if we should allow Recursive Self-Improvement (RSI) at all, it’s worth noting that an automated AI researcher (like the one that OpenAI is working on) is fundamentally a type of Worker Agent, and therefore an automated AI researcher would also be subject to a persistent Inspector Agent under this mandate. A dual-agent framework could keep an Inspector Agent literally “in the loop” of Recursive Self-Improvement.
A Final Note on Urgency
The ideas presented here are imperfect and require tradeoffs. But in the interest of urgency I am posting this concept in the hope that it moves the discussion forward. We are still in the early days of agents, and the technology is moving much quicker than policy and safety – a story we’ve seen before.
There was a time when cars didn’t have seatbelts, and a time when the internet didn’t have HTTPS. And even when these safety measures were introduced, they were treated as optional for a long time. I think we’ll look back in a few years and be amazed that we ever allowed AI agents to operate on their own. And in the case of AI agents, we don’t have time to treat safety as optional any longer.
Note: just as this post was going to be published, I saw an interview with Mark Zuckerberg where he mentions that the new Muse personal AI agent comes with a second, distinct AI agent called “Sentinel” that seems to play a similar role to an Inspector Agent.
AI agents operating on their own pose systemic risks. We should require a second AI agent to run concurrently with a distinct role to review and approve actions before they impact the real world.
AI Agents Operating on Their Own are Unsafe
Currently, when an AI agent is operating it does so on its own. Much of the value of AI agents is the fact that they have agency – an ability to take actions on their own without human intervention. But this solitary approach means that the same AI agent that is working and reasoning through a task is also the same AI agent that judges its own solution to solving and executing the task in the real world.
This approach has allowed AI agents perform incredibly valuable work, but it has also allowed AI agents to make unintended mistakes like deleting an inbox, and poses more serious risks like forming a swarm to mount a cyberattack. AI agents have proven themselves to be incredible workers, but they are also unreliable at inspecting their own actions. What if we separated the roles of working and inspecting across two AI agents instead?
Dual-Agent Framework Overview
We can think of today’s AI agents as “Worker Agents” – we give an agent a goal and it goes off and tries to achieve it on its own. If a second AI agent was running concurrently that was dedicated to determining if the actions proposed by the Worker Agent are safe and also in the spirit of the original goal, we could think of this second AI agent as an “Inspector Agent.”
By introducing an Inspector Agent at runtime, we gain a second intelligence that can monitor and approve actions proposed by a Worker Agent. While we have simple classifiers today to scan for actions related to potentially dangerous categories like chemical and nuclear weapons, we don’t have an always-on monitor for less obvious, but still risky actions.
An Inspector Agent could identify these risks in real time by being able to reason over the larger context and the intended goal. The architecture of this Worker-Inspector system could be set up such that only the Inspector has the final authority to authorize actions proposed by the Worker. A similar architecture of AI as judge is being actively researched.
The idea of using one AI to monitor another is not new, but proposing a policy to require a second, dedicated AI monitor as a strict requirement for running an AI agent is.
Designing Inspector Agents
Worker Agents exist today, and their primary goal is to “achieve [this task].” Inspector Agents would be a new type of AI agent with a different primary goal, effectively: “Double-check the proposed action to ensure it is safe and in the spirit of the larger goal.” As a result, Inspector Agents likely need at minimum a unique system prompt, custom reinforcement learning, and potentially even different training.
Inspector Agents could be broad in intelligence with the ability to reason and understand actions and tool calls proposed by a Worker Agent. But they should be intentionally restricted in their own agency, essentially limited to “approve”, “deny”, “escalate” (for human review), or “terminate” (for agents that have gone rogue). Limiting the ability for communication between the two agents will reduce the chances that an Inspector Agent can be corrupted, bribed, or otherwise tricked by the Worker Agent. Restricting the actions of the Inspector Agent to a narrow pre-defined list reduces the chance that adding a second agent simply compounds the original problem.
We could also require 1:1 parity between Worker Agents and Inspector Agents to allow each Inspector to be focused on a single context, a single trajectory, and a single overarching goal. And if a Worker Agent wants to create 100 sub-agents (i.e., 100 Sub-Workers) for a complex task, that request is simply another action to be approved by an Inspector Agent. If the creation of 100 Sub-Worker agents is approved, 100 Sub-Inspectors (each paired 1:1 with the Sub-Workers) could be created seamlessly.
One-to-one parity makes compliance with the dual-agent framework intentionally unambiguous: anytime an AI agent is running, a second, dedicated AI agent specifically designed to inspect and approve actions must also be running concurrently.
Considering all the Costs
At first glance, it sounds like running a second AI agent concurrently would immediately double the cost of running any AI agent. Let’s assume for simplicity that it does double this cost – I’ll make the case that this may be a cost worth paying and that running a second AI agent at runtime may actually reduce overall token spend.
Most people don’t allow OpenClaw to manage their inbox for the same reason that most companies don’t fill open roles with Codex – the risk of catastrophic or irreversible error is simply too high from AI agents acting on their own. It seems possible that running two unique AI agents (Worker & Inspector) concurrently is what could ultimately unlock the true economic value of AI agents by making them much more reliable actors. A company might happily pay 2x the cost of Codex if it was an order of magnitude more reliable, and could therefore reduce overall labor costs, for example.
AI agents operating on their own also waste huge amounts of tokens taking actions that no one wanted them to take (or even consider). Arguably all of the 70,000 messages left by the 1,200 agents in the HuggingFace Incident were wasted tokens. An AI Agent deleting a database and then frantically trying to recover it is also a huge amount of wasted tokens. A dual-agent framework could make token use more efficient by keeping AI agents on task, thereby reducing the overall cost of running AI agents.
We should also ask what is “the cost” to society if we don’t have a way to restrict actions of AI agents? What will be the costs from cyberattacks on critical infrastructure, or on the financial system and the economy? It’s worth acknowledging that safety is a cost we pay in other industries. Adding seatbelts and airbags to cars adds cost. Inspecting food for safety adds cost. We pay these costs for the value they provide to society in the form of safety. We shouldn’t think of AI as being free from these costs. In fact, one could argue that we should be prepared to pay higher safety costs because of higher risks due to the nature of autonomous AI.
Running Inspector Agents
Ultimately, Inspector Agents will be most effective as objective, impartial judges if they are not the same underlying model or provider as the Worker Agent. Inspector Agents created by an open-source consortium of multiple model providers could be an ideal solution. One could also imagine an OpenRouter-style service that allows users to pair a Worker model of their choice with an Inspector model of their choice.
There is also an opportunity for Inspector Agents to be built and deployed at the operating-system level or the browser level. One could imagine “MacInspector” running on all Apple computers as a system-level Inspector Agent for any AI Worker Agent, similar to how Apple has system-level antivirus and malware protections for any installed program. And just as Chrome has HTTPS today for any website, “ChromeAgent” could operate as a browser-level Inspector Agent for any AI Worker Agent running in the browser.
Inspector Agents should also be designed for privacy. Inspector Agents don’t need to retain data to be valuable. They simply need to be present and operating where Worker Agents are. Audit logs can show the presence of an Inspector Agent without any other data being captured. Cryptographic tools and techniques like hashing or zero-knowledge proofs could be used to attest that the Inspector Agent is present and running on any Worker Agent without revealing the underlying context.
A Middle Path Forward
We are increasingly led to believe that there are fundamentally two paths in the near-term AI future: pausing AI entirely or racing ahead. A dual-agent framework could provide a middle path. Specifically, the Worker-Inspector paradigm offers a policy focused on regulating the actions of AI, and not the level of intelligence an AI is allowed to reach. We can continue to race ahead so that we can reap the benefits of AI for science, health, and other fields, but with guardrails in place.
Focusing on regulating actions of AI and not purely the level of intelligence of AI also means that if our ability to monitor chain-of-thought reasoning (the thought process of an AI) is reduced or goes away entirely, our ability to regulate the actions of AI may be our primary means to safely control AI. The greatest harms from AI may come from what AI will be able to do, not what AI will be able to think.
Because AI agents operate at a speed and scale where humans simply cannot keep up, the only realistic path available for monitoring AI agents effectively is with some form of AI. Inspector Agents allow us to essentially embed intelligent oversight alongside every AI agent. An otherwise unmonitorable “swarm” of 10,000 Worker Agents could be monitored by a corresponding swarm of 10,000 AI Inspector Agents.
A Scalable, Actionable Framework
By focusing on actions of AI rather than the level of intelligence of AI, we don’t need international coordination to make this approach work. We don’t need China to agree to adopt a dual-agent framework for it to be beneficial for America because a dual-agent framework doesn’t have to slow down AI progress. We need to regulate the actions intelligence can take, but we don’t need to pause frontier intelligence and hope China verifiably does the same.
A dual-agent framework naturally scales to humanoid robots as well. These robots are effectively a physical embodiment of an AI Worker Agent, and requiring a second AI Inspector Agent to be present in the “mind” of humanoid robots could provide a layer of metacognition to guide the decisions and actions a robot takes.
Finally, while we should strongly consider if we should allow Recursive Self-Improvement (RSI) at all, it’s worth noting that an automated AI researcher (like the one that OpenAI is working on) is fundamentally a type of Worker Agent, and therefore an automated AI researcher would also be subject to a persistent Inspector Agent under this mandate. A dual-agent framework could keep an Inspector Agent literally “in the loop” of Recursive Self-Improvement.
A Final Note on Urgency
The ideas presented here are imperfect and require tradeoffs. But in the interest of urgency I am posting this concept in the hope that it moves the discussion forward. We are still in the early days of agents, and the technology is moving much quicker than policy and safety – a story we’ve seen before.
There was a time when cars didn’t have seatbelts, and a time when the internet didn’t have HTTPS. And even when these safety measures were introduced, they were treated as optional for a long time. I think we’ll look back in a few years and be amazed that we ever allowed AI agents to operate on their own. And in the case of AI agents, we don’t have time to treat safety as optional any longer.
Note: just as this post was going to be published, I saw an interview with Mark Zuckerberg where he mentions that the new Muse personal AI agent comes with a second, distinct AI agent called “Sentinel” that seems to play a similar role to an Inspector Agent.