As Artificial Intelligence (AI) causes an erosion in the economic opportunities available for employment and a philosophical erosion with regard to how much we are able to contribute to those tasks efficiently, this post proposes the idea of ingraining human involvement as a design objective for future AI. The post only introduces the idea as opposed to providing a technical outline of how this can be implemented. I propose that this idea of human-computer interaction can be achieved through techniques in mechanistic interpretability and indicate that greater human agency might help in our broader goals towards building safe and aligned AI.
Context, Proposal and Mechanism:
While the idea of Artificial General Intelligence (AGI) was first spoken about by Mark Gubrud in 1997 I think many of us truly realised what that AGI might look like earlier this year with the proliferation of agents and release of models like GPT 5.6 and Fable 5. Till about last year, while AI had already started to make modular tasks like programming and parts of the research pipeline like literature review more efficient, I think the idea of creating a closed-loop AI-enterprise for the same activities, be it research or creating apps still seemed somewhat distant, but there is a growing body of evidence now that this is not going to be too distant a reality. The disproof of the Erdos unit-distance conjecture, and more recently the ten advances in mathematics and theoretical computer science are a case in point. This has naturally increased questions about whether humans are developing this technology, creating conditions for omnicide either unconditionally by relegating ourselves to the number-2 spot in the intelligence hierarchy or more intentionally by enabling bad actors (including possibly AI itself) with intelligence so powerful as to cause mass destruction. In this post, I will be looking primarily at the former case, which I consider to be a more novel problem, in that more historical precedent exists for the latter.[1]
In talking about the first case, I admit that it is more than a cause for mass extinction; a superior AI that greatly curtails the extent of human inquiry and satisfaction from human-led endeavours is more likely to cause a philosophical omnicide as opposed to an actual one, but even that has important consequences. So, the question remains: how do we benefit from the advances of AI-led innovation while ensuring that all humans will continue to have a sense of purpose? The core idea underlying the solution discussed in this post is to optimise not only for performance or safety in AI training but also for maximising active human involvement, which we propose can be done through reinforcement learning and mechanistic interpretability.
I think this method of more productive human-collaboration is something that shouldn’t be left to just the user but is something that frontier labs should actively start building at the beginning into the UI and later into the models themselves. I see mechanistic interpretability as serving a role in enabling this transformation, in the process, helping turn interpretability from a seemingly esoteric science to an application that can start having more direct relevance to users of AI bots and agents. Thus far, interpretability’s functional use has been to help in achieving alignment, which is a very important goal, but I think we should expand its role beyond the obvious use cases for AI safety to also include cases to enable more active human involvement in AI outputs.
Humans can potentially serve two roles in working with AI: one as a collaborator and the second as a verifier. In the former case, humans actively work with AI in its output, while in the latter, the human checks the AI’s output thoroughly. As far as direct human-AI collaboration goes, the greatest potential is in non-modular and ill-defined tasks, where the humans can have a difference of opinion (by the very nature of the problem) and guide the AI (potentially individualised AI bots) in accordance with their opinions on the subject. In this regard, I believe techniques from mechanistic interpretability can act as a metacognitive filter for LLMs whereby they can decide specific steps of thinking or execution processes where the marginal value of human collaboration is greatest. The type of questions that LLMs ask can be improved gradually with reinforcement learning to ensure that human collaboration in the entire process remains active.
A potential use case of this type of collaboration could be in historical writing, where an AI can summarise facts, but the human can be selective about what to include and actively evaluate the possible historical narratives that AI suggests or form a completely new narrative that she can debate the plausibilty of with AI, which can later be written up by the human. This will help in ensuring heterogeneity of beliefs as opposed to content that echoes similar opinions written in a syntactically diverse style. That said, there are certainly tasks where human involvement will not be necessary and autonomous AI agents may perform those tasks autonomously; an example of such a task may be translating a script from one programming language to another.
One method for better human-AI collaboration is actively questioning the human where there is possible subjectivity in the prompts – and I mean actual questions where human text responses are required as opposed to the questions currently posed by frontier models, which have multiple-choice options with a model recommendation which the human is likely to choose even if it’s not the best option in reality. A more complete picture of how LLMs work might help users better control the nature of interaction with AI, possibly by steering internal vectors towards goals that correlate positively to promoting human choice in AI-generated output. Another interesting interpretability question in this regard is with relevance to whether the questions that current models ask are questions out of genuine deference to human opinion or whether this is some form of sycophantic behaviour to create an illusion of agency for the human. This might help us in ensuring that collaborative models are engaging with humans out of true intent.
Additionally, to incentivise the humans to actively engage with AI, we might create a discrete report from user logs to determine the share of human contribution and AI contribution and the nature of that interaction. Currently, console logs provide a very imprecise and cumbersome method to understand the nature of human interactions with AI; a more clearly defined metric might help another human better judge the nature of collaboration, which can serve as an important method, especially in education. This metric might measure factors like novel ideas proposed by humans or the number of times the human contradicts the advice of the bot or actively debates with it.
Looking at the second possible role of a human as the verifier, a similar sort of metric could be used to determine whether a human fully understands what the AI is doing to such an extent that they can take responsibility for its decision. This might prove particularly useful in fields like medicine and law, where there are fundamental questions of a person taking responsibility for the model’s decision, which I don’t think can be taken till the human completely understands the output of the model.
Additionally, it might be interesting to see whether openly reporting interpretability metrics like a model entropy as a proxy for the probability of hallucination or a sycophancy tracker in the UI while working with the AI bot might help humans better vet the output of the model. The idea is that if humans are able to openly see the way in which the model is formulating its responses, it’s likely that the human in the loop will be able to better understand the way in which the chatbot comes to its answers and possibly correct or at the very least notice fallacies in the reasoning process. The idea here is to understand how the model works beyond simple chains-of-thought, which themselves are generative and therefore risk the chance of not being faithful to the model’s thinking process.
I believe that future research in AI safety, at least in the short term, should help achieve these objectives as a way of softening the economic impact of AI and philosophical implications of the technology on the nature of human ambition beyond the goal of AI aligned with human values. A valid question on this proposed system: why bother with it? When AI can potentially discover scientific ideas on its own and build apps, why should the human be forced to come in the loop and spend their cognitive capital?
I think there are two primary arguments for this: firstly, I think there are genuine areas of human-AI collaboration where the task being performed is not modular or very objective. Wherever there are questions about creating a framework anew, there are decisions to be made on what to optimise for which in order to retain human agency, which should be decisions that are thought of by the human in the loop. So, the idea is not necessarily that the human is doing a particular activity better than the AI; it’s just that the human has beliefs and thought processes that are different from that of the AI, and therefore that’s sufficient reason for the human to be actively involved in the process.
Beyond this reason, I think humans should partake in this loop as their moral duty to prevent civilisational collapse, and it’s something that frontier labs and governments can help incentivise for more and more people. In the short term, there are also some efficiency and output gains from the proposed nature of interaction, as there are still limits on model context and the ways in which agent interface with other applications through Model Context Protocol (MCP) can be inefficient.[2]
The idea of optimising RL and interpretability for human-AI collaboration is not something that can readily be applied. There are a large number of open research questions, especially on the interpretability side of things which need addressing. There are questions about how reliable entropy is as a proxy for model hallucination. How do we optimally identify circuits within the transformers that we can steer to ensure our objectives are met? And how exactly can we develop a deterministic method to judge human-AI collaboration? There may also be concerns that the Human-AI collaboration metric may fall prey to Goodhart's law and be gamed, an important consideration while determining the exact mechanism of the metric. But I believe these are far more tractable problems than the larger issue of how we will manage a future population of 10 billion people who are being displaced from their jobs and feel threatened.[3]
In conclusion, to navigate the AI revolution, there is a need to work proactively in ensuring that humans continue to preserve their role in the decision-making process. Achieving this goal requires going beyond ethical or regulatory guidelines; it should incorporate mechanistic interpretability and collective feedback loops to ensure human choice remains all-important in the face of advanced autonomous digital agents.
Like AI-enabled warfare, biological, chemical and nuclear weapons all had the potential to cause omnicide, but they haven’t due to formal human collaboration banning them or due to constructs like Mutually Assured Destruction (MAD.) Even in the case of corporate consolidation of important resources, there is precedent in anti-trust law and state-led dissolution of corporate giants.
These are only arguments that can be used in the short term, as model contexts are likely to get larger in the future and MCP is also likely to be integrated more seamlessly, thus making the model as fast as a human while working on another application.
Humans feeling threatened can erode institutions and lead to mass hysteria. Witch Trials in Europe and America are an example of the scale of destruction that civilisational threats can cause.
Summary:
As Artificial Intelligence (AI) causes an erosion in the economic opportunities available for employment and a philosophical erosion with regard to how much we are able to contribute to those tasks efficiently, this post proposes the idea of ingraining human involvement as a design objective for future AI. The post only introduces the idea as opposed to providing a technical outline of how this can be implemented. I propose that this idea of human-computer interaction can be achieved through techniques in mechanistic interpretability and indicate that greater human agency might help in our broader goals towards building safe and aligned AI.
Context, Proposal and Mechanism:
While the idea of Artificial General Intelligence (AGI) was first spoken about by Mark Gubrud in 1997 I think many of us truly realised what that AGI might look like earlier this year with the proliferation of agents and release of models like GPT 5.6 and Fable 5. Till about last year, while AI had already started to make modular tasks like programming and parts of the research pipeline like literature review more efficient, I think the idea of creating a closed-loop AI-enterprise for the same activities, be it research or creating apps still seemed somewhat distant, but there is a growing body of evidence now that this is not going to be too distant a reality. The disproof of the Erdos unit-distance conjecture, and more recently the ten advances in mathematics and theoretical computer science are a case in point. This has naturally increased questions about whether humans are developing this technology, creating conditions for omnicide either unconditionally by relegating ourselves to the number-2 spot in the intelligence hierarchy or more intentionally by enabling bad actors (including possibly AI itself) with intelligence so powerful as to cause mass destruction. In this post, I will be looking primarily at the former case, which I consider to be a more novel problem, in that more historical precedent exists for the latter.[1]
In talking about the first case, I admit that it is more than a cause for mass extinction; a superior AI that greatly curtails the extent of human inquiry and satisfaction from human-led endeavours is more likely to cause a philosophical omnicide as opposed to an actual one, but even that has important consequences. So, the question remains: how do we benefit from the advances of AI-led innovation while ensuring that all humans will continue to have a sense of purpose? The core idea underlying the solution discussed in this post is to optimise not only for performance or safety in AI training but also for maximising active human involvement, which we propose can be done through reinforcement learning and mechanistic interpretability.
I think this method of more productive human-collaboration is something that shouldn’t be left to just the user but is something that frontier labs should actively start building at the beginning into the UI and later into the models themselves. I see mechanistic interpretability as serving a role in enabling this transformation, in the process, helping turn interpretability from a seemingly esoteric science to an application that can start having more direct relevance to users of AI bots and agents. Thus far, interpretability’s functional use has been to help in achieving alignment, which is a very important goal, but I think we should expand its role beyond the obvious use cases for AI safety to also include cases to enable more active human involvement in AI outputs.
Humans can potentially serve two roles in working with AI: one as a collaborator and the second as a verifier. In the former case, humans actively work with AI in its output, while in the latter, the human checks the AI’s output thoroughly. As far as direct human-AI collaboration goes, the greatest potential is in non-modular and ill-defined tasks, where the humans can have a difference of opinion (by the very nature of the problem) and guide the AI (potentially individualised AI bots) in accordance with their opinions on the subject. In this regard, I believe techniques from mechanistic interpretability can act as a metacognitive filter for LLMs whereby they can decide specific steps of thinking or execution processes where the marginal value of human collaboration is greatest. The type of questions that LLMs ask can be improved gradually with reinforcement learning to ensure that human collaboration in the entire process remains active.
A potential use case of this type of collaboration could be in historical writing, where an AI can summarise facts, but the human can be selective about what to include and actively evaluate the possible historical narratives that AI suggests or form a completely new narrative that she can debate the plausibilty of with AI, which can later be written up by the human. This will help in ensuring heterogeneity of beliefs as opposed to content that echoes similar opinions written in a syntactically diverse style. That said, there are certainly tasks where human involvement will not be necessary and autonomous AI agents may perform those tasks autonomously; an example of such a task may be translating a script from one programming language to another.
One method for better human-AI collaboration is actively questioning the human where there is possible subjectivity in the prompts – and I mean actual questions where human text responses are required as opposed to the questions currently posed by frontier models, which have multiple-choice options with a model recommendation which the human is likely to choose even if it’s not the best option in reality. A more complete picture of how LLMs work might help users better control the nature of interaction with AI, possibly by steering internal vectors towards goals that correlate positively to promoting human choice in AI-generated output. Another interesting interpretability question in this regard is with relevance to whether the questions that current models ask are questions out of genuine deference to human opinion or whether this is some form of sycophantic behaviour to create an illusion of agency for the human. This might help us in ensuring that collaborative models are engaging with humans out of true intent.
Additionally, to incentivise the humans to actively engage with AI, we might create a discrete report from user logs to determine the share of human contribution and AI contribution and the nature of that interaction. Currently, console logs provide a very imprecise and cumbersome method to understand the nature of human interactions with AI; a more clearly defined metric might help another human better judge the nature of collaboration, which can serve as an important method, especially in education. This metric might measure factors like novel ideas proposed by humans or the number of times the human contradicts the advice of the bot or actively debates with it.
Looking at the second possible role of a human as the verifier, a similar sort of metric could be used to determine whether a human fully understands what the AI is doing to such an extent that they can take responsibility for its decision. This might prove particularly useful in fields like medicine and law, where there are fundamental questions of a person taking responsibility for the model’s decision, which I don’t think can be taken till the human completely understands the output of the model.
Additionally, it might be interesting to see whether openly reporting interpretability metrics like a model entropy as a proxy for the probability of hallucination or a sycophancy tracker in the UI while working with the AI bot might help humans better vet the output of the model. The idea is that if humans are able to openly see the way in which the model is formulating its responses, it’s likely that the human in the loop will be able to better understand the way in which the chatbot comes to its answers and possibly correct or at the very least notice fallacies in the reasoning process. The idea here is to understand how the model works beyond simple chains-of-thought, which themselves are generative and therefore risk the chance of not being faithful to the model’s thinking process.
I believe that future research in AI safety, at least in the short term, should help achieve these objectives as a way of softening the economic impact of AI and philosophical implications of the technology on the nature of human ambition beyond the goal of AI aligned with human values. A valid question on this proposed system: why bother with it? When AI can potentially discover scientific ideas on its own and build apps, why should the human be forced to come in the loop and spend their cognitive capital?
I think there are two primary arguments for this: firstly, I think there are genuine areas of human-AI collaboration where the task being performed is not modular or very objective. Wherever there are questions about creating a framework anew, there are decisions to be made on what to optimise for which in order to retain human agency, which should be decisions that are thought of by the human in the loop. So, the idea is not necessarily that the human is doing a particular activity better than the AI; it’s just that the human has beliefs and thought processes that are different from that of the AI, and therefore that’s sufficient reason for the human to be actively involved in the process.
Beyond this reason, I think humans should partake in this loop as their moral duty to prevent civilisational collapse, and it’s something that frontier labs and governments can help incentivise for more and more people. In the short term, there are also some efficiency and output gains from the proposed nature of interaction, as there are still limits on model context and the ways in which agent interface with other applications through Model Context Protocol (MCP) can be inefficient.[2]
The idea of optimising RL and interpretability for human-AI collaboration is not something that can readily be applied. There are a large number of open research questions, especially on the interpretability side of things which need addressing. There are questions about how reliable entropy is as a proxy for model hallucination. How do we optimally identify circuits within the transformers that we can steer to ensure our objectives are met? And how exactly can we develop a deterministic method to judge human-AI collaboration? There may also be concerns that the Human-AI collaboration metric may fall prey to Goodhart's law and be gamed, an important consideration while determining the exact mechanism of the metric. But I believe these are far more tractable problems than the larger issue of how we will manage a future population of 10 billion people who are being displaced from their jobs and feel threatened.[3]
In conclusion, to navigate the AI revolution, there is a need to work proactively in ensuring that humans continue to preserve their role in the decision-making process. Achieving this goal requires going beyond ethical or regulatory guidelines; it should incorporate mechanistic interpretability and collective feedback loops to ensure human choice remains all-important in the face of advanced autonomous digital agents.
Like AI-enabled warfare, biological, chemical and nuclear weapons all had the potential to cause omnicide, but they haven’t due to formal human collaboration banning them or due to constructs like Mutually Assured Destruction (MAD.) Even in the case of corporate consolidation of important resources, there is precedent in anti-trust law and state-led dissolution of corporate giants.
These are only arguments that can be used in the short term, as model contexts are likely to get larger in the future and MCP is also likely to be integrated more seamlessly, thus making the model as fast as a human while working on another application.
Humans feeling threatened can erode institutions and lead to mass hysteria. Witch Trials in Europe and America are an example of the scale of destruction that civilisational threats can cause.