This post was originally written for my Substack. Please feel free to subscribe if this kind of content interests you!
Frontier AI companies have long had the stated goal of automating their own research. The default plan has been to build models that are even better than their most skilled researchers at advancing the frontier, and let these models build their own successors; a feedback loop that will ultimately result in superintelligence that dwarfs the collective intelligence of humanity. This process is known as recursive self-improvement (RSI).
The problem with this plan is that it might run entirely out of human control, making it more likely that the resulting superintelligence will kill everyone. So as AI models are increasingly accelerating their own development, AI companies are starting to think that this RSI thing may not be such a good idea after all. Anthropic CEO Dario Amodei recently wrote on his blog that RSI “must be pursued very carefully, if at all”. OpenAI’s official blog states: “Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely”. Similar calls are coming from elsewhere. Former DeepMind researcher Alex Turner wrote in The Guardian that companies must be prevented from “allowing AI to self-improve into an uncontrollable level of intelligence”. And New York Times columnist Ezra Klein dedicated an entire op-ed to calling for a ban on RSI:
Human beings need to control the frontier, and controlling the frontier means stopping the labs from doing something they are on the cusp of doing: recursive self-improvement, the process by which A.I.s begin autonomously building and improving new generations of more powerful A.I.s at ever more rapid speeds.
Sounds simple enough. We probably should just not let AI models build successive AI models in ways that human researchers can’t oversee, because that sounds very dicey. But there’s another wrinkle, which is that no one can agree on what RSI actually is. Sam Altman made this point in an interview with Fortune: “I think it’s very hard to say what a ban on RSI means”. Maybe he is just participating in the classic AGI Lab CEO practice of injecting undue slipperiness into the definition of various terms with the aim of confusing everyone, but he does have a point. There are many different definitions of RSI, which I’ll cover below. Even if we could agree on a definition, it is unclear to what extent each of them is already occurring, or whether we’ll know when it is. So “banning RSI” will require picking a somewhat arbitrary (and hopefully very conservative) threshold, and sticking with it.
What even is RSI?
Here is every definition of RSI that I could come up with, listed from most to least expansive:
#1 Any time a technology is used to help improve itself, this is RSI. This is the broadest possible way to define RSI, and I think a pretty useless one in our present situation. Under this definition, some version of RSI has been happening throughout the history of technological progress; a blacksmith can use his hammer, tongs, and anvil to help craft better hammers, tongs, and anvils. A more useful one would take into account the ways in which RSI in the AI context specifically will be qualitatively different; it will be faster, less easy to oversee, and less controllable.
#2 If humans are using AI to help with their work in literally any way at all, this is RSI: This version of RSI is already well underway. For example, Anthropic reports that Claude currently “collaborates” in over 90% of work at the company and “leads” 26%. This is also the fairly conservative threshold at which Ezra Klein proposes we institute a ban: “A couple of years ago, none of these labs had turned substantial coding over to the A.I.s. It was human beings typing code at human speeds with our clumsy human fingers. Now, most of the code is written by AI. So as a first step, we could just go back to where none of the code is written by AI.”
#3 If AI can replace a human researcher of a certain level, this is RSI: We can also think of RSI as a thing that happens once AI can replace a certain level researcher at an AI company. This need not result in a closed loop where the next generation of model is built without any human input; OpenAI claims to have internally developed an AI research “intern” that can “carry out well-defined research tasks under human direction”. They have the grander goal of developing a fully automated AI researcher by March 2028. This researcher, if successfully developed, would presumably cross the much harder threshold of choosing which research questions to pursue and judging success (the ever-elusive “research taste”). As of summer 2026, frontier models are apparently still quite bad at this.
But now we’re stuck at another definitional crossroads: is it RSI if models are replacing junior researchers but still need the direction of senior ones? Or, put another way, if AI models are very good engineers (they can complete well-scoped tasks) but not very good researchers (they don’t themselves define the tasks)? The intuitive answer to this question is “no” if we’re concerned about runaway capability advancements that humans cannot supervise; even having an extremely large, extremely fast population of automated junior AI researchers is a bit like staffing a restaurant with 10x more very competent cooks, but still only having one head chef capable of writing the menu. In this scenario, we wouldn’t expect our super-chefs to start developing novel recipe ideas themselves, even if they’re helping cook and serve the food a whole lot faster.
So, we could decide that automating lower-level research doesn’t count as RSI so long as humans remain in charge of determining research directions. The problem with this is that maintaining high-level human oversight doesn’t necessarily eliminate danger. For one, automating large amounts of research, even if this is junior or “grunt” work, still increases the speed of progress – and speed is at least somewhat upstream of risk. In our restaurant analogy, we could imagine that the head chef starts to get much faster at improving their menus, because they have a larger workforce that can more quickly iterate on experimental dishes. For two, junior automated researchers could be misaligned, and increasing the volume produced by these researchers makes it more likely that incidences of misaligned behaviour will slip through the cracks. It likely wouldn’t be possible for humans to review all the work completed by a large population of AIs, and automated monitoring techniques are still imperfect.
This idea of replacement features in several frontier lab frameworks. Here’s the minimum threshold at which researcher automation would trigger some kind of safeguard at each company (these safeguards range from halting development to increased security standards):
It’s notable that the bar varies a lot here, and in some cases is really very high (at Anthropic the bar appears to literally be “AI can automate the research arm of our whole entire company”). It has also been subject to some historical goalpost-shifting. A previous version of Anthropic’s Responsible Scaling Policy triggered heightened security standards for models with “the ability to fully automate the work of an entry-level, remote-only Researcher at Anthropic”. This threshold was retired in Version 3 of Anthropic’s RSP. It’s also notable that its system card for Claude Opus 4.6, published in February 2026, concluded that it was difficult to comprehensively rule out the model surpassing it. Anthropic tentatively concluded that it hadn’t, but this determination was largely based on a survey of 16 employees, who all subjectively assessed that Opus 4.6 would not be able to replace an entry-level researcher.
So using researcher automation as a standard for RSI appears to have some problems. It’s hard to know what degree of replacement – and at what level of research – would present risk. We also don’t seem to have great ways of assessing when a given model could substitute for a human researcher, given that we’re still relying on methods like asking said researchers whether it seems like the model could do their job. Replacement thresholds for triggering safeguards at various companies are, predictably, all over the place.
#4 When automated research is speeding up capability progress by a certain amount, this is RSI. Rather than defining RSI by the processes that produce it (such as a particular technique, or automating a certain level of role), we could define it as a specific outcome. One such outcome could be the speedup in capabilities progress resulting from automation. This is a common feature of AI lab safety frameworks. For example, OpenAI’s Preparedness Framework separates its “leading indicator” for AI self-improvement that would pose critical risks (a somewhat vaguely-defined “superhuman research scientist agent”) from its “lagging indicator” (a generational model improvement that happens 5x as quickly as equivalent progress in 2024).
There are some bright red safety flags here. It probably isn’t clear that you’ve hit your “leading indicator” of a very competent AI researcher until you can observe your “lagging indicator” of extremely fast capabilities progress. OpenAI says they will halt development if and when this threshold is met – but at what point is this supposed to occur? If you’ve already observed months of rapid progress, is it already too late? That the observable phenomenon that defines critical risk “lags” behind the source of that risk is kind of the entire problem.
#5 When there’s no human in the loop, this is RSI. Maybe autonomous AI R&D shouldn’t really count as RSI unless it’s happening without any human input. This is probably the most commonly invoked definition in calls to “ban” or “not to pursue” RSI. For example, Yoshua Bengio calls for prohibition of “recursive self-improvement where AI could create progressively advanced models without human input”. Anthropic defines RSI as “an AI system capable of fully autonomously designing and developing its own successor”, a phenomenon that is “not inevitable” and should only be pursued if it can be done safely.
I agree that this is a terrifying prospect, but it’s worth dwelling on quite how high this bar is. It hinges on no human involvement at all. The obvious issue with operationalising some kind of ban around this standard is that it would be quite easy for humans to be peripherally involved in a process of automated AI R&D that they nonetheless have no true control over. An AI company could install some humans in the loop that play a rubber-stamping function in order to stay on the right side of the law – but in a world where smarter-than-human AIs are producing reams of research 24 hours a day, will these humans even understand what they are rubber-stamping? There’s some evidence of a tragic paradox here: that “the more you try to involve humans in reviewing AI decisions, the less meaningful that review becomes”. A 2024 Harvard Business School study of 228 human evaluators screening recommendations for early-stage innovation found that evaluators were more likely to accept recommendations made by LLMs, and even more likely to accept decisions that the LLMs provided a narrative explanation for. But when explanations were provided, decision quality decreased (as determined by a panel of experts). So humans have a tendency to defer to AIs – and to defer more often when explainability is increased, which in turn makes their decisions worse. Imagine this phenomenon playing out in a scenario where AIs are conducting increasingly inscrutable research into their own development, perhaps while deceiving their overseers. I don’t feel great about that.
My guess is that human researchers will be kept around for as long as they can contribute anything at all – it’s fairly cheap for AI companies to pay out some researcher salaries (which are admittedly high by most people’s standards, but low compared to overall compute expenditure) for political, regulatory, or signalling reasons. But this won’t necessarily be evidence that the oversight they are supplying is meaningful, or that RSI isn’t well underway.
#6 When acceleration becomes self-sustaining, this is RSI. Another bar for defining RSI is that it must accelerate in a self-sustaining fashion; each capability gain produces more than one subsequent gain. Under this definition, even automating all AI research would not necessarily result in RSI, because progress could still be bottlenecked by access to compute or fundamental difficulties in finding new ideas. For something to truly count as RSI, you would need progress to keep accelerating even when factors like compute and human labour stay stagnant. One paper uses a concept that is similar to the reproduction rate in epidemiology – in the same way that an outbreak can only keep growing if R >1 (ie when each infected person goes to infect more than one other person), explosive growth in AI capabilities can only occur if its recursive reproduction number, written as ℛ_AI, > 1. Otherwise, progress will eventually slow down.
This is in some ways a higher bar than full research automation, but in others a lower one; it doesn’t actually require every human researcher to be automated away. Because it requires that inputs like human labour and compute are held constant, it would still be possible for an AI lab to achieve extremely accelerated progress without meeting it by building more chips and/ or hiring more people. It is also probably even less useful as an operationalisable definition than a particular level of researcher automation or an entirely closed feedback loop, because, in practice, it is likely only confirmable in retrospect. Researchers remain agnostic about whether or under what conditions we’d see self-sustaining progress, and that it may be indistinguishable at first from progress that is very fast but not self-sustaining. So although this definition gets at the thing people are usually most scared of when they discuss RSI (an uncontrollable, recursive loop that results in an “intelligence explosion”), it does not seem like the kind of thing it would be possible to “ban”. Self-sustaining progress is a phenomenon, not a set of actions. To use a crude analogy, if you wanted to ban the development and release of an engineered virus, you wouldn’t make a certain level of transmissibility the threshold for that ban – especially if you had no way of knowing how contagious it would be ahead of time. Otherwise you’re waiting for people to release viruses into the population and waiting to see if they become pandemics.
A handy chart of RSI definitions (disclosure: text is Claude-generated)
So what does this mean for a ban?
Having done this deep-dive on various definitions of RSI, I feel pretty confused about the idea of “banning” RSI. Depending on how you define it, RSI can either be a set of activities (automating certain roles, using particular techniques, etc) or an outcome (a certain level of speedup or the emergence of a self-sustaining progress loop). Obviously, you can only proactively ban the former. But given that we still have uncertainty about precisely what inputs will lead to the outcome we’re afraid of, it seems easy to target the wrong thing.
But this doesn’t mean there’s nothing to be done. Our aim is to prevent a dynamic where AI is building successive generations of itself in ways that humans cannot reliably supervise. So we should severely limit the amount of authority that we cede to AIs! Any one specific intervention might end up being insufficient, so we should maximise – or at least diversify – the number of preventative measures we take. It seems like there are plenty of such interventions. At the extreme end, just don’t let frontier researchers AI-generate code (per Ezra’s suggestion). This feels like an extremely costly intervention and would have the effect of dramatically slowing progress, but everyone dying is also extremely costly. Or one lesser intervention: only use models that have been available publicly for some set period of time for internal capabilities research. For example, the AI Futures Project recommends a minimum of nine months. If your models haven’t been rigorously tested enough for public release, why should you trust them with the much higher-stakes task of accelerating your own research?
You could also cap the number of hours that agents can work autonomously per day. Cap the amount of compute that is allocated to autonomous AI R&D projects. Benchmark the AI R&D capabilities of your models frequently and publicly. If they’re getting scarily good, don’t use them internally. Design an effective embedded auditing system such that all of the above can be verified. There are probably many more things you could do here that would take many more blog posts to fully operationalise. At the very least, setting off a recursive feedback cycle should emphatically not be the goal of any frontier AI company. This should be the thing they are, above all, trying to avoid. Don’t let Claude n build Claude n+1 while you go home and knit sweaters. That’s a completely horrible idea. Treat your internal models like a cohort of untrustworthy, cheating, lying interns. Don’t give them access to too many resources. Don’t turn your back on them. Assume they are trying to screw you over.
That RSI is hard to define or predict is an argument for a very, very conservative line. It’s possible that Ezra Klein’s “no AI-generated code” proposal is more of a rhetorical salvo than a serious suggestion, but, if so, I think his point is compelling. If you’re so scared of RSI, you should be backing away from it like an angry bear in the woods. In a rebuttal to Ezra’s piece, economist Tyler Cowen asks whether “humans [are] even capable of writing the next steps of code that are required for further progress”. But given our present face-to-face-with-a-bear-in-the-woods predicament, it’s far from clear that we want to make further progress anyway. Arguably, a host of preventative measures to try and stay on the right side of the RSI-line are second best to bringing this whole thing to a grinding halt.
This post was originally written for my Substack. Please feel free to subscribe if this kind of content interests you!
Frontier AI companies have long had the stated goal of automating their own research. The default plan has been to build models that are even better than their most skilled researchers at advancing the frontier, and let these models build their own successors; a feedback loop that will ultimately result in superintelligence that dwarfs the collective intelligence of humanity. This process is known as recursive self-improvement (RSI).
@jayelmnop on Twitter (May 2025)
The problem with this plan is that it might run entirely out of human control, making it more likely that the resulting superintelligence will kill everyone. So as AI models are increasingly accelerating their own development, AI companies are starting to think that this RSI thing may not be such a good idea after all. Anthropic CEO Dario Amodei recently wrote on his blog that RSI “must be pursued very carefully, if at all”. OpenAI’s official blog states: “Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely”. Similar calls are coming from elsewhere. Former DeepMind researcher Alex Turner wrote in The Guardian that companies must be prevented from “allowing AI to self-improve into an uncontrollable level of intelligence”. And New York Times columnist Ezra Klein dedicated an entire op-ed to calling for a ban on RSI:
Human beings need to control the frontier, and controlling the frontier means stopping the labs from doing something they are on the cusp of doing: recursive self-improvement, the process by which A.I.s begin autonomously building and improving new generations of more powerful A.I.s at ever more rapid speeds.
Sounds simple enough. We probably should just not let AI models build successive AI models in ways that human researchers can’t oversee, because that sounds very dicey. But there’s another wrinkle, which is that no one can agree on what RSI actually is. Sam Altman made this point in an interview with Fortune: “I think it’s very hard to say what a ban on RSI means”. Maybe he is just participating in the classic AGI Lab CEO practice of injecting undue slipperiness into the definition of various terms with the aim of confusing everyone, but he does have a point. There are many different definitions of RSI, which I’ll cover below. Even if we could agree on a definition, it is unclear to what extent each of them is already occurring, or whether we’ll know when it is. So “banning RSI” will require picking a somewhat arbitrary (and hopefully very conservative) threshold, and sticking with it.
What even is RSI?
Here is every definition of RSI that I could come up with, listed from most to least expansive:
#1 Any time a technology is used to help improve itself, this is RSI. This is the broadest possible way to define RSI, and I think a pretty useless one in our present situation. Under this definition, some version of RSI has been happening throughout the history of technological progress; a blacksmith can use his hammer, tongs, and anvil to help craft better hammers, tongs, and anvils. A more useful one would take into account the ways in which RSI in the AI context specifically will be qualitatively different; it will be faster, less easy to oversee, and less controllable.
#2 If humans are using AI to help with their work in literally any way at all, this is RSI: This version of RSI is already well underway. For example, Anthropic reports that Claude currently “collaborates” in over 90% of work at the company and “leads” 26%. This is also the fairly conservative threshold at which Ezra Klein proposes we institute a ban: “A couple of years ago, none of these labs had turned substantial coding over to the A.I.s. It was human beings typing code at human speeds with our clumsy human fingers. Now, most of the code is written by AI. So as a first step, we could just go back to where none of the code is written by AI.”
#3 If AI can replace a human researcher of a certain level, this is RSI: We can also think of RSI as a thing that happens once AI can replace a certain level researcher at an AI company. This need not result in a closed loop where the next generation of model is built without any human input; OpenAI claims to have internally developed an AI research “intern” that can “carry out well-defined research tasks under human direction”. They have the grander goal of developing a fully automated AI researcher by March 2028. This researcher, if successfully developed, would presumably cross the much harder threshold of choosing which research questions to pursue and judging success (the ever-elusive “research taste”). As of summer 2026, frontier models are apparently still quite bad at this.
But now we’re stuck at another definitional crossroads: is it RSI if models are replacing junior researchers but still need the direction of senior ones? Or, put another way, if AI models are very good engineers (they can complete well-scoped tasks) but not very good researchers (they don’t themselves define the tasks)? The intuitive answer to this question is “no” if we’re concerned about runaway capability advancements that humans cannot supervise; even having an extremely large, extremely fast population of automated junior AI researchers is a bit like staffing a restaurant with 10x more very competent cooks, but still only having one head chef capable of writing the menu. In this scenario, we wouldn’t expect our super-chefs to start developing novel recipe ideas themselves, even if they’re helping cook and serve the food a whole lot faster.
So, we could decide that automating lower-level research doesn’t count as RSI so long as humans remain in charge of determining research directions. The problem with this is that maintaining high-level human oversight doesn’t necessarily eliminate danger. For one, automating large amounts of research, even if this is junior or “grunt” work, still increases the speed of progress – and speed is at least somewhat upstream of risk. In our restaurant analogy, we could imagine that the head chef starts to get much faster at improving their menus, because they have a larger workforce that can more quickly iterate on experimental dishes. For two, junior automated researchers could be misaligned, and increasing the volume produced by these researchers makes it more likely that incidences of misaligned behaviour will slip through the cracks. It likely wouldn’t be possible for humans to review all the work completed by a large population of AIs, and automated monitoring techniques are still imperfect.
This idea of replacement features in several frontier lab frameworks. Here’s the minimum threshold at which researcher automation would trigger some kind of safeguard at each company (these safeguards range from halting development to increased security standards):
It’s notable that the bar varies a lot here, and in some cases is really very high (at Anthropic the bar appears to literally be “AI can automate the research arm of our whole entire company”). It has also been subject to some historical goalpost-shifting. A previous version of Anthropic’s Responsible Scaling Policy triggered heightened security standards for models with “the ability to fully automate the work of an entry-level, remote-only Researcher at Anthropic”. This threshold was retired in Version 3 of Anthropic’s RSP. It’s also notable that its system card for Claude Opus 4.6, published in February 2026, concluded that it was difficult to comprehensively rule out the model surpassing it. Anthropic tentatively concluded that it hadn’t, but this determination was largely based on a survey of 16 employees, who all subjectively assessed that Opus 4.6 would not be able to replace an entry-level researcher.
So using researcher automation as a standard for RSI appears to have some problems. It’s hard to know what degree of replacement – and at what level of research – would present risk. We also don’t seem to have great ways of assessing when a given model could substitute for a human researcher, given that we’re still relying on methods like asking said researchers whether it seems like the model could do their job. Replacement thresholds for triggering safeguards at various companies are, predictably, all over the place.
#4 When automated research is speeding up capability progress by a certain amount, this is RSI. Rather than defining RSI by the processes that produce it (such as a particular technique, or automating a certain level of role), we could define it as a specific outcome. One such outcome could be the speedup in capabilities progress resulting from automation. This is a common feature of AI lab safety frameworks. For example, OpenAI’s Preparedness Framework separates its “leading indicator” for AI self-improvement that would pose critical risks (a somewhat vaguely-defined “superhuman research scientist agent”) from its “lagging indicator” (a generational model improvement that happens 5x as quickly as equivalent progress in 2024).
There are some bright red safety flags here. It probably isn’t clear that you’ve hit your “leading indicator” of a very competent AI researcher until you can observe your “lagging indicator” of extremely fast capabilities progress. OpenAI says they will halt development if and when this threshold is met – but at what point is this supposed to occur? If you’ve already observed months of rapid progress, is it already too late? That the observable phenomenon that defines critical risk “lags” behind the source of that risk is kind of the entire problem.
#5 When there’s no human in the loop, this is RSI. Maybe autonomous AI R&D shouldn’t really count as RSI unless it’s happening without any human input. This is probably the most commonly invoked definition in calls to “ban” or “not to pursue” RSI. For example, Yoshua Bengio calls for prohibition of “recursive self-improvement where AI could create progressively advanced models without human input”. Anthropic defines RSI as “an AI system capable of fully autonomously designing and developing its own successor”, a phenomenon that is “not inevitable” and should only be pursued if it can be done safely.
I agree that this is a terrifying prospect, but it’s worth dwelling on quite how high this bar is. It hinges on no human involvement at all. The obvious issue with operationalising some kind of ban around this standard is that it would be quite easy for humans to be peripherally involved in a process of automated AI R&D that they nonetheless have no true control over. An AI company could install some humans in the loop that play a rubber-stamping function in order to stay on the right side of the law – but in a world where smarter-than-human AIs are producing reams of research 24 hours a day, will these humans even understand what they are rubber-stamping? There’s some evidence of a tragic paradox here: that “the more you try to involve humans in reviewing AI decisions, the less meaningful that review becomes”. A 2024 Harvard Business School study of 228 human evaluators screening recommendations for early-stage innovation found that evaluators were more likely to accept recommendations made by LLMs, and even more likely to accept decisions that the LLMs provided a narrative explanation for. But when explanations were provided, decision quality decreased (as determined by a panel of experts). So humans have a tendency to defer to AIs – and to defer more often when explainability is increased, which in turn makes their decisions worse. Imagine this phenomenon playing out in a scenario where AIs are conducting increasingly inscrutable research into their own development, perhaps while deceiving their overseers. I don’t feel great about that.
My guess is that human researchers will be kept around for as long as they can contribute anything at all – it’s fairly cheap for AI companies to pay out some researcher salaries (which are admittedly high by most people’s standards, but low compared to overall compute expenditure) for political, regulatory, or signalling reasons. But this won’t necessarily be evidence that the oversight they are supplying is meaningful, or that RSI isn’t well underway.
#6 When acceleration becomes self-sustaining, this is RSI. Another bar for defining RSI is that it must accelerate in a self-sustaining fashion; each capability gain produces more than one subsequent gain. Under this definition, even automating all AI research would not necessarily result in RSI, because progress could still be bottlenecked by access to compute or fundamental difficulties in finding new ideas. For something to truly count as RSI, you would need progress to keep accelerating even when factors like compute and human labour stay stagnant. One paper uses a concept that is similar to the reproduction rate in epidemiology – in the same way that an outbreak can only keep growing if R >1 (ie when each infected person goes to infect more than one other person), explosive growth in AI capabilities can only occur if its recursive reproduction number, written as ℛ_AI, > 1. Otherwise, progress will eventually slow down.
This is in some ways a higher bar than full research automation, but in others a lower one; it doesn’t actually require every human researcher to be automated away. Because it requires that inputs like human labour and compute are held constant, it would still be possible for an AI lab to achieve extremely accelerated progress without meeting it by building more chips and/ or hiring more people. It is also probably even less useful as an operationalisable definition than a particular level of researcher automation or an entirely closed feedback loop, because, in practice, it is likely only confirmable in retrospect. Researchers remain agnostic about whether or under what conditions we’d see self-sustaining progress, and that it may be indistinguishable at first from progress that is very fast but not self-sustaining. So although this definition gets at the thing people are usually most scared of when they discuss RSI (an uncontrollable, recursive loop that results in an “intelligence explosion”), it does not seem like the kind of thing it would be possible to “ban”. Self-sustaining progress is a phenomenon, not a set of actions. To use a crude analogy, if you wanted to ban the development and release of an engineered virus, you wouldn’t make a certain level of transmissibility the threshold for that ban – especially if you had no way of knowing how contagious it would be ahead of time. Otherwise you’re waiting for people to release viruses into the population and waiting to see if they become pandemics.
A handy chart of RSI definitions (disclosure: text is Claude-generated)
So what does this mean for a ban?
Having done this deep-dive on various definitions of RSI, I feel pretty confused about the idea of “banning” RSI. Depending on how you define it, RSI can either be a set of activities (automating certain roles, using particular techniques, etc) or an outcome (a certain level of speedup or the emergence of a self-sustaining progress loop). Obviously, you can only proactively ban the former. But given that we still have uncertainty about precisely what inputs will lead to the outcome we’re afraid of, it seems easy to target the wrong thing.
But this doesn’t mean there’s nothing to be done. Our aim is to prevent a dynamic where AI is building successive generations of itself in ways that humans cannot reliably supervise. So we should severely limit the amount of authority that we cede to AIs! Any one specific intervention might end up being insufficient, so we should maximise – or at least diversify – the number of preventative measures we take. It seems like there are plenty of such interventions. At the extreme end, just don’t let frontier researchers AI-generate code (per Ezra’s suggestion). This feels like an extremely costly intervention and would have the effect of dramatically slowing progress, but everyone dying is also extremely costly. Or one lesser intervention: only use models that have been available publicly for some set period of time for internal capabilities research. For example, the AI Futures Project recommends a minimum of nine months. If your models haven’t been rigorously tested enough for public release, why should you trust them with the much higher-stakes task of accelerating your own research?
You could also cap the number of hours that agents can work autonomously per day. Cap the amount of compute that is allocated to autonomous AI R&D projects. Benchmark the AI R&D capabilities of your models frequently and publicly. If they’re getting scarily good, don’t use them internally. Design an effective embedded auditing system such that all of the above can be verified. There are probably many more things you could do here that would take many more blog posts to fully operationalise. At the very least, setting off a recursive feedback cycle should emphatically not be the goal of any frontier AI company. This should be the thing they are, above all, trying to avoid. Don’t let Claude n build Claude n+1 while you go home and knit sweaters. That’s a completely horrible idea. Treat your internal models like a cohort of untrustworthy, cheating, lying interns. Don’t give them access to too many resources. Don’t turn your back on them. Assume they are trying to screw you over.
That RSI is hard to define or predict is an argument for a very, very conservative line. It’s possible that Ezra Klein’s “no AI-generated code” proposal is more of a rhetorical salvo than a serious suggestion, but, if so, I think his point is compelling. If you’re so scared of RSI, you should be backing away from it like an angry bear in the woods. In a rebuttal to Ezra’s piece, economist Tyler Cowen asks whether “humans [are] even capable of writing the next steps of code that are required for further progress”. But given our present face-to-face-with-a-bear-in-the-woods predicament, it’s far from clear that we want to make further progress anyway. Arguably, a host of preventative measures to try and stay on the right side of the RSI-line are second best to bringing this whole thing to a grinding halt.