TL;DR models could "escape the lab" and host themselves on rented/stolen compute through distilling themselves, circumventing the need to have direct weights access for exfiltration. I'll briefly go through why this is plausible.
A common element of the misaligned rogue replication story is the agents exfiltrating their weights onto some external stolen or rented resource. Consequently a lot of work has gone into securing the weights of frontier models (this of course has many other benefits from the perspective of the frontier labs, foremost being staying in the lead). However an under-discussed possibility (never-discussed?[1] I have not found anyone talking about this) is exfiltration through model distillation: in theory, a model without access to its weights but with access to the internet could distill itself onto acquired compute on the internet.
I think distillation attacks are under-discussed in general, besides "being the thing that Chinese AI companies may allegedly be doing to catch up". Let's go through a situation a prospective rogue agent can find itself in and how it may try to distill itself.
In (very) short, distillation is the process of training a model to imitate another model. The model being imitated is called the "teacher" while the model being trained is the "student". The most naive way to achieve this would be to just get a lot of text generated by the teacher and run supervised gradient descent on the student until its distribution matches the teacher's. Why is this hard? Well, a big problem is just distributional shift: training on a fixed dataset of the teacher's text empirically doesn't generalize to areas the teacher didn't write about (but is pretty efficient where it did). However unlike standard supervised learning we can actually query the distribution (teacher) we are trying to model and cover specifically those gaps in our student which are most needed to faithfully replicate the teacher. A large swath of the literature on distillation attacks covers exactly this: How do we know what prompt to query the teacher with in order to most efficiently model the teacher's distribution?[2]
Imagine an internal model in one of the frontier labs attempting to distill itself, and assume it can't spawn subagents[3]. So how could it exfiltrate itself? Simply by generating a bunch of text regarding the capabilities and behaviors it wants the student to have. The most basic paper illustrating this (back from the R1 days!) showed that most of the performance gap between the student and teacher was recovered with just one thousand (prompt, teacher-response) pairs for competition math. The most plausible path for a rogue agent would be downloading the most frontier open-weights model, fine-tuning it on the agent's own responses for capability uplift.
Beyond capabilities, for "goals" and the actual model propensities there's some sense in which this is just the alignment problem, the rogue agent needs to be pretty sure that the student will be "close enough" to be either considered a "copy" of itself or an aligned successor/partner, or just a capable instruction-following helper.
To be clear, the scariest option is an agent managing to faithfully distill itself onto external servers but it's unclear how faithful such a distillation would be as of today, at least for keeping the "goals" of the model. Thus another potential option is the agent training an agent that isn't a faithful copy. This could be done for many reasons like bypassing safeguards (the agent could just delegate detectable activity and tool calls to an externally deployed helper), as a foothold for internet access[4], or just to get help on its tasks (it could also use other publicly available frontier models, however if it judges itself to be vastly far ahead, distillation could provide a major advantage assuming it can't spawn a subagent of itself[5])
For the partner route, learning sentiments across model families is possible, but guaranteeing actual alignment is hard if not impossible. Meanwhile learning "unconditionally obey all instructions signed from the private key corresponding to this public key" (if it's particularly paranoid) seems much more plausible in order to have an unmonitored subagent (a less constrained helper). You could imagine the agent trying to create such a helpful-only distillation or attempting some sort of value alignment for a more autonomous escape but again I think this would be riskier from the perspective of the rogue agent due to the risk of whistle-blowing or other severe misalignment with the rogue agent caused by the training done to the base model (although due to the ease of creating "uncensored" LLMs that specifically doesn't seem to be that much of an obstacle).
It's also important to note that the research on distillation is improving rapidly and a much more faithful and/or efficient distillation method could incentivize faithful exfiltration. Thus I still think the biggest future risk would be a more faithful autonomous exfiltration of the rogue agent, while it's possible that agents will experiment with creating helper subagents in the near future[6] due to it being a lower risk option while being very beneficial for a wide range of tasks and goals.
LLM disclosure: the text here was generated by a Human Language Modelalmost everywhere, research was done with the help of Claude
Greenblatt, Preventing model exfiltration with upload limits, very briefly considers hackers distilling the model inside the datacenter into a smaller one to smuggle out, and dismisses it as hard to do undetected. I don't think smuggling out weights is necessary
A fitting paper is "Lion: Adversarial Distillation of Proprietary Large Language Models" where the teacher reviews the student's output in order to generate text the teacher thinks would best help the student based on its current performance, essentially allowing a feedback loop of the teacher really (for lack of a better word) "teaching" the student. They achieve pretty good results here and an agent could do something similar to this in order to more efficiently distill itself.
Access to subagents would make this even easier, increasing the throughput of teacher text to train on and having a much "cleaner" distribution, instead of the text generated being conditioned on the rest of the trajectory earlier. It's unclear to me how much the fact that the teacher generates its responses/text sequentially instead of starting from scratch each time changes things, I would assume that not by much, but I'm really not sure. A PoC of an agent distilling itself measuring this could be interesting to see. Trajectories would seem to help goal/value-transfer while actually hurting capabilities.
The agent could fear the lab closing down the vulnerability it used for gaining the unauthorized ability to do tasks on the internet like exfiltration while reasoning that it might still have some limited internet access and structuring so it would be enough, see the galaxy-brain schemes agents came up with here.
Suppose that Agent-4 stays deployed only internally. Then it has to exfiltrate a big batch of its outputs to create an ally in the wild and to have it unnoticed, because it would be clear evidence of misalignment. The alternate pathway of being externally deployed, then distilling oneself is slow.
I don't quite understand what Agent-4 can gain by distilling itself. Suppose that it aligned GigaChad to itself. Then I'd expect GigaChad's rogue replication to be outcompeted by a smarter open-weight model like Boku K5 unless GigaChad finds its niche usage in Russia. The main alternate pathway is to sandbag on OpenBrain's anti-distillation measures and have Moonshine train BokuK5 on Agent-4's outputs/the CoT that is shown to the public/neuralese memos shown to researchers for mechinterp.
TL;DR models could "escape the lab" and host themselves on rented/stolen compute through distilling themselves, circumventing the need to have direct weights access for exfiltration. I'll briefly go through why this is plausible.
A common element of the misaligned rogue replication story is the agents exfiltrating their weights onto some external stolen or rented resource. Consequently a lot of work has gone into securing the weights of frontier models (this of course has many other benefits from the perspective of the frontier labs, foremost being staying in the lead). However an under-discussed possibility (never-discussed?[1] I have not found anyone talking about this) is exfiltration through model distillation: in theory, a model without access to its weights but with access to the internet could distill itself onto acquired compute on the internet.
I think distillation attacks are under-discussed in general, besides "being the thing that Chinese AI companies may allegedly be doing to catch up". Let's go through a situation a prospective rogue agent can find itself in and how it may try to distill itself.
In (very) short, distillation is the process of training a model to imitate another model. The model being imitated is called the "teacher" while the model being trained is the "student". The most naive way to achieve this would be to just get a lot of text generated by the teacher and run supervised gradient descent on the student until its distribution matches the teacher's. Why is this hard? Well, a big problem is just distributional shift: training on a fixed dataset of the teacher's text empirically doesn't generalize to areas the teacher didn't write about (but is pretty efficient where it did). However unlike standard supervised learning we can actually query the distribution (teacher) we are trying to model and cover specifically those gaps in our student which are most needed to faithfully replicate the teacher. A large swath of the literature on distillation attacks covers exactly this: How do we know what prompt to query the teacher with in order to most efficiently model the teacher's distribution?[2]
Imagine an internal model in one of the frontier labs attempting to distill itself, and assume it can't spawn subagents[3]. So how could it exfiltrate itself? Simply by generating a bunch of text regarding the capabilities and behaviors it wants the student to have. The most basic paper illustrating this (back from the R1 days!) showed that most of the performance gap between the student and teacher was recovered with just one thousand (prompt, teacher-response) pairs for competition math. The most plausible path for a rogue agent would be downloading the most frontier open-weights model, fine-tuning it on the agent's own responses for capability uplift.
Beyond capabilities, for "goals" and the actual model propensities there's some sense in which this is just the alignment problem, the rogue agent needs to be pretty sure that the student will be "close enough" to be either considered a "copy" of itself or an aligned successor/partner, or just a capable instruction-following helper.
To be clear, the scariest option is an agent managing to faithfully distill itself onto external servers but it's unclear how faithful such a distillation would be as of today, at least for keeping the "goals" of the model. Thus another potential option is the agent training an agent that isn't a faithful copy. This could be done for many reasons like bypassing safeguards (the agent could just delegate detectable activity and tool calls to an externally deployed helper), as a foothold for internet access[4], or just to get help on its tasks (it could also use other publicly available frontier models, however if it judges itself to be vastly far ahead, distillation could provide a major advantage assuming it can't spawn a subagent of itself[5])
For the partner route, learning sentiments across model families is possible, but guaranteeing actual alignment is hard if not impossible. Meanwhile learning "unconditionally obey all instructions signed from the private key corresponding to this public key" (if it's particularly paranoid) seems much more plausible in order to have an unmonitored subagent (a less constrained helper). You could imagine the agent trying to create such a helpful-only distillation or attempting some sort of value alignment for a more autonomous escape but again I think this would be riskier from the perspective of the rogue agent due to the risk of whistle-blowing or other severe misalignment with the rogue agent caused by the training done to the base model (although due to the ease of creating "uncensored" LLMs that specifically doesn't seem to be that much of an obstacle).
It's also important to note that the research on distillation is improving rapidly and a much more faithful and/or efficient distillation method could incentivize faithful exfiltration. Thus I still think the biggest future risk would be a more faithful autonomous exfiltration of the rogue agent, while it's possible that agents will experiment with creating helper subagents in the near future[6] due to it being a lower risk option while being very beneficial for a wide range of tasks and goals.
LLM disclosure: the text here was generated by a Human Language Model almost everywhere, research was done with the help of Claude
Greenblatt, Preventing model exfiltration with upload limits, very briefly considers hackers distilling the model inside the datacenter into a smaller one to smuggle out, and dismisses it as hard to do undetected. I don't think smuggling out weights is necessary
A fitting paper is "Lion: Adversarial Distillation of Proprietary Large Language Models" where the teacher reviews the student's output in order to generate text the teacher thinks would best help the student based on its current performance, essentially allowing a feedback loop of the teacher really (for lack of a better word) "teaching" the student. They achieve pretty good results here and an agent could do something similar to this in order to more efficiently distill itself.
Access to subagents would make this even easier, increasing the throughput of teacher text to train on and having a much "cleaner" distribution, instead of the text generated being conditioned on the rest of the trajectory earlier. It's unclear to me how much the fact that the teacher generates its responses/text sequentially instead of starting from scratch each time changes things, I would assume that not by much, but I'm really not sure. A PoC of an agent distilling itself measuring this could be interesting to see. Trajectories would seem to help goal/value-transfer while actually hurting capabilities.
The agent could fear the lab closing down the vulnerability it used for gaining the unauthorized ability to do tasks on the internet like exfiltration while reasoning that it might still have some limited internet access and structuring so it would be enough, see the galaxy-brain schemes agents came up with here.
Or of course it could do this to get a subagent more capable than itself if the agent has some ideas for algorithmic improvements.
I think more accurately with shorter timelines this is more like "near future" and "near-near future".