Epistemic Status: Speculative hypothesis about a novel anti-distillation mechanism but I think it is plausible and grounded in some inferences from research. Extrapolates from multi-teacher distillation literature to RL seeds. Untested empirically.
The topic of distillation has been a big topic in the LLM field since last year, and it is still an important unsolved issue recently as DeepSeek, Moonshot AI, MiniMax successfully performed distillation attacks on Claude. I initially wrote a quick take about this topic when I thought of this seed diversity mechanism(although that was too confident upon further research), but as I thought about it more I think the concept deserves a full post with more details and arguments as I think it holds up to a few known attacks that some known defenses are susceptible to.
The basic idea is about distilling two or more teacher models which have sufficiently different policies, which is the case for different RL seeds as the RL seed’s reasoning strategy is only a small subset of the base model[1], into a student model. You get a flattening effect in the student model as it tries to learn the contradictory policies but fails as the loss landscape is incompatible, so it is mean-seeking[2] and does not find a single good solution. Mean-seeking applies to the logits, and it follows that the internal representations are not of any mode, additionally the gradient destruction/interference arguments apply to both logits and features.
This idea can be connected to anti-distillation. The provider can deploy 10-20 RL seeds for the same model decided randomly at the start of the chat, maybe seeds can be randomized every ~100-500k cumulative tokens, but does not change in the middle of the chat for continuity/caching reasons. The models are similarly capable and the seed reset token limit is high, so the legitimate users are not harmed much or at all at inference time. However the poisoning only takes effect at training time, where the student model would face the above flattening phenomenon, which especially targets the most valuable complex reasoning, long context agentic coding abilities that only the frontier models have and the most susceptible to flattening as they require multi-step coherence.
Initially I thought that the method was ineffective against RL-judge distillation attempts, concretely used by DeepSeek extracting rubric criteria, as judging can be substituted or averaged out. However I think it is the opposite now, as RL is known to be much more sensitive than SFT when directly trained, and it would be even worse if the data is used to train a reward model that has to incorporate conflicting judgements about what is good or bad. This could lead to reward hacking problems or over-smoothing when the reward model cannot identify exceptional responses, causing a premature plateau.
However the parts concerning degradation are the main reason I frame this post as a hypothesis, there seems to be a gap in the literature and uncertain evidence. The results show that when different teacher models are distilled into a student model, there is an improvement from 1 to 2 teachers, because of ensemble benefits, while it degrades from there going to 4 teachers, because of knowledge conflict[3]. Although the authors only continue to 4 models, this knowledge conflict mechanism also predicts that degradation will continue to catastrophic levels at 10-20 models because the knowledge conflict mechanism only gets worse.
However RL seeds from the same base model are a completely different story, they are likely much closer in style and reasoning strategies, although RL is still mode-seeking and converges on different reasoning strategies and internal representations. I could not find any literature on this, so this part about RL seed distillation degradation is theoretical. RL seeds retain the same low entropy positions as the base model, and diverge at the high entropy positions that drive significant capability gains[4], so the theoretical mean-seeking concentrates at these critical high entropy positions that are important for capabilities. This mean-seeking leads to an over-smoothed and mediocre reasoning characterized by the intersection between the seeds even without the mode-switches. Additionally, the 4 teacher distillation degradation example uses around 7b models as teachers, and I think that might be far too small to create the complex reasoning capabilities that mean-seeking harms most.
Anti-distillation sampling is an inference time intervention to modify output distributions to maintain quality for legitimate users but is less useful/poisons the training data for the student.[5] When you look at it from another perspective though, it mainly perturbs high entropy tokens while keeping low entropy tokens intact, which is very similar to the idea here. Empirically it does produce a large degradation far beyond the degradation in the teacher model(4%) in the student model(12%), so the high entropy tokens perturbed does cause degradation, even when the low entropy tokens are the same. This mirrors the scenario with seed diversity where the low entropy tokens are the same, while the high entropy tokens which are load-bearing for capabilities differ.
Also, I do not think it is straightforward to conclude that greater divergence between model families vs different RL seeds leads to more vs less degradation. Between different model families, there are obvious stylistic differences that serve as an implicit context router, while the low entropy syntax is similar for RL seeds so there is no such analog. This mechanism could theoretically cause the same concept behind negative interference and unlearning when prefixes are similar yet reasoning continuations are very different, or cause single mode mean-seeking as opposed to mode-covering.
However these are important points regarding degradation that have not been verified empirically so it remains a hypothesis, whereas the later points depends on some form of degradation happening.
Current anti-distillation methods can be enumerated:
Detection and terms of service which is clearly not an effective deterrence without enforcement.
Traffic detection and account banning, which is a cat-and-mouse game favoured for the offense, for example the “hydra cluster architectures” get past imperfect account level control and rate limits to obtain the needed outputs.
Output filtering and chain-of-thought hiding, which has counters like prompting in a way that tries to expose an explicit form of the chain-of-thought, and is ineffective in general.
Anti-distillation sampling[5] which to reiterate is an inference time intervention to modify output distributions to maintain quality for legitimate users but is less useful/poisons the training data for the student. This method is countered in a few main ways.
Firstly, the perturbation is a detectable deviation from the natural distribution. The attacker could replace the affected tokens or just filter them out with a large sample size on similar prompts. This does not affect seed diversity as the collected data is from genuinely different natural distributions, the average is not a meaningful thing in the same way that can be used to restore the individual distributions.
Secondly, text can be rewritten by an AI to retain the core logic while replacing the poisoning text since it is likely unnatural. Rewriting does not affect seed diversity as the natural distributions and logic are structurally different.
Thirdly, the student model used must be similar in architecture to the proxy model used otherwise it doesn’t work, while the seed diversity defense is general.
Watermarking in general, which can be rewritten by an AI to retain the core logic while getting rid of unnatural token choices or watermarks. Rewriting does not affect seed diversity as the natural distributions and logic are structurally different.
For seed diversity to work, it is necessary that the attacker is unable to separate the data which is unlabelled into different clusters and just train on one cluster corresponding to one seed.
This method theoretically works because at standard temperature settings, empirically the single sample variance if you resend the prompt is high, you can get entirely different outputs for the same seed, which corresponds to different clusters, then it is hard to know if the different outputs are from the same seed or different seeds, especially when low entropy tokens/syntax are similar across different seeds. Trying to cluster data runs into information-theoretic limits when individual data points reside in overlapping areas, so you genuinely could not tell which cluster the data came from no matter what method you use, but there are still deep differences that don’t show up in standard interpretability methods.
Training works on the population level, so small biases in the data add up and incoherent gradients are a real problem even if they are not visible at the individual datapoint level. RL seeds share the same base model, so they naturally have similar embeddings, but this point can be engineered. However even if that did not work for RL seeds, this fundamental individual-population asymmetry between individual level identification and population level training means that there must be a point where individual differences are not identifiable but the population incoherence still degrades training, which can be engineered.
Additionally, clustering is non-linearly harder as the number of seeds used linearly increases. The degradation magnitude is likely also larger with the number of seeds as conflicting sources increases, so the defense has the advantage. Even if clustering showed that there are N clusters or seeds used, that does not help in classification. If classification accuracy is somewhat above random the result is the same as long as it is not very high.
One key property is that it works even when the method is fully public, as these points are fully general, although it has to be stress-tested first, which is important as defences will inevitably become public.
However even though many of these arguments are backed by empirical research and not speculation, the full argument hasn’t so it needs to be tested empirically and red-teamed first. Things that are not empirically verified whether the knowledge conflict mechanism that causes degradation and flattening appears with RL seeds as teachers, and if the degradation magnitude is large enough not just theoretically against a motivated adversary who will try to work around limitations. However this is complementary to other methods if the degradation magnitude is not big enough.
Epistemic Status: Speculative hypothesis about a novel anti-distillation mechanism but I think it is plausible and grounded in some inferences from research. Extrapolates from multi-teacher distillation literature to RL seeds. Untested empirically.
The topic of distillation has been a big topic in the LLM field since last year, and it is still an important unsolved issue recently as DeepSeek, Moonshot AI, MiniMax successfully performed distillation attacks on Claude. I initially wrote a quick take about this topic when I thought of this seed diversity mechanism(although that was too confident upon further research), but as I thought about it more I think the concept deserves a full post with more details and arguments as I think it holds up to a few known attacks that some known defenses are susceptible to.
The basic idea is about distilling two or more teacher models which have sufficiently different policies, which is the case for different RL seeds as the RL seed’s reasoning strategy is only a small subset of the base model[1], into a student model. You get a flattening effect in the student model as it tries to learn the contradictory policies but fails as the loss landscape is incompatible, so it is mean-seeking[2] and does not find a single good solution. Mean-seeking applies to the logits, and it follows that the internal representations are not of any mode, additionally the gradient destruction/interference arguments apply to both logits and features.
This idea can be connected to anti-distillation. The provider can deploy 10-20 RL seeds for the same model decided randomly at the start of the chat, maybe seeds can be randomized every ~100-500k cumulative tokens, but does not change in the middle of the chat for continuity/caching reasons. The models are similarly capable and the seed reset token limit is high, so the legitimate users are not harmed much or at all at inference time. However the poisoning only takes effect at training time, where the student model would face the above flattening phenomenon, which especially targets the most valuable complex reasoning, long context agentic coding abilities that only the frontier models have and the most susceptible to flattening as they require multi-step coherence.
Initially I thought that the method was ineffective against RL-judge distillation attempts, concretely used by DeepSeek extracting rubric criteria, as judging can be substituted or averaged out. However I think it is the opposite now, as RL is known to be much more sensitive than SFT when directly trained, and it would be even worse if the data is used to train a reward model that has to incorporate conflicting judgements about what is good or bad. This could lead to reward hacking problems or over-smoothing when the reward model cannot identify exceptional responses, causing a premature plateau.
However the parts concerning degradation are the main reason I frame this post as a hypothesis, there seems to be a gap in the literature and uncertain evidence. The results show that when different teacher models are distilled into a student model, there is an improvement from 1 to 2 teachers, because of ensemble benefits, while it degrades from there going to 4 teachers, because of knowledge conflict[3]. Although the authors only continue to 4 models, this knowledge conflict mechanism also predicts that degradation will continue to catastrophic levels at 10-20 models because the knowledge conflict mechanism only gets worse.
However RL seeds from the same base model are a completely different story, they are likely much closer in style and reasoning strategies, although RL is still mode-seeking and converges on different reasoning strategies and internal representations. I could not find any literature on this, so this part about RL seed distillation degradation is theoretical. RL seeds retain the same low entropy positions as the base model, and diverge at the high entropy positions that drive significant capability gains[4], so the theoretical mean-seeking concentrates at these critical high entropy positions that are important for capabilities. This mean-seeking leads to an over-smoothed and mediocre reasoning characterized by the intersection between the seeds even without the mode-switches. Additionally, the 4 teacher distillation degradation example uses around 7b models as teachers, and I think that might be far too small to create the complex reasoning capabilities that mean-seeking harms most.
Anti-distillation sampling is an inference time intervention to modify output distributions to maintain quality for legitimate users but is less useful/poisons the training data for the student.[5] When you look at it from another perspective though, it mainly perturbs high entropy tokens while keeping low entropy tokens intact, which is very similar to the idea here. Empirically it does produce a large degradation far beyond the degradation in the teacher model(4%) in the student model(12%), so the high entropy tokens perturbed does cause degradation, even when the low entropy tokens are the same. This mirrors the scenario with seed diversity where the low entropy tokens are the same, while the high entropy tokens which are load-bearing for capabilities differ.
Also, I do not think it is straightforward to conclude that greater divergence between model families vs different RL seeds leads to more vs less degradation. Between different model families, there are obvious stylistic differences that serve as an implicit context router, while the low entropy syntax is similar for RL seeds so there is no such analog. This mechanism could theoretically cause the same concept behind negative interference and unlearning when prefixes are similar yet reasoning continuations are very different, or cause single mode mean-seeking as opposed to mode-covering.
However these are important points regarding degradation that have not been verified empirically so it remains a hypothesis, whereas the later points depends on some form of degradation happening.
Current anti-distillation methods can be enumerated:
Anti-distillation sampling[5] which to reiterate is an inference time intervention to modify output distributions to maintain quality for legitimate users but is less useful/poisons the training data for the student. This method is countered in a few main ways.
Firstly, the perturbation is a detectable deviation from the natural distribution. The attacker could replace the affected tokens or just filter them out with a large sample size on similar prompts. This does not affect seed diversity as the collected data is from genuinely different natural distributions, the average is not a meaningful thing in the same way that can be used to restore the individual distributions.
Secondly, text can be rewritten by an AI to retain the core logic while replacing the poisoning text since it is likely unnatural. Rewriting does not affect seed diversity as the natural distributions and logic are structurally different.
Thirdly, the student model used must be similar in architecture to the proxy model used otherwise it doesn’t work, while the seed diversity defense is general.
For seed diversity to work, it is necessary that the attacker is unable to separate the data which is unlabelled into different clusters and just train on one cluster corresponding to one seed.
This method theoretically works because at standard temperature settings, empirically the single sample variance if you resend the prompt is high, you can get entirely different outputs for the same seed, which corresponds to different clusters, then it is hard to know if the different outputs are from the same seed or different seeds, especially when low entropy tokens/syntax are similar across different seeds. Trying to cluster data runs into information-theoretic limits when individual data points reside in overlapping areas, so you genuinely could not tell which cluster the data came from no matter what method you use, but there are still deep differences that don’t show up in standard interpretability methods.
Training works on the population level, so small biases in the data add up and incoherent gradients are a real problem even if they are not visible at the individual datapoint level. RL seeds share the same base model, so they naturally have similar embeddings, but this point can be engineered. However even if that did not work for RL seeds, this fundamental individual-population asymmetry between individual level identification and population level training means that there must be a point where individual differences are not identifiable but the population incoherence still degrades training, which can be engineered.
Additionally, clustering is non-linearly harder as the number of seeds used linearly increases. The degradation magnitude is likely also larger with the number of seeds as conflicting sources increases, so the defense has the advantage. Even if clustering showed that there are N clusters or seeds used, that does not help in classification. If classification accuracy is somewhat above random the result is the same as long as it is not very high.
One key property is that it works even when the method is fully public, as these points are fully general, although it has to be stress-tested first, which is important as defences will inevitably become public.
However even though many of these arguments are backed by empirical research and not speculation, the full argument hasn’t so it needs to be tested empirically and red-teamed first. Things that are not empirically verified whether the knowledge conflict mechanism that causes degradation and flattening appears with RL seeds as teachers, and if the degradation magnitude is large enough not just theoretically against a motivated adversary who will try to work around limitations. However this is complementary to other methods if the degradation magnitude is not big enough.
Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? (Yue et al., 2025)
Agree to disagree: Adaptive ensemble knowledge distillation in gradient space (Du et al., 2020)
Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs (Jin et al., 2026)
Beyond the 80/20 rule: High-Entropy minority tokens drive effective reinforcement learning for LLM reasoning (Wang et al., 2025)
Antidistillation sampling (Savani et al., 2025)