This is a special post for quick takes by Asdfer. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
With the recent news about Claude's distillation attempt I was thinking of a structural way to counter it.
There is this phenomenon about high RL seed variance where if you take two RL seeds which are good at different types of questions, and distill both of them into a single student model, you would expect the student model to score well on the union set. But this isn't the case, the student model actually scores worse than either teacher because of the flattening effect, it tries to mimic both teachers' features, but since the loss landscape is incompatible, it is just mean-seeking and fails to find a good solution.
This seems applicable to anti-distillation as well, It is feasible to deploy two or more RL seeds for the same model decided randomly at the start of the chat(but does not change in the middle of the chat for continuity/user experience). The models are similarly capable so there should be little or no harm to legitimate users at inference time. Distillation attempts would face the above flattening and degradation phenomenon.
One condition for this to work is low distinguishability, where the attacker must not be able to accurately determine which text was generated by which model. Distillation attempts usually make use of short contexts for encapsulation, as the earlier context may corrupt future distillation data, therefore there is less data that can distinguish the models. A standard temperature of 1 also significantly increases obfuscation. Another attack factor using a probe matching a response with a known model is also disrupted by these standard non-zero temperature settings. And even if some data is distinguishable it is likely that some data simply cannot be, which makes the mechanism still effective. Even then, the strongest point is actually the lack of ground truth to even do classification training, as the model identifier does not even need to be revealed at all throughout.
This is likely to be less effective for distillation attempts which use the model as an RL judge, but this case is a significant minority and may be less effective compared to the standard distillation usage.