Antidistillation Sampling
Yash Savani, Asher Trockman, Zhili Feng, Avi Schwarzschild, Alex Robey, Marc Finzi, J. Zico Kolter
Antidistillation Sampling: 60 upvotes on Hugging Face Daily Papers, #2 of 23 papers on 2025-04-18. Day-by-day upvote history.
Frontier models that generate extended reasoning traces inadvertently produce rich token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. Antidistillation sampling provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's practical utility. For further details, see https://antidistillation.com.
Paper page on Hugging Face · arXiv
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.