Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Wenze Lin, Jiyuan Long, Jiale Zhao, Shenzhi Wang, Xitai Jiang, Ce Luo, Rui Lan, Qianli Ma, Fukang Wen, Hui Wu, Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Gao Huang

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?: 1 upvotes on Hugging Face Daily Papers, #102 of 109 papers on 2026-09-29. Day-by-day upvote history.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.