XYX
XYX on Hugging Face Daily Papers: 4 papers, 0 in the top 3 of their day, 26 upvotes.
- Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training 10 upvotes, #30 of 2026-05-13
- PACED: Distillation at the Frontier of Student Competence 4 upvotes, #32 of 2026-03-13
- On-Policy Self-Distillation for Reasoning Compression 5 upvotes, #17 of 2026-03-06
- Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning 5 upvotes, #19 of 2026-02-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.