Kai Yang
Kai Yang on Hugging Face Daily Papers: 5 papers, 0 in the top 3 of their day, 146 upvotes.
- Debiased Model-based Representations for Sample-efficient Continuous Control 9 upvotes, #33 of 2026-05-13
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation 57 upvotes, #4 of 2026-02-13
- EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control 5 upvotes, #18 of 2025-11-21
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners 29 upvotes, #10 of 2025-10-01
- Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model 28 upvotes, #6 of 2023-11-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.