Choi
Choi on Hugging Face Daily Papers: 2 papers, 0 in the top 3 of their day, 16 upvotes.
- KL for a KL: On-Policy Distillation with Control Variate Baseline 19 upvotes, #14 of 2026-05-14
- Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States 16 upvotes, #20 of 2026-05-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.