garyzhang
garyzhang on Hugging Face Daily Papers: 5 papers, 0 in the top 3 of their day, 74 upvotes.
- On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models 52 upvotes, #5 of 2026-02-09
- Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends 6 upvotes, #31 of 2025-10-03
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting 6 upvotes, #12 of 2025-08-21
- Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models 9 upvotes, #27 of 2025-05-26
- The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective 10 upvotes, #12 of 2024-07-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.