Yanxi Chen
Yanxi Chen on Hugging Face Daily Papers: 9 papers, 0 in the top 3 of their day, 97 upvotes.
- Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning 8 upvotes, #27 of 2026-06-23
- SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees 4 upvotes, #30 of 2026-02-09
- On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models 52 upvotes, #5 of 2026-02-09
- Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends 6 upvotes, #31 of 2025-10-03
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting 6 upvotes, #12 of 2025-08-21
- Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models 9 upvotes, #27 of 2025-05-26
- A Simple and Provable Scaling Law for the Test-Time Compute of Large Language Models 4 upvotes, #22 of 2024-12-03
- EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models 4 upvotes, #11 of 2024-02-02
- EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism 7 upvotes, #7 of 2023-12-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.