Yanxi Chen

Yanxi Chen on Hugging Face Daily Papers: 9 papers, 0 in the top 3 of their day, 97 upvotes.

  1. Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning 8 upvotes, #27 of 2026-06-23
  2. SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees 4 upvotes, #30 of 2026-02-09
  3. On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models 52 upvotes, #5 of 2026-02-09
  4. Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends 6 upvotes, #31 of 2025-10-03
  5. On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting 6 upvotes, #12 of 2025-08-21
  6. Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models 9 upvotes, #27 of 2025-05-26
  7. A Simple and Provable Scaling Law for the Test-Time Compute of Large Language Models 4 upvotes, #22 of 2024-12-03
  8. EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models 4 upvotes, #11 of 2024-02-02
  9. EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism 7 upvotes, #7 of 2023-12-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.