Chujie Zheng

Chujie Zheng on Hugging Face Daily Papers: 16 papers, 7 in the top 3 of their day, 1,565 upvotes.

  1. QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents 43 upvotes, #15 of 2026-09-29
  2. Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding 24 upvotes, #11 of 2026-06-23
  3. Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models 11 upvotes, #19 of 2026-02-09
  4. Stabilizing Reinforcement Learning with LLMs: Formulation and Practices 83 upvotes, #4 of 2025-12-02
  5. Soft Adaptive Policy Optimization 33 upvotes, #6 of 2025-11-26
  6. Group Sequence Policy Optimization 257 upvotes, #1 of 2025-07-25
  7. A Survey on Latent Reasoning 78 upvotes, #2 of 2025-07-09
  8. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning 144 upvotes, #1 of 2025-06-03
  9. BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs 11 upvotes, #23 of 2025-05-22
  10. Qwen3 Technical Report 152 upvotes, #1 of 2025-05-19
  11. WorldPM: Scaling Human Preference Modeling 33 upvotes, #5 of 2025-05-16
  12. SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 92 upvotes, #3 of 2025-02-21
  13. The Lessons of Developing Process Reward Models in Mathematical Reasoning 83 upvotes, #1 of 2025-01-14
  14. ProcessBench: Identifying Process Errors in Mathematical Reasoning 61 upvotes, #2 of 2024-12-10
  15. Yi-Lightning Technical Report 21 upvotes, #6 of 2024-12-02
  16. Safe Unlearning: A Surprisingly Effective and Generalizable Solution to Defend Against Jailbreak Attacks 9 upvotes, #12 of 2024-07-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.