Zichen

Zichen on Hugging Face Daily Papers: 12 papers, 4 in the top 3 of their day, 541 upvotes.

  1. GEM: A Gym for Agentic LLMs 79 upvotes, #2 of 2025-10-02
  2. Language Models Can Learn from Verbal Feedback Without Scalar Rewards 64 upvotes, #6 of 2025-09-29
  3. Variational Reasoning for Language Models 66 upvotes, #5 of 2025-09-29
  4. SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning 43 upvotes, #2 of 2025-07-01
  5. SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis 51 upvotes, #4 of 2025-06-04
  6. Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
  7. Optimizing Anytime Reasoning via Budget Relative Policy Optimization 34 upvotes, #3 of 2025-05-21
  8. Efficient Process Reward Model Training via Active Learning 13 upvotes, #14 of 2025-04-16
  9. Understanding R1-Zero-Like Training: A Critical Perspective 36 upvotes, #7 of 2025-04-03
  10. Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs 11 upvotes, #14 of 2025-02-18
  11. Sample-Efficient Alignment for LLMs 10 upvotes, #6 of 2024-11-06
  12. Bootstrapping Language Models with DPO Implicit Rewards 34 upvotes, #3 of 2024-06-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.