Zichen
Zichen on Hugging Face Daily Papers: 12 papers, 4 in the top 3 of their day, 541 upvotes.
- GEM: A Gym for Agentic LLMs 79 upvotes, #2 of 2025-10-02
- Language Models Can Learn from Verbal Feedback Without Scalar Rewards 64 upvotes, #6 of 2025-09-29
- Variational Reasoning for Language Models 66 upvotes, #5 of 2025-09-29
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning 43 upvotes, #2 of 2025-07-01
- SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis 51 upvotes, #4 of 2025-06-04
- Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization 34 upvotes, #3 of 2025-05-21
- Efficient Process Reward Model Training via Active Learning 13 upvotes, #14 of 2025-04-16
- Understanding R1-Zero-Like Training: A Critical Perspective 36 upvotes, #7 of 2025-04-03
- Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs 11 upvotes, #14 of 2025-02-18
- Sample-Efficient Alignment for LLMs 10 upvotes, #6 of 2024-11-06
- Bootstrapping Language Models with DPO Implicit Rewards 34 upvotes, #3 of 2024-06-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.