Xiangxin Zhou
Xiangxin Zhou on Hugging Face Daily Papers: 15 papers, 1 in the top 3 of their day, 435 upvotes.
- Best Practice Critic Optimization 17 upvotes, #9 of 2026-08-26
- Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning 35 upvotes, #8 of 2026-07-22
- MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators 17 upvotes, #16 of 2026-07-17
- Exploring the Design Space of Reward Backpropagation for Flow Matching 10 upvotes, #24 of 2026-06-23
- Rethinking the Divergence Regularization in LLM RL 33 upvotes, #11 of 2026-06-10
- Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 42 upvotes, #8 of 2026-06-10
- Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models 41 upvotes, #9 of 2026-06-10
- Reinforcing Few-step Generators via Reward-Tilted Distribution Matching 5 upvotes, #40 of 2026-05-26
- Rethinking the Trust Region in LLM Reinforcement Learning 30 upvotes, #12 of 2026-02-05
- Defeating the Training-Inference Mismatch via FP16 27 upvotes, #6 of 2025-11-03
- GEM: A Gym for Agentic LLMs 79 upvotes, #2 of 2025-10-02
- Variational Reasoning for Language Models 66 upvotes, #5 of 2025-09-29
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use 9 upvotes, #9 of 2025-08-11
- Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
- ProteinBench: A Holistic Evaluation of Protein Foundation Models 6 upvotes, #12 of 2024-09-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.