Xiangxin Zhou

Xiangxin Zhou on Hugging Face Daily Papers: 15 papers, 1 in the top 3 of their day, 435 upvotes.

  1. Best Practice Critic Optimization 17 upvotes, #9 of 2026-08-26
  2. Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning 35 upvotes, #8 of 2026-07-22
  3. MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators 17 upvotes, #16 of 2026-07-17
  4. Exploring the Design Space of Reward Backpropagation for Flow Matching 10 upvotes, #24 of 2026-06-23
  5. Rethinking the Divergence Regularization in LLM RL 33 upvotes, #11 of 2026-06-10
  6. Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 42 upvotes, #8 of 2026-06-10
  7. Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models 41 upvotes, #9 of 2026-06-10
  8. Reinforcing Few-step Generators via Reward-Tilted Distribution Matching 5 upvotes, #40 of 2026-05-26
  9. Rethinking the Trust Region in LLM Reinforcement Learning 30 upvotes, #12 of 2026-02-05
  10. Defeating the Training-Inference Mismatch via FP16 27 upvotes, #6 of 2025-11-03
  11. GEM: A Gym for Agentic LLMs 79 upvotes, #2 of 2025-10-02
  12. Variational Reasoning for Language Models 66 upvotes, #5 of 2025-09-29
  13. OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use 9 upvotes, #9 of 2025-08-11
  14. Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
  15. ProteinBench: A Holistic Evaluation of Protein Foundation Models 6 upvotes, #12 of 2024-09-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.