Runze Liu

Runze Liu on Hugging Face Daily Papers: 6 papers, 2 in the top 3 of their day, 412 upvotes.

  1. ASPO: Asymmetric Importance Sampling Policy Optimization 13 upvotes, #13 of 2025-10-08
  2. Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models 12 upvotes, #24 of 2025-10-01
  3. A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
  4. Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR 19 upvotes, #14 of 2025-07-22
  5. GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning 12 upvotes, #15 of 2025-04-04
  6. Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 128 upvotes, #1 of 2025-02-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.