Runze Liu
Runze Liu on Hugging Face Daily Papers: 6 papers, 2 in the top 3 of their day, 412 upvotes.
- ASPO: Asymmetric Importance Sampling Policy Optimization 13 upvotes, #13 of 2025-10-08
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models 12 upvotes, #24 of 2025-10-01
- A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
- Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR 19 upvotes, #14 of 2025-07-22
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning 12 upvotes, #15 of 2025-04-04
- Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 128 upvotes, #1 of 2025-02-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.