Ran Xu
Ran Xu on Hugging Face Daily Papers: 5 papers, 0 in the top 3 of their day, 52 upvotes.
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training 14 upvotes, #31 of 2026-02-03
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs 9 upvotes, #18 of 2025-11-26
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment 7 upvotes, #33 of 2025-10-10
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play 6 upvotes, #48 of 2025-09-30
- MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale 4 upvotes, #30 of 2025-06-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.