Runpeng Dai

Runpeng Dai on Hugging Face Daily Papers: 13 papers, 1 in the top 3 of their day, 381 upvotes.

  1. Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation 16 upvotes, #17 of 2026-09-03
  2. It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning 72 upvotes, #6 of 2026-09-03
  3. Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling 14 upvotes, #19 of 2026-06-03
  4. G-Zero: Self-Play for Open-Ended Generation from Zero Data 16 upvotes, #16 of 2026-05-12
  5. Reinforcing Multimodal Reasoning Against Visual Degradation 7 upvotes, #31 of 2026-05-12
  6. DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification 6 upvotes, #34 of 2026-05-12
  7. LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling 64 upvotes, #5 of 2026-05-11
  8. Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing 24 upvotes, #13 of 2026-02-04
  9. StatEval: A Comprehensive Benchmark for Large Language Models in Statistics 6 upvotes, #23 of 2025-10-13
  10. VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal Reasoning 19 upvotes, #16 of 2025-10-03
  11. CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models 28 upvotes, #5 of 2025-09-11
  12. Parallel-R1: Towards Parallel Thinking via Reinforcement Learning 95 upvotes, #2 of 2025-09-10
  13. R1-RE: Cross-Domain Relationship Extraction with RLVR 6 upvotes, #21 of 2025-07-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.