Runpeng Dai
Runpeng Dai on Hugging Face Daily Papers: 13 papers, 1 in the top 3 of their day, 381 upvotes.
- Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation 16 upvotes, #17 of 2026-09-03
- It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning 72 upvotes, #6 of 2026-09-03
- Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling 14 upvotes, #19 of 2026-06-03
- G-Zero: Self-Play for Open-Ended Generation from Zero Data 16 upvotes, #16 of 2026-05-12
- Reinforcing Multimodal Reasoning Against Visual Degradation 7 upvotes, #31 of 2026-05-12
- DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification 6 upvotes, #34 of 2026-05-12
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling 64 upvotes, #5 of 2026-05-11
- Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing 24 upvotes, #13 of 2026-02-04
- StatEval: A Comprehensive Benchmark for Large Language Models in Statistics 6 upvotes, #23 of 2025-10-13
- VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal Reasoning 19 upvotes, #16 of 2025-10-03
- CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models 28 upvotes, #5 of 2025-09-11
- Parallel-R1: Towards Parallel Thinking via Reinforcement Learning 95 upvotes, #2 of 2025-09-10
- R1-RE: Cross-Domain Relationship Extraction with RLVR 6 upvotes, #21 of 2025-07-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.