Bingxiang He
Bingxiang He on Hugging Face Daily Papers: 18 papers, 6 in the top 3 of their day, 1,104 upvotes.
- Diffusion Reward Models 35 upvotes, #22 of 2026-09-29
- StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? 19 upvotes, #34 of 2026-09-09
- Rethinking On-Policy Distillation of Large Language Models II: One Training Example 91 upvotes, #8 of 2026-09-04
- Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation 21 upvotes, #10 of 2026-08-27
- SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? 14 upvotes, #16 of 2026-08-27
- On-policy Distillation with Verifiable Reward 17 upvotes, #9 of 2026-08-26
- MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement 8 upvotes, #22 of 2026-08-19
- PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments 28 upvotes, #11 of 2026-08-18
- Weak-to-Strong Generalization via Direct On-Policy Distillation 134 upvotes, #1 of 2026-07-14
- NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? 62 upvotes, #2 of 2026-06-24
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe 85 upvotes, #2 of 2026-04-15
- How Far Can Unsupervised RLVR Scale LLM Training? 52 upvotes, #4 of 2026-03-10
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents 20 upvotes, #4 of 2025-11-06
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe 46 upvotes, #4 of 2025-09-24
- A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
- MiniCPM4: Ultra-Efficient LLMs on End Devices 78 upvotes, #3 of 2025-06-10
- Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.