Bingxiang He

Bingxiang He on Hugging Face Daily Papers: 18 papers, 6 in the top 3 of their day, 1,104 upvotes.

  1. Diffusion Reward Models 35 upvotes, #22 of 2026-09-29
  2. StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? 19 upvotes, #34 of 2026-09-09
  3. Rethinking On-Policy Distillation of Large Language Models II: One Training Example 91 upvotes, #8 of 2026-09-04
  4. Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation 21 upvotes, #10 of 2026-08-27
  5. SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? 14 upvotes, #16 of 2026-08-27
  6. On-policy Distillation with Verifiable Reward 17 upvotes, #9 of 2026-08-26
  7. MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement 8 upvotes, #22 of 2026-08-19
  8. PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments 28 upvotes, #11 of 2026-08-18
  9. Weak-to-Strong Generalization via Direct On-Policy Distillation 134 upvotes, #1 of 2026-07-14
  10. NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? 62 upvotes, #2 of 2026-06-24
  11. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe 85 upvotes, #2 of 2026-04-15
  12. How Far Can Unsupervised RLVR Scale LLM Training? 52 upvotes, #4 of 2026-03-10
  13. JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
  14. CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents 20 upvotes, #4 of 2025-11-06
  15. MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe 46 upvotes, #4 of 2025-09-24
  16. A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
  17. MiniCPM4: Ultra-Efficient LLMs on End Devices 78 upvotes, #3 of 2025-06-10
  18. Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.