Kaiyan Zhang

Kaiyan Zhang on Hugging Face Daily Papers: 24 papers, 12 in the top 3 of their day, 1,689 upvotes.

  1. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering 181 upvotes, #4 of 2026-07-31
  2. NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? 62 upvotes, #2 of 2026-06-24
  3. EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions 79 upvotes, #2 of 2026-06-23
  4. Post-Trained MoE Can Skip Half Experts via Self-Distillation 30 upvotes, #10 of 2026-05-19
  5. How Far Can Unsupervised RLVR Scale LLM Training? 52 upvotes, #4 of 2026-03-10
  6. JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
  7. Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models 12 upvotes, #24 of 2025-10-01
  8. FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
  9. SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning 73 upvotes, #3 of 2025-09-12
  10. A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
  11. Towards a Unified View of Large Language Model Post-Training 67 upvotes, #3 of 2025-09-05
  12. SSRL: Self-Search Reinforcement Learning 88 upvotes, #2 of 2025-08-18
  13. TTRL: Test-Time Reinforcement Learning 96 upvotes, #2 of 2025-04-23
  14. GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning 12 upvotes, #15 of 2025-04-04
  15. Video-T1: Test-Time Scaling for Video Generation 84 upvotes, #2 of 2025-03-25
  16. Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models 26 upvotes, #6 of 2025-03-17
  17. Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 128 upvotes, #1 of 2025-02-11
  18. Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04
  19. MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding 19 upvotes, #6 of 2025-01-31
  20. Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization 35 upvotes, #1 of 2024-12-25
  21. How to Synthesize Text Data without Model Collapse? 46 upvotes, #4 of 2024-12-20
  22. Free Process Rewards without Process Labels 26 upvotes, #4 of 2024-12-04
  23. Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices 14 upvotes, #8 of 2024-10-16
  24. Towards Building Specialized Generalist AI with System 1 and System 2 Fusion 9 upvotes, #14 of 2024-07-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.