Kaiyan Zhang
Kaiyan Zhang on Hugging Face Daily Papers: 24 papers, 12 in the top 3 of their day, 1,689 upvotes.
- Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering 181 upvotes, #4 of 2026-07-31
- NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? 62 upvotes, #2 of 2026-06-24
- EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions 79 upvotes, #2 of 2026-06-23
- Post-Trained MoE Can Skip Half Experts via Self-Distillation 30 upvotes, #10 of 2026-05-19
- How Far Can Unsupervised RLVR Scale LLM Training? 52 upvotes, #4 of 2026-03-10
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models 12 upvotes, #24 of 2025-10-01
- FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning 73 upvotes, #3 of 2025-09-12
- A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
- Towards a Unified View of Large Language Model Post-Training 67 upvotes, #3 of 2025-09-05
- SSRL: Self-Search Reinforcement Learning 88 upvotes, #2 of 2025-08-18
- TTRL: Test-Time Reinforcement Learning 96 upvotes, #2 of 2025-04-23
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning 12 upvotes, #15 of 2025-04-04
- Video-T1: Test-Time Scaling for Video Generation 84 upvotes, #2 of 2025-03-25
- Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models 26 upvotes, #6 of 2025-03-17
- Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 128 upvotes, #1 of 2025-02-11
- Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding 19 upvotes, #6 of 2025-01-31
- Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization 35 upvotes, #1 of 2024-12-25
- How to Synthesize Text Data without Model Collapse? 46 upvotes, #4 of 2024-12-20
- Free Process Rewards without Process Labels 26 upvotes, #4 of 2024-12-04
- Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices 14 upvotes, #8 of 2024-10-16
- Towards Building Specialized Generalist AI with System 1 and System 2 Fusion 9 upvotes, #14 of 2024-07-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.