zhu
zhu on Hugging Face Daily Papers: 11 papers, 6 in the top 3 of their day, 839 upvotes.
- FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning 73 upvotes, #3 of 2025-09-12
- A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
- Towards a Unified View of Large Language Model Post-Training 67 upvotes, #3 of 2025-09-05
- SSRL: Self-Search Reinforcement Learning 88 upvotes, #2 of 2025-08-18
- Reasoning with Exploration: An Entropy Perspective 26 upvotes, #8 of 2025-06-18
- Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space 26 upvotes, #11 of 2025-05-20
- TTRL: Test-Time Reinforcement Learning 96 upvotes, #2 of 2025-04-23
- Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models 26 upvotes, #6 of 2025-03-17
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding 19 upvotes, #6 of 2025-01-31
- How to Synthesize Text Data without Model Collapse? 46 upvotes, #4 of 2024-12-20
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.