zhu

zhu on Hugging Face Daily Papers: 11 papers, 6 in the top 3 of their day, 839 upvotes.

  1. FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
  2. SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning 73 upvotes, #3 of 2025-09-12
  3. A Survey of Reinforcement Learning for Large Reasoning Models 156 upvotes, #1 of 2025-09-11
  4. Towards a Unified View of Large Language Model Post-Training 67 upvotes, #3 of 2025-09-05
  5. SSRL: Self-Search Reinforcement Learning 88 upvotes, #2 of 2025-08-18
  6. Reasoning with Exploration: An Entropy Perspective 26 upvotes, #8 of 2025-06-18
  7. Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space 26 upvotes, #11 of 2025-05-20
  8. TTRL: Test-Time Reinforcement Learning 96 upvotes, #2 of 2025-04-23
  9. Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models 26 upvotes, #6 of 2025-03-17
  10. MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding 19 upvotes, #6 of 2025-01-31
  11. How to Synthesize Text Data without Model Collapse? 46 upvotes, #4 of 2024-12-20

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.