Zhiyuan Hu

Zhiyuan Hu on Hugging Face Daily Papers: 6 papers, 3 in the top 3 of their day, 490 upvotes.

  1. Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs 140 upvotes, #3 of 2026-01-16
  2. Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning 82 upvotes, #4 of 2026-01-16
  3. Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models 113 upvotes, #1 of 2025-05-16
  4. JudgeLRM: Large Reasoning Models as a Judge 55 upvotes, #2 of 2025-04-02
  5. Natural Language Reinforcement Learning 25 upvotes, #5 of 2024-11-22
  6. LongRecipe: Recipe for Efficient Long Context Generalization in Large Languge Models 37 upvotes, #4 of 2024-09-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.