Zhiyuan Hu
Zhiyuan Hu on Hugging Face Daily Papers: 6 papers, 3 in the top 3 of their day, 490 upvotes.
- Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs 140 upvotes, #3 of 2026-01-16
- Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning 82 upvotes, #4 of 2026-01-16
- Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models 113 upvotes, #1 of 2025-05-16
- JudgeLRM: Large Reasoning Models as a Judge 55 upvotes, #2 of 2025-04-02
- Natural Language Reinforcement Learning 25 upvotes, #5 of 2024-11-22
- LongRecipe: Recipe for Efficient Long Context Generalization in Large Languge Models 37 upvotes, #4 of 2024-09-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.