Jian Hu
Jian Hu on Hugging Face Daily Papers: 5 papers, 4 in the top 3 of their day, 423 upvotes.
- Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text 89 upvotes, #2 of 2026-02-02
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 118 upvotes, #1 of 2025-06-02
- LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs 19 upvotes, #13 of 2025-04-22
- REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models 79 upvotes, #1 of 2025-01-08
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework 31 upvotes, #3 of 2024-05-21
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.