Ganqu Cui
Ganqu Cui on Hugging Face Daily Papers: 13 papers, 7 in the top 3 of their day, 775 upvotes.
- V-GameGym: Visual Game Generation for Code Large Language Models 9 upvotes, #17 of 2025-09-26
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe 46 upvotes, #4 of 2025-09-24
- FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models 114 upvotes, #1 of 2025-05-29
- TTRL: Test-Time Reinforcement Learning 96 upvotes, #2 of 2025-04-23
- Learning to Reason under Off-Policy Guidance 77 upvotes, #1 of 2025-04-22
- A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond 39 upvotes, #3 of 2025-03-31
- UltraIF: Advancing Instruction Following from the Wild 20 upvotes, #9 of 2025-02-07
- Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04
- Free Process Rewards without Process Labels 26 upvotes, #4 of 2024-12-04
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies 14 upvotes, #5 of 2024-04-10
- Advancing LLM Reasoning Generalists with Preference Trees 36 upvotes, #2 of 2024-04-03
- RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback 11 upvotes, #12 of 2023-12-05
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.