Ganqu Cui

Ganqu Cui on Hugging Face Daily Papers: 13 papers, 7 in the top 3 of their day, 775 upvotes.

  1. V-GameGym: Visual Game Generation for Code Large Language Models 9 upvotes, #17 of 2025-09-26
  2. MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe 46 upvotes, #4 of 2025-09-24
  3. FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
  4. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models 114 upvotes, #1 of 2025-05-29
  5. TTRL: Test-Time Reinforcement Learning 96 upvotes, #2 of 2025-04-23
  6. Learning to Reason under Off-Policy Guidance 77 upvotes, #1 of 2025-04-22
  7. A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond 39 upvotes, #3 of 2025-03-31
  8. UltraIF: Advancing Instruction Following from the Wild 20 upvotes, #9 of 2025-02-07
  9. Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04
  10. Free Process Rewards without Process Labels 26 upvotes, #4 of 2024-12-04
  11. MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies 14 upvotes, #5 of 2024-04-10
  12. Advancing LLM Reasoning Generalists with Preference Trees 36 upvotes, #2 of 2024-04-03
  13. RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback 11 upvotes, #12 of 2023-12-05

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.