Baolin Peng

Baolin Peng on Hugging Face Daily Papers: 8 papers, 3 in the top 3 of their day, 332 upvotes.

  1. Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math 37 upvotes, #3 of 2025-05-01
  2. Reinforcement Learning for Reasoning in Large Language Models with One Training Example 88 upvotes, #1 of 2025-04-30
  3. Magma: A Foundation Model for Multimodal AI Agents 46 upvotes, #5 of 2025-02-19
  4. Improving Autonomous AI Agents with Reflective Tree Search and Self-Learning 8 upvotes, #17 of 2024-10-04
  5. LiteSearch: Efficacious Tree Search for LLM 34 upvotes, #4 of 2024-07-02
  6. Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing 44 upvotes, #1 of 2024-04-19
  7. Teaching Language Models to Self-Improve through Interactive Demonstrations 12 upvotes, #8 of 2023-10-23
  8. Stabilizing RLHF through Advantage Model and Selective Rehearsal 10 upvotes, #6 of 2023-09-20

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.