HuangMeow

HuangMeow on Hugging Face Daily Papers: 2 papers, 2 in the top 3 of their day, 73 upvotes.

  1. Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning 204 upvotes, #1 of 2026-03-12
  2. VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training 215 upvotes, #2 of 2026-02-23

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.