Minjae Oh

Minjae Oh on Hugging Face Daily Papers: 4 papers, 0 in the top 3 of their day, 60 upvotes.

  1. The Low-Rank Structure of VLA Reinforcement Learning 16 upvotes, #33 of 2026-10-01
  2. SHAPE of Chain-of-Thought in Math Reasoning 31 upvotes, #9 of 2026-09-01
  3. KL for a KL: On-Policy Distillation with Control Variate Baseline 19 upvotes, #14 of 2026-05-14
  4. Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States 16 upvotes, #20 of 2026-05-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.