Rishabh Agarwal

Rishabh Agarwal on Hugging Face Daily Papers: 3 papers, 1 in the top 3 of their day, 167 upvotes.

  1. Training Language Models to Self-Correct via Reinforcement Learning 117 upvotes, #1 of 2024-09-20
  2. On scalable oversight with weak LLMs judging strong LLMs 11 upvotes, #10 of 2024-07-08
  3. Transformers Can Achieve Length Generalization But Not Robustly 14 upvotes, #4 of 2024-02-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.