Aviral Kumar
Aviral Kumar on Hugging Face Daily Papers: 8 papers, 3 in the top 3 of their day, 341 upvotes.
- WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks 15 upvotes, #13 of 2026-01-07
- Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning 36 upvotes, #4 of 2025-03-12
- Training Language Models to Self-Correct via Reinforcement Learning 117 upvotes, #1 of 2024-09-20
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters 15 upvotes, #6 of 2024-08-07
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning 17 upvotes, #11 of 2024-06-21
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RL 10 upvotes, #6 of 2024-03-07
- Robotic Offline RL from Internet Videos via Value-Function Pre-Training 9 upvotes, #3 of 2023-09-25
- Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions 24 upvotes, #3 of 2023-09-20
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.