Harshit Sikchi
Harshit Sikchi on Hugging Face Daily Papers: 3 papers, 1 in the top 3 of their day, 41 upvotes.
- RL Zero: Zero-Shot Language to Behaviors without any Supervision 4 upvotes, #16 of 2024-12-09
- Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms 9 upvotes, #8 of 2024-06-06
- Contrastive Prefence Learning: Learning from Human Feedback without RL 25 upvotes, #1 of 2023-10-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.