Harshit Sikchi

Harshit Sikchi on Hugging Face Daily Papers: 3 papers, 1 in the top 3 of their day, 41 upvotes.

  1. RL Zero: Zero-Shot Language to Behaviors without any Supervision 4 upvotes, #16 of 2024-12-09
  2. Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms 9 upvotes, #8 of 2024-06-06
  3. Contrastive Prefence Learning: Learning from Human Feedback without RL 25 upvotes, #1 of 2023-10-23

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.