Tianqi Liu

Tianqi Liu on Hugging Face Daily Papers: 9 papers, 2 in the top 3 of their day, 173 upvotes.

  1. RRM: Robust Reward Model Training Mitigates Reward Hacking 3 upvotes, #17 of 2024-09-25
  2. Building Math Agents with Multi-Turn Iterative Preference Learning 14 upvotes, #9 of 2024-09-06
  3. PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs 4 upvotes, #12 of 2024-06-06
  4. Offline Regularised Reinforcement Learning for Large Language Models Alignment 11 upvotes, #6 of 2024-05-30
  5. Direct Language Model Alignment from Online AI Feedback 36 upvotes, #4 of 2024-02-08
  6. LiPO: Listwise Preference Optimization through Learning-to-Rank 20 upvotes, #6 of 2024-02-06
  7. Gemini: A Family of Highly Capable Multimodal Models 50 upvotes, #2 of 2023-12-20
  8. Statistical Rejection Sampling Improves Preference Optimization 15 upvotes, #4 of 2023-09-14
  9. SLiC-HF: Sequence Likelihood Calibration with Human Feedback 7 upvotes, #3 of 2023-05-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.