Tianqi Liu
Tianqi Liu on Hugging Face Daily Papers: 9 papers, 2 in the top 3 of their day, 173 upvotes.
- RRM: Robust Reward Model Training Mitigates Reward Hacking 3 upvotes, #17 of 2024-09-25
- Building Math Agents with Multi-Turn Iterative Preference Learning 14 upvotes, #9 of 2024-09-06
- PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs 4 upvotes, #12 of 2024-06-06
- Offline Regularised Reinforcement Learning for Large Language Models Alignment 11 upvotes, #6 of 2024-05-30
- Direct Language Model Alignment from Online AI Feedback 36 upvotes, #4 of 2024-02-08
- LiPO: Listwise Preference Optimization through Learning-to-Rank 20 upvotes, #6 of 2024-02-06
- Gemini: A Family of Highly Capable Multimodal Models 50 upvotes, #2 of 2023-12-20
- Statistical Rejection Sampling Improves Preference Optimization 15 upvotes, #4 of 2023-09-14
- SLiC-HF: Sequence Likelihood Calibration with Human Feedback 7 upvotes, #3 of 2023-05-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.