Wei Shen
Wei Shen on Hugging Face Daily Papers: 6 papers, 2 in the top 3 of their day, 190 upvotes.
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning 54 upvotes, #3 of 2025-05-20
- Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model 18 upvotes, #8 of 2025-04-24
- LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs 19 upvotes, #13 of 2025-04-22
- Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback 43 upvotes, #2 of 2025-03-31
- Policy Filtration in RLHF to Fine-Tune LLM for Code Generation 5 upvotes, #10 of 2024-09-17
- LongRecipe: Recipe for Efficient Long Context Generalization in Large Languge Models 37 upvotes, #4 of 2024-09-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.