Jason Weston
Jason Weston on Hugging Face Daily Papers: 23 papers, 6 in the top 3 of their day, 846 upvotes.
- Jointly Reinforcing Diversity and Quality in Language Model Generations 25 upvotes, #13 of 2025-09-03
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning 20 upvotes, #8 of 2025-05-16
- Multi-Token Attention 39 upvotes, #3 of 2025-04-02
- SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks 8 upvotes, #16 of 2025-03-20
- Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback 14 upvotes, #7 of 2025-01-24
- Byte Latent Transformer: Patches Scale Better Than Tokens 74 upvotes, #1 of 2024-12-17
- Training Large Language Models to Reason in a Continuous Latent Space 54 upvotes, #3 of 2024-12-10
- Adaptive Decoding via Latent Preference Optimization 10 upvotes, #9 of 2024-11-19
- Thinking LLMs: General Instruction Following with Thought Generation 7 upvotes, #15 of 2024-10-15
- Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources 15 upvotes, #4 of 2024-09-13
- Better Alignment with Instruction Back-and-Forth Translation 13 upvotes, #5 of 2024-08-09
- Self-Taught Evaluators 16 upvotes, #6 of 2024-08-06
- Iterative Reasoning Preference Optimization 35 upvotes, #5 of 2024-05-01
- Reverse Training to Nurse the Reversal Curse 10 upvotes, #13 of 2024-03-21
- Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 29 upvotes, #3 of 2024-03-13
- Self-Rewarding Language Models 156 upvotes, #1 of 2024-01-19
- System 2 Attention (is something you might need too) 43 upvotes, #4 of 2023-11-21
- The ART of LLM Refinement: Ask, Refine, and Trust 11 upvotes, #9 of 2023-11-15
- Branch-Solve-Merge Improves Large Language Model Evaluation and Generation 8 upvotes, #7 of 2023-10-24
- Chain-of-Verification Reduces Hallucination in Large Language Models 39 upvotes, #5 of 2023-09-21
- Self-Alignment with Instruction Backtranslation 43 upvotes, #1 of 2023-08-14
- Leveraging Implicit Feedback from Deployment Data in Dialogue 5 upvotes, #9 of 2023-07-27
- System-Level Natural Language Feedback 11 upvotes, #5 of 2023-06-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.