Sainbayar Sukhbaatar

Sainbayar Sukhbaatar on Hugging Face Daily Papers: 19 papers, 6 in the top 3 of their day, 684 upvotes.

  1. StepWiser: Stepwise Generative Judges for Wiser Reasoning 19 upvotes, #10 of 2025-08-28
  2. Self-Challenging Language Model Agents 9 upvotes, #27 of 2025-06-04
  3. Multi-Token Attention 39 upvotes, #3 of 2025-04-02
  4. SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks 8 upvotes, #16 of 2025-03-20
  5. Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback 14 upvotes, #7 of 2025-01-24
  6. Training Large Language Models to Reason in a Continuous Latent Space 54 upvotes, #3 of 2024-12-10
  7. Adaptive Decoding via Latent Preference Optimization 10 upvotes, #9 of 2024-11-19
  8. Self-Consistency Preference Optimization 12 upvotes, #4 of 2024-11-07
  9. Thinking LLMs: General Instruction Following with Thought Generation 7 upvotes, #15 of 2024-10-15
  10. Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge 19 upvotes, #14 of 2024-07-30
  11. Iterative Reasoning Preference Optimization 35 upvotes, #5 of 2024-05-01
  12. Reverse Training to Nurse the Reversal Curse 10 upvotes, #13 of 2024-03-21
  13. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 29 upvotes, #3 of 2024-03-13
  14. Teaching Large Language Models to Reason with Reinforcement Learning 32 upvotes, #3 of 2024-03-08
  15. Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping 47 upvotes, #2 of 2024-02-23
  16. Self-Rewarding Language Models 156 upvotes, #1 of 2024-01-19
  17. System 2 Attention (is something you might need too) 43 upvotes, #4 of 2023-11-21
  18. Improving Open Language Models by Learning from Organic Interactions 3 upvotes, #12 of 2023-06-09
  19. Large Language Model Programs 2 upvotes, #8 of 2023-05-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.