Xiang Yue

Xiang Yue on Hugging Face Daily Papers: 30 papers, 10 in the top 3 of their day, 1,168 upvotes.

  1. On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models 31 upvotes, #5 of 2025-12-09
  2. The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution 44 upvotes, #6 of 2025-10-30
  3. Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents 24 upvotes, #8 of 2025-10-29
  4. Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning 62 upvotes, #2 of 2025-07-02
  5. VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
  6. The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think 24 upvotes, #6 of 2025-05-16
  7. VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge 10 upvotes, #20 of 2025-04-16
  8. ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
  9. VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 20 upvotes, #13 of 2025-03-14
  10. Small Models Struggle to Learn from Strong Reasoners 27 upvotes, #7 of 2025-02-20
  11. Demystifying Long Chain-of-Thought Reasoning in LLMs 49 upvotes, #2 of 2025-02-06
  12. Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate 50 upvotes, #1 of 2025-01-30
  13. Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 22 upvotes, #5 of 2025-01-24
  14. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
  15. Evaluating Language Models as Synthetic Data Generators 39 upvotes, #5 of 2024-12-06
  16. Teach Multimodal LLMs to Comprehend Electrocardiographic Images 22 upvotes, #3 of 2024-10-28
  17. JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation 12 upvotes, #7 of 2024-10-23
  18. Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages 41 upvotes, #8 of 2024-10-22
  19. MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures 72 upvotes, #2 of 2024-10-18
  20. Harnessing Webpage UIs for Text-Rich Visual Understanding 28 upvotes, #7 of 2024-10-18
  21. MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks 34 upvotes, #5 of 2024-10-15
  22. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-09-05
  23. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark 35 upvotes, #1 of 2024-06-04
  24. Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization 30 upvotes, #3 of 2024-05-27
  25. MuPT: A Generative Symbolic Music Pretrained Transformer 14 upvotes, #5 of 2024-04-10
  26. CodeEditorBench: Evaluating Code Editing Capability of Large Language Models 14 upvotes, #7 of 2024-04-05
  27. Long-context LLMs Struggle with Long In-context Learning 28 upvotes, #3 of 2024-04-03
  28. StructLM: Towards Building Generalist Models for Structured Knowledge Grounding 27 upvotes, #7 of 2024-02-27
  29. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement 84 upvotes, #1 of 2024-02-23
  30. Data Engineering for Scaling Language Models to 128K Context 24 upvotes, #7 of 2024-02-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.