Xiang Yue
Xiang Yue on Hugging Face Daily Papers: 30 papers, 10 in the top 3 of their day, 1,168 upvotes.
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models 31 upvotes, #5 of 2025-12-09
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution 44 upvotes, #6 of 2025-10-30
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents 24 upvotes, #8 of 2025-10-29
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning 62 upvotes, #2 of 2025-07-02
- VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
- The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think 24 upvotes, #6 of 2025-05-16
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge 10 upvotes, #20 of 2025-04-16
- ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 20 upvotes, #13 of 2025-03-14
- Small Models Struggle to Learn from Strong Reasoners 27 upvotes, #7 of 2025-02-20
- Demystifying Long Chain-of-Thought Reasoning in LLMs 49 upvotes, #2 of 2025-02-06
- Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate 50 upvotes, #1 of 2025-01-30
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 22 upvotes, #5 of 2025-01-24
- MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
- Evaluating Language Models as Synthetic Data Generators 39 upvotes, #5 of 2024-12-06
- Teach Multimodal LLMs to Comprehend Electrocardiographic Images 22 upvotes, #3 of 2024-10-28
- JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation 12 upvotes, #7 of 2024-10-23
- Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages 41 upvotes, #8 of 2024-10-22
- MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures 72 upvotes, #2 of 2024-10-18
- Harnessing Webpage UIs for Text-Rich Visual Understanding 28 upvotes, #7 of 2024-10-18
- MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks 34 upvotes, #5 of 2024-10-15
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-09-05
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark 35 upvotes, #1 of 2024-06-04
- Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization 30 upvotes, #3 of 2024-05-27
- MuPT: A Generative Symbolic Music Pretrained Transformer 14 upvotes, #5 of 2024-04-10
- CodeEditorBench: Evaluating Code Editing Capability of Large Language Models 14 upvotes, #7 of 2024-04-05
- Long-context LLMs Struggle with Long In-context Learning 28 upvotes, #3 of 2024-04-03
- StructLM: Towards Building Generalist Models for Structured Knowledge Grounding 27 upvotes, #7 of 2024-02-27
- OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement 84 upvotes, #1 of 2024-02-23
- Data Engineering for Scaling Language Models to 128K Context 24 upvotes, #7 of 2024-02-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.