minghao
minghao on Hugging Face Daily Papers: 35 papers, 8 in the top 3 of their day, 1,485 upvotes.
- SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models 3 upvotes, #33 of 2026-08-06
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents 37 upvotes, #9 of 2026-06-23
- MMAE: A Massive Multitask Audio Editing Benchmark 44 upvotes, #6 of 2026-06-08
- TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation 14 upvotes, #25 of 2026-06-02
- WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models 22 upvotes, #10 of 2026-04-21
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding 23 upvotes, #15 of 2026-03-31
- Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization 22 upvotes, #8 of 2026-02-27
- EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies 9 upvotes, #25 of 2026-02-12
- SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature 7 upvotes, #8 of 2026-01-20
- The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning 48 upvotes, #4 of 2026-01-12
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents 42 upvotes, #7 of 2025-12-16
- A^2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning 22 upvotes, #11 of 2025-10-20
- COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes 13 upvotes, #18 of 2025-10-17
- Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures 10 upvotes, #23 of 2025-10-17
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems 26 upvotes, #12 of 2025-10-14
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs 45 upvotes, #4 of 2025-10-14
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing 15 upvotes, #11 of 2025-10-02
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video? 3 upvotes, #64 of 2025-09-30
- Towards Personalized Deep Research: Benchmarks and Evaluations 27 upvotes, #13 of 2025-09-30
- VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
- Reverse-Engineered Reasoning for Open-Ended Generation 142 upvotes, #1 of 2025-09-09
- Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL 114 upvotes, #1 of 2025-08-20
- VeriGUI: Verifiable Long-Chain GUI Dataset 137 upvotes, #2 of 2025-08-07
- CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization 38 upvotes, #6 of 2025-07-09
- OAgents: An Empirical Study of Building Effective Agents 34 upvotes, #4 of 2025-06-24
- Scaling Test-time Compute for LLM Agents 57 upvotes, #2 of 2025-06-18
- FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models 27 upvotes, #7 of 2025-05-06
- COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values 41 upvotes, #5 of 2025-04-09
- YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 92 upvotes, #3 of 2025-02-21
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations 19 upvotes, #5 of 2024-12-10
- AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions 36 upvotes, #2 of 2024-10-30
- A Comparative Study on Reasoning Patterns of OpenAI's o1 Model 15 upvotes, #15 of 2024-10-18
- PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents 20 upvotes, #10 of 2024-06-21
- MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series 41 upvotes, #1 of 2024-05-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.