HAODONG DUAN
HAODONG DUAN on Hugging Face Daily Papers: 38 papers, 19 in the top 3 of their day, 2,213 upvotes.
- SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence 34 upvotes, #7 of 2026-01-07
- MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization 103 upvotes, #2 of 2025-10-10
- SPARK: Synergistic Policy And Reward Co-Evolving Framework 16 upvotes, #21 of 2025-09-29
- A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers 135 upvotes, #2 of 2025-09-01
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 170 upvotes, #1 of 2025-08-26
- Intern-S1: A Scientific Multimodal Foundation Model 242 upvotes, #1 of 2025-08-22
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents 28 upvotes, #3 of 2025-07-28
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence 3 upvotes, #56 of 2025-05-30
- Visual Agentic Reinforcement Fine-Tuning 31 upvotes, #6 of 2025-05-21
- MM-IFEngine: Towards Multimodal Instruction Following 31 upvotes, #6 of 2025-04-11
- Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing 67 upvotes, #3 of 2025-04-04
- LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? 31 upvotes, #6 of 2025-03-27
- Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM 41 upvotes, #4 of 2025-03-19
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
- Visual-RFT: Visual Reinforcement Fine-Tuning 62 upvotes, #2 of 2025-03-04
- OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference 67 upvotes, #1 of 2025-02-26
- VideoRoPE: What Makes for Good Video Rotary Position Embedding? 60 upvotes, #3 of 2025-02-10
- Redundancy Principles for MLLMs Benchmarks 25 upvotes, #4 of 2025-01-27
- InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 39 upvotes, #6 of 2025-01-22
- Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement 14 upvotes, #13 of 2025-01-22
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? 36 upvotes, #5 of 2025-01-13
- BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 34 upvotes, #3 of 2025-01-07
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
- MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
- MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models 34 upvotes, #1 of 2024-10-24
- CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution 55 upvotes, #3 of 2024-10-22
- ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs 13 upvotes, #8 of 2024-10-17
- GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI 82 upvotes, #2 of 2024-08-09
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models 11 upvotes, #6 of 2024-07-17
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
- MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning 18 upvotes, #6 of 2024-06-26
- MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding 27 upvotes, #7 of 2024-06-21
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions 61 upvotes, #1 of 2024-06-07
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
- InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
- JourneyDB: A Benchmark for Generative Image Understanding 20 upvotes, #5 of 2023-07-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.