Conghui He
Conghui He on Hugging Face Daily Papers: 32 papers, 12 in the top 3 of their day, 1,714 upvotes.
- Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs 181 upvotes, #1 of 2026-01-27
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning 3 upvotes, #23 of 2025-12-09
- ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning 9 upvotes, #17 of 2025-09-26
- From Uniform to Heterogeneous: Tailoring Policy Optimization to Every Token's Nature 2 upvotes, #29 of 2025-09-23
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression 141 upvotes, #2 of 2025-05-27
- CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges 16 upvotes, #4 of 2025-04-29
- FUSION: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding 38 upvotes, #5 of 2025-04-15
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
- GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation 54 upvotes, #4 of 2025-04-04
- Lumina-Image 2.0: A Unified and Efficient Image Generative Framework 18 upvotes, #11 of 2025-03-28
- LEMMA: Learning from Errors for MatheMatical Advancement in LLMs 13 upvotes, #14 of 2025-03-25
- MathFusion: Enhancing Mathematic Problem-solving of LLM through Instruction Fusion 22 upvotes, #17 of 2025-03-21
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer 20 upvotes, #7 of 2025-03-20
- LEGION: Learning to Ground and Explain for Synthetic Image Detection 19 upvotes, #8 of 2025-03-20
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? 36 upvotes, #5 of 2025-01-13
- GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training 2 upvotes, #24 of 2024-12-17
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
- Chimera: Improving Generalist Model with Domain-Specific Experts 9 upvotes, #19 of 2024-12-11
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations 19 upvotes, #5 of 2024-12-10
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
- OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation 18 upvotes, #7 of 2024-12-04
- MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models 34 upvotes, #1 of 2024-10-24
- PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction 42 upvotes, #1 of 2024-10-23
- LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models 53 upvotes, #1 of 2024-10-15
- Multi-Agent Collaborative Data Selection for Efficient LLM Pretraining 17 upvotes, #6 of 2024-10-14
- MinerU: An Open-Source Solution for Precise Document Content Extraction 22 upvotes, #4 of 2024-09-30
- CDM: A Reliable Metric for Fair and Accurate Formula Recognition Evaluation 18 upvotes, #7 of 2024-09-06
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
- InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
- SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models 17 upvotes, #8 of 2024-02-09
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.