Bin Wang
Bin Wang on Hugging Face Daily Papers: 15 papers, 5 in the top 3 of their day, 867 upvotes.
- MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale 117 upvotes, #3 of 2026-04-07
- Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale 125 upvotes, #1 of 2026-03-27
- MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding 131 upvotes, #1 of 2026-03-25
- InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery 68 upvotes, #7 of 2026-02-10
- TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition 9 upvotes, #23 of 2025-12-03
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing 100 upvotes, #4 of 2025-09-29
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations 19 upvotes, #5 of 2024-12-10
- OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation 18 upvotes, #7 of 2024-12-04
- Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction 28 upvotes, #5 of 2024-10-29
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception 20 upvotes, #5 of 2024-10-17
- MinerU: An Open-Source Solution for Precise Document Content Extraction 22 upvotes, #4 of 2024-09-30
- CDM: A Reliable Metric for Fair and Accurate Formula Recognition Evaluation 18 upvotes, #7 of 2024-09-06
- InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
- Parrot Captions Teach CLIP to Spot Text 11 upvotes, #9 of 2023-12-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.