Boqiang Zhang
Boqiang Zhang on Hugging Face Daily Papers: 5 papers, 3 in the top 3 of their day, 378 upvotes.
- Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders 105 upvotes, #1 of 2026-03-09
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models 19 upvotes, #14 of 2025-12-19
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources 98 upvotes, #2 of 2025-09-26
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 75 upvotes, #3 of 2025-01-23
- VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM 40 upvotes, #4 of 2025-01-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.