HaochenWang
HaochenWang on Hugging Face Daily Papers: 7 papers, 2 in the top 3 of their day, 213 upvotes.
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs 17 upvotes, #10 of 2025-11-11
- PairUni: Pairwise Training for Unified Multimodal Language Models 13 upvotes, #16 of 2025-10-30
- Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence 52 upvotes, #3 of 2025-10-24
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 35 upvotes, #9 of 2025-10-22
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology 44 upvotes, #3 of 2025-07-11
- VGR: Visual Grounded Reasoning 19 upvotes, #15 of 2025-06-17
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer 15 upvotes, #10 of 2025-04-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.