沈云航 Yunhang Shen
沈云航 Yunhang Shen on Hugging Face Daily Papers: 11 papers, 5 in the top 3 of their day, 541 upvotes.
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 232 upvotes, #1 of 2026-04-08
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion 48 upvotes, #4 of 2026-03-11
- Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision 41 upvotes, #4 of 2026-01-28
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting 41 upvotes, #5 of 2025-10-28
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model 8 upvotes, #12 of 2025-05-07
- Aligning Multimodal LLM with Human Preference: A Survey 21 upvotes, #8 of 2025-03-19
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction 36 upvotes, #2 of 2025-01-06
- VITA: Towards Open-Source Interactive Omni Multimodal LLM 40 upvotes, #1 of 2024-08-12
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 13 upvotes, #3 of 2024-06-03
- A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise 14 upvotes, #5 of 2023-12-20
- Woodpecker: Hallucination Correction for Multimodal Large Language Models 18 upvotes, #3 of 2023-10-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.