caohaoyu

caohaoyu on Hugging Face Daily Papers: 6 papers, 1 in the top 3 of their day, 390 upvotes.

  1. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 232 upvotes, #1 of 2026-04-08
  2. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion 48 upvotes, #4 of 2026-03-11
  3. RISE-Video: Can Video Generators Decode Implicit World Rules? 26 upvotes, #9 of 2026-02-06
  4. Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision 41 upvotes, #4 of 2026-01-28
  5. VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting 41 upvotes, #5 of 2025-10-28
  6. VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model 8 upvotes, #12 of 2025-05-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.