Yi-Fan Zhang

Yi-Fan Zhang on Hugging Face Daily Papers: 21 papers, 6 in the top 3 of their day, 1,071 upvotes.

  1. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 232 upvotes, #1 of 2026-04-08
  2. Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? 36 upvotes, #5 of 2026-04-06
  3. Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis 46 upvotes, #10 of 2026-04-01
  4. Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch 45 upvotes, #7 of 2025-12-03
  5. VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting 41 upvotes, #5 of 2025-10-28
  6. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing 49 upvotes, #4 of 2025-09-30
  7. RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
  8. BaseReward: A Strong Baseline for Multimodal Reward Model 21 upvotes, #5 of 2025-09-22
  9. Kwai Keye-VL 1.5 Technical Report 33 upvotes, #11 of 2025-09-03
  10. Thyme: Think Beyond Images 76 upvotes, #3 of 2025-08-18
  11. Kwai Keye-VL Technical Report 118 upvotes, #1 of 2025-07-03
  12. MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios 38 upvotes, #12 of 2025-05-28
  13. R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning 22 upvotes, #9 of 2025-05-06
  14. MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models 13 upvotes, #9 of 2025-04-07
  15. Aligning Multimodal LLM with Human Preference: A Survey 21 upvotes, #8 of 2025-03-19
  16. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment 30 upvotes, #6 of 2025-02-17
  17. VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction 36 upvotes, #2 of 2025-01-06
  18. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
  19. MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? 25 upvotes, #2 of 2024-08-26
  20. VITA: Towards Open-Source Interactive Omni Multimodal LLM 40 upvotes, #1 of 2024-08-12
  21. Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models 10 upvotes, #16 of 2024-06-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.