Yuhao Dong

Yuhao Dong on Hugging Face Daily Papers: 30 papers, 5 in the top 3 of their day, 1,883 upvotes.

  1. PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models 18 upvotes, #11 of 2026-07-29
  2. Kimi K3: Open Frontier Intelligence 458 upvotes, #1 of 2026-07-28
  3. VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding 168 upvotes, #2 of 2026-07-17
  4. ViQ: Text-Aligned Visual Quantized Representations at Any Resolution 38 upvotes, #6 of 2026-06-26
  5. S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence 39 upvotes, #6 of 2026-06-19
  6. From Pixels to Words -- Towards Native One-Vision Models at Scale 72 upvotes, #4 of 2026-05-28
  7. LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
  8. Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos 22 upvotes, #12 of 2026-05-20
  9. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 232 upvotes, #1 of 2026-04-08
  10. FileGram: Grounding Agent Personalization in File-System Behavioral Traces 40 upvotes, #9 of 2026-04-07
  11. PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning 18 upvotes, #12 of 2026-04-02
  12. Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models 12 upvotes, #22 of 2026-03-24
  13. VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining 21 upvotes, #14 of 2026-03-20
  14. Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition 28 upvotes, #15 of 2026-02-10
  15. Kimi K2.5: Visual Agentic Intelligence 219 upvotes, #2 of 2026-02-03
  16. Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark 9 upvotes, #21 of 2025-10-16
  17. RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
  18. Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding 20 upvotes, #13 of 2025-07-22
  19. High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning 11 upvotes, #17 of 2025-07-09
  20. ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models 22 upvotes, #5 of 2025-06-30
  21. Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning 42 upvotes, #7 of 2025-06-17
  22. EgoLife: Towards Egocentric Life Assistant 35 upvotes, #4 of 2025-03-07
  23. Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment 21 upvotes, #7 of 2025-02-07
  24. Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives 23 upvotes, #4 of 2025-01-10
  25. Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 19 upvotes, #7 of 2024-11-22
  26. 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion 17 upvotes, #8 of 2024-09-20
  27. Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 22 upvotes, #5 of 2024-09-20
  28. Coarse Correspondence Elicit 3D Spacetime Understanding in Multimodal Language Model 20 upvotes, #4 of 2024-08-02
  29. Efficient Inference of Vision Instruction-Following Models with Elastic Cache 15 upvotes, #8 of 2024-07-26
  30. Octopus: Embodied Vision-Language Programmer from Environmental Feedback 37 upvotes, #2 of 2023-10-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.