Xiangyu Zeng

Xiangyu Zeng on Hugging Face Daily Papers: 10 papers, 4 in the top 3 of their day, 604 upvotes.

  1. OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction 221 upvotes, #1 of 2026-10-02
  2. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs 165 upvotes, #2 of 2026-07-21
  3. VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance 9 upvotes, #20 of 2026-07-17
  4. VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding 168 upvotes, #2 of 2026-07-17
  5. DOPD: Dual On-policy Distillation 103 upvotes, #2 of 2026-07-01
  6. RIVER: A Real-Time Interaction Benchmark for Video LLMs 5 upvotes, #14 of 2026-03-05
  7. Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale 2 upvotes, #70 of 2025-09-30
  8. VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning 9 upvotes, #11 of 2025-04-10
  9. Make Your Training Flexible: Towards Deployment-Efficient Video Models 5 upvotes, #38 of 2025-03-21
  10. Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment 18 upvotes, #4 of 2024-12-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.