Xiangyu Zeng
Xiangyu Zeng on Hugging Face Daily Papers: 10 papers, 4 in the top 3 of their day, 604 upvotes.
- OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction 221 upvotes, #1 of 2026-10-02
- TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs 165 upvotes, #2 of 2026-07-21
- VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance 9 upvotes, #20 of 2026-07-17
- VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding 168 upvotes, #2 of 2026-07-17
- DOPD: Dual On-policy Distillation 103 upvotes, #2 of 2026-07-01
- RIVER: A Real-Time Interaction Benchmark for Video LLMs 5 upvotes, #14 of 2026-03-05
- Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale 2 upvotes, #70 of 2025-09-30
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning 9 upvotes, #11 of 2025-04-10
- Make Your Training Flexible: Towards Deployment-Efficient Video Models 5 upvotes, #38 of 2025-03-21
- Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment 18 upvotes, #4 of 2024-12-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.