Xin Zhou
Xin Zhou on Hugging Face Daily Papers: 7 papers, 4 in the top 3 of their day, 647 upvotes.
- Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving 378 upvotes, #2 of 2026-09-02
- SimWAM: A Simple World Action Model for End-to-End Autonomous Driving 105 upvotes, #1 of 2026-08-10
- Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments 138 upvotes, #2 of 2026-05-29
- HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation 71 upvotes, #5 of 2026-05-07
- When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models 114 upvotes, #5 of 2026-04-10
- Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models 152 upvotes, #2 of 2026-03-30
- MINIMA: Modality Invariant Image Matching 3 upvotes, #10 of 2025-01-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.