Zhiding Yu
Zhiding Yu on Hugging Face Daily Papers: 17 papers, 5 in the top 3 of their day, 670 upvotes.
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding 135 upvotes, #1 of 2026-05-27
- Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence 17 upvotes, #15 of 2026-05-01
- ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents 14 upvotes, #19 of 2026-03-20
- Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline 4 upvotes, #18 of 2026-03-06
- PhyCritic: Multimodal Critic Models for Physical AI 51 upvotes, #4 of 2026-02-12
- OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation 18 upvotes, #12 of 2026-01-23
- Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning 50 upvotes, #6 of 2026-01-15
- NVIDIA Nemotron Nano V2 VL 25 upvotes, #5 of 2025-11-07
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models 65 upvotes, #2 of 2025-04-22
- Slow-Fast Architecture for Video Multi-Modal Large Language Models 8 upvotes, #12 of 2025-04-07
- Token-Efficient Long Video Understanding for Multimodal LLMs 79 upvotes, #2 of 2025-03-07
- QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation 10 upvotes, #12 of 2025-02-10
- StreamChat: Chatting with Streaming Video 17 upvotes, #9 of 2024-12-12
- Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders 76 upvotes, #1 of 2024-08-29
- LITA: Language Instructed Temporal-Localization Assistant 16 upvotes, #2 of 2024-03-29
- T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching 11 upvotes, #9 of 2024-02-23
- FocalFormer3D : Focusing on Hard Instance for 3D Object Detection 10 upvotes, #4 of 2023-08-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.