Zilong Huang
Zilong Huang on Hugging Face Daily Papers: 14 papers, 7 in the top 3 of their day, 733 upvotes.
- Let ViT Speak: Generative Language-Image Pre-training 32 upvotes, #3 of 2026-05-04
- Mixture-of-Depths Attention 77 upvotes, #7 of 2026-03-17
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation 13 upvotes, #15 of 2026-03-13
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 35 upvotes, #9 of 2025-10-22
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology 44 upvotes, #3 of 2025-07-11
- Seed1.5-VL Technical Report 136 upvotes, #1 of 2025-05-13
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer 15 upvotes, #10 of 2025-04-16
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding 28 upvotes, #6 of 2025-04-16
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation 47 upvotes, #2 of 2025-04-14
- Video Depth Anything: Consistent Depth Estimation for Super-Long Videos 21 upvotes, #11 of 2025-01-22
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos 40 upvotes, #4 of 2025-01-08
- Depth Anything V2 83 upvotes, #1 of 2024-06-14
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data 64 upvotes, #1 of 2024-01-22
- BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs 29 upvotes, #3 of 2023-07-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.