Zilong Huang

Zilong Huang on Hugging Face Daily Papers: 14 papers, 7 in the top 3 of their day, 733 upvotes.

  1. Let ViT Speak: Generative Language-Image Pre-training 32 upvotes, #3 of 2026-05-04
  2. Mixture-of-Depths Attention 77 upvotes, #7 of 2026-03-17
  3. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation 13 upvotes, #15 of 2026-03-13
  4. Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 35 upvotes, #9 of 2025-10-22
  5. Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology 44 upvotes, #3 of 2025-07-11
  6. Seed1.5-VL Technical Report 136 upvotes, #1 of 2025-05-13
  7. The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer 15 upvotes, #10 of 2025-04-16
  8. Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding 28 upvotes, #6 of 2025-04-16
  9. GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation 47 upvotes, #2 of 2025-04-14
  10. Video Depth Anything: Consistent Depth Estimation for Super-Long Videos 21 upvotes, #11 of 2025-01-22
  11. Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos 40 upvotes, #4 of 2025-01-08
  12. Depth Anything V2 83 upvotes, #1 of 2024-06-14
  13. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data 64 upvotes, #1 of 2024-01-22
  14. BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs 29 upvotes, #3 of 2023-07-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.