Xin Li

Xin Li on Hugging Face Daily Papers: 9 papers, 5 in the top 3 of their day, 402 upvotes.

  1. VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 75 upvotes, #3 of 2025-01-23
  2. VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM 40 upvotes, #4 of 2025-01-03
  3. 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining 91 upvotes, #1 of 2025-01-03
  4. Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective 6 upvotes, #15 of 2024-10-17
  5. The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio 29 upvotes, #3 of 2024-10-17
  6. SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages 52 upvotes, #3 of 2024-07-30
  7. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 28 upvotes, #8 of 2024-06-13
  8. CLEX: Continuous Length Extrapolation for Large Language Models 10 upvotes, #10 of 2023-10-26
  9. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding 20 upvotes, #2 of 2023-06-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.