Xin Li
Xin Li on Hugging Face Daily Papers: 9 papers, 5 in the top 3 of their day, 402 upvotes.
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 75 upvotes, #3 of 2025-01-23
- VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM 40 upvotes, #4 of 2025-01-03
- 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining 91 upvotes, #1 of 2025-01-03
- Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective 6 upvotes, #15 of 2024-10-17
- The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio 29 upvotes, #3 of 2024-10-17
- SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages 52 upvotes, #3 of 2024-07-30
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 28 upvotes, #8 of 2024-06-13
- CLEX: Continuous Length Extrapolation for Large Language Models 10 upvotes, #10 of 2023-10-26
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding 20 upvotes, #2 of 2023-06-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.