Yanwei Li

Yanwei Li on Hugging Face Daily Papers: 11 papers, 5 in the top 3 of their day, 520 upvotes.

  1. Visual Spatial Tuning 46 upvotes, #2 of 2025-11-10
  2. Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 35 upvotes, #9 of 2025-10-22
  3. How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective 28 upvotes, #6 of 2025-09-24
  4. Seed1.5-VL Technical Report 136 upvotes, #1 of 2025-05-13
  5. Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding 28 upvotes, #6 of 2025-04-16
  6. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency 27 upvotes, #9 of 2025-02-14
  7. Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition 43 upvotes, #4 of 2024-12-13
  8. LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
  9. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 13 upvotes, #3 of 2024-06-03
  10. Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models 33 upvotes, #2 of 2024-03-28
  11. GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction 5 upvotes, #4 of 2023-05-31

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.