Bo Li

Bo Li on Hugging Face Daily Papers: 17 papers, 8 in the top 3 of their day, 664 upvotes.

  1. Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
  2. LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training 38 upvotes, #8 of 2025-09-29
  3. MMSearch-R1: Incentivizing LMMs to Search 56 upvotes, #3 of 2025-06-25
  4. EgoLife: Towards Egocentric Life Assistant 35 upvotes, #4 of 2025-03-07
  5. Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 22 upvotes, #5 of 2025-01-24
  6. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
  7. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
  8. Large Multi-modal Models Can Interpret Features in Large Multi-modal Models 14 upvotes, #7 of 2024-11-25
  9. MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures 72 upvotes, #2 of 2024-10-18
  10. Video Instruction Tuning With Synthetic Data 33 upvotes, #5 of 2024-10-04
  11. LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
  12. LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models 28 upvotes, #5 of 2024-07-18
  13. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 34 upvotes, #3 of 2024-07-11
  14. Long Context Transfer from Language to Vision 30 upvotes, #4 of 2024-06-25
  15. OtterHD: A High-Resolution Multi-modality Model 34 upvotes, #1 of 2023-11-08
  16. Octopus: Embodied Vision-Language Programmer from Environmental Feedback 37 upvotes, #2 of 2023-10-13
  17. MIMIC-IT: Multi-Modal In-Context Instruction Tuning 12 upvotes, #2 of 2023-06-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.