Bo Li
Bo Li on Hugging Face Daily Papers: 17 papers, 8 in the top 3 of their day, 664 upvotes.
- Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training 38 upvotes, #8 of 2025-09-29
- MMSearch-R1: Incentivizing LMMs to Search 56 upvotes, #3 of 2025-06-25
- EgoLife: Towards Egocentric Life Assistant 35 upvotes, #4 of 2025-03-07
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 22 upvotes, #5 of 2025-01-24
- MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 42 upvotes, #3 of 2024-12-09
- MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
- Large Multi-modal Models Can Interpret Features in Large Multi-modal Models 14 upvotes, #7 of 2024-11-25
- MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures 72 upvotes, #2 of 2024-10-18
- Video Instruction Tuning With Synthetic Data 33 upvotes, #5 of 2024-10-04
- LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models 28 upvotes, #5 of 2024-07-18
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 34 upvotes, #3 of 2024-07-11
- Long Context Transfer from Language to Vision 30 upvotes, #4 of 2024-06-25
- OtterHD: A High-Resolution Multi-modality Model 34 upvotes, #1 of 2023-11-08
- Octopus: Embodied Vision-Language Programmer from Environmental Feedback 37 upvotes, #2 of 2023-10-13
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning 12 upvotes, #2 of 2023-06-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.