Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

shuhao gu, Jialing Zhang, Siyuan Zhou, Kevin Yu, Xingzhaohu, ldwang, Caozhou, Jintao Jia, Zhuoyi Zhang, Grace Wang, 胡振崇, bowenzhang, lijijie, Dong Liang, Yingli Zhao, Yulong Ao, Yaoqi Liu, Fangxiang Feng, Guang Liu

Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data: 18 upvotes on Hugging Face Daily Papers, #6 of 16 papers on 2024-10-28. Day-by-day upvote history.

Vision-Language Models (VLMs) have recently made significant progress, but the limited scale and quality of open-source instruction data hinder their performance compared to closed-source models. In this work, we address this limitation by introducing Infinity-MM, a large-scale multimodal instruction dataset with 40 million samples, enhanced through rigorous quality filtering and deduplication. We also propose a synthetic instruction generation method based on open-source VLMs, using detailed image annotations and diverse question generation. Using this data, we trained a 2-billion-parameter VLM, Aquila-VL-2B, achieving state-of-the-art (SOTA) performance for models of similar scale. This demonstrates that expanding instruction data and generating synthetic data can significantly improve the performance of open-source models.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.