Descriptive Caption Enhancement with Visual Specialists for Multimodal Perception

yanpeng_sun, Jing Hao, Ke Zhu, Jiang-Jiang Liu, yuxiang zhao, Xiaofan Li, Gang Zhang, Zechao Li, Jingdong Wang

Descriptive Caption Enhancement with Visual Specialists for Multimodal Perception: 7 upvotes on Hugging Face Daily Papers, #11 of 17 papers on 2024-12-20. Day-by-day upvote history.

Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods either distill the caption from the LMM models or construct the captions from the internet images or by human. We propose to leverage off-the-shelf visual specialists, which were trained from annotated images initially not for image captioning, for enhancing the image caption. Our approach, named DCE, explores object low-level and fine-grained attributes (e.g., depth, emotion and fine-grained categories) and object relations (e.g., relative location and human-object-interaction (HOI)), and combine the attributes into the descriptive caption. Experiments demonstrate that such visual specialists are able to improve the performance for visual understanding tasks as well as reasoning that benefits from more accurate visual understanding. We will release the source code and the pipeline so that other visual specialists are easily combined into the pipeline. The complete source code of DCE pipeline and datasets will be available at https://github.com/syp2ysy/DCE.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.