VideoChat: Chat-Centric Video Understanding

Kunchang Li, yinanhe, Yi Wang, Yizhuo Li, wangwenhai, Ping Luo, Yali Wang, Limin Wang, Yu Qiao

VideoChat: Chat-Centric Video Understanding: 3 upvotes on Hugging Face Daily Papers, #3 of 12 papers on 2023-05-11. Day-by-day upvote history.

In this study, we initiate an exploration into video understanding by introducing VideoChat, an end-to-end chat-centric video understanding system. It integrates video foundation models and large language models via a learnable neural interface, excelling in spatiotemporal reasoning, event localization, and causal relationship inference. To instructively tune this system, we propose a video-centric instruction dataset, composed of thousands of videos matched with detailed descriptions and conversations. This dataset emphasizes spatiotemporal reasoning and causal relationships, providing a valuable asset for training chat-centric video understanding systems. Preliminary qualitative experiments reveal our system's potential across a broad spectrum of video applications and set the standard for future research. Access our code and data at https://github.com/OpenGVLab/Ask-Anything

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.