MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

weimin wang, Jiawei Liu, Zhijie Lin, Yan Jiangqiao, Shuo Chen, Chetwin Low, Tuyen Hoang, Jie Wu, Jun Hao Liew, PeRFlow, Zhou, Jiashi Feng

MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation: 49 upvotes on Hugging Face Daily Papers, #1 of 8 papers on 2024-01-10. Day-by-day upvote history.

The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator, reference image embedding module and frame interpolation module into an end-to-end video generation pipeline. Benefiting from these architecture designs, MagicVideo-V2 can generate an aesthetically pleasing, high-resolution video with remarkable fidelity and smoothness. It demonstrates superior performance over leading Text-to-Video systems such as Runway, Pika 1.0, Morph, Moon Valley and Stable Video Diffusion model via user evaluation at large scale.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.