Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction

Biao Gong, Cheng Zou, zheng, Hu Yu, chenjingdong , jianxinsun, Junbo Zhao, Jun Zhou, Kaixiang Ji, Lixiang Ru, wang, qingpei.gqp, Rui Liu, Willie Chai, xinyu xiao, Ziyuan

Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction: 15 upvotes on Hugging Face Daily Papers, #14 of 22 papers on 2025-05-06. Day-by-day upvote history.

We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides an open-source implementation of the integrated MetaQueries and M2-omni framework, while introducing the novel multi-scale learnable tokens and multi-scale representation alignment strategy. By leveraging a fixed MLLM and a learnable diffusion model, Ming-Lite-Uni enables native multimodal AR models to perform both text-to-image generation and instruction based image editing tasks, expanding their capabilities beyond pure visual understanding. Our experimental results demonstrate the strong performance of Ming-Lite-Uni and illustrate the impressive fluid nature of its interactive process. All code and model weights are open-sourced to foster further exploration within the community. Notably, this work aligns with concurrent multimodal AI milestones - such as ChatGPT-4o with native image generation updated in March 25, 2025 - underscoring the broader significance of unified models like Ming-Lite-Uni on the path toward AGI. Ming-Lite-Uni is in alpha stage and will soon be further refined.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.