Kunchang Li
Kunchang Li on Hugging Face Daily Papers: 14 papers, 5 in the top 3 of their day, 496 upvotes.
- Scaling Properties of Text Conditioning in Visual Generation 38 upvotes, #7 of 2026-08-03
- Emerging Properties in Unified Multimodal Pretraining 124 upvotes, #1 of 2025-05-21
- Seed1.5-VL Technical Report 136 upvotes, #1 of 2025-05-13
- Make Your Training Flexible: Towards Deployment-Efficient Video Models 5 upvotes, #38 of 2025-03-21
- Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment 18 upvotes, #4 of 2024-12-30
- Causal Diffusion Transformers for Generative Modeling 23 upvotes, #7 of 2024-12-17
- Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel 5 upvotes, #17 of 2024-12-12
- TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration 3 upvotes, #31 of 2024-10-18
- InternVideo2: Scaling Video Foundation Models for Multimodal Video Understanding 14 upvotes, #3 of 2024-03-25
- Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding 10 upvotes, #9 of 2024-03-15
- VideoMamba: State Space Model for Efficient Video Understanding 20 upvotes, #5 of 2024-03-12
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation 26 upvotes, #3 of 2023-07-14
- VideoChat: Chat-Centric Video Understanding 3 upvotes, #3 of 2023-05-11
- InternChat: Solving Vision-Centric Tasks by Interacting with Chatbots Beyond Language 5 upvotes, #5 of 2023-05-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.