KAIPENG ZHANG

KAIPENG ZHANG on Hugging Face Daily Papers: 16 papers, 6 in the top 3 of their day, 820 upvotes.

  1. Yume: An Interactive World Generation Model 77 upvotes, #1 of 2025-07-24
  2. Sekai: A Video Dataset towards World Exploration 60 upvotes, #1 of 2025-06-19
  3. A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation 7 upvotes, #14 of 2025-06-16
  4. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
  5. LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis 25 upvotes, #6 of 2025-03-28
  6. MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scale Reinforcement Learning 53 upvotes, #3 of 2025-03-11
  7. GATE OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation 17 upvotes, #9 of 2024-12-03
  8. TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts 6 upvotes, #9 of 2024-10-24
  9. Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation 42 upvotes, #5 of 2024-10-10
  10. MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models 56 upvotes, #1 of 2024-08-07
  11. Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model 39 upvotes, #1 of 2024-07-26
  12. Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality 8 upvotes, #13 of 2024-06-17
  13. Adapting LLaMA Decoder to Vision Transformer 12 upvotes, #6 of 2024-04-11
  14. SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models 17 upvotes, #8 of 2024-02-09
  15. OneLLM: One Framework to Align All Modalities with Language 23 upvotes, #6 of 2023-12-06
  16. ImageBind-LLM: Multi-modality Instruction Tuning 17 upvotes, #7 of 2023-09-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.