Le Zhuo
Le Zhuo on Hugging Face Daily Papers: 15 papers, 4 in the top 3 of their day, 398 upvotes.
- PICABench: How Far Are We from Physically Realistic Image Editing? 60 upvotes, #3 of 2025-10-21
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding 49 upvotes, #3 of 2025-10-09
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals 15 upvotes, #12 of 2025-10-07
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT 39 upvotes, #3 of 2025-05-02
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning 15 upvotes, #14 of 2025-04-23
- VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning 40 upvotes, #5 of 2025-04-11
- OmniCaptioner: One Captioner to Rule Them All 17 upvotes, #8 of 2025-04-10
- Lumina-Image 2.0: A Unified and Efficient Image Generative Framework 18 upvotes, #11 of 2025-03-28
- IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models 13 upvotes, #8 of 2025-01-24
- Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation 7 upvotes, #14 of 2024-12-16
- VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection 10 upvotes, #8 of 2024-11-25
- I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow 5 upvotes, #14 of 2024-10-14
- PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions 21 upvotes, #4 of 2024-09-24
- LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation 19 upvotes, #4 of 2024-08-29
- Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining 29 upvotes, #2 of 2024-08-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.