Dongzhi Jiang
Dongzhi Jiang on Hugging Face Daily Papers: 16 papers, 5 in the top 3 of their day, 511 upvotes.
- GenClaw: Code-Driven Agentic Image Generation 38 upvotes, #9 of 2026-05-29
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation 36 upvotes, #6 of 2026-03-10
- Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation 22 upvotes, #22 of 2026-02-03
- Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation 42 upvotes, #3 of 2025-12-12
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards 24 upvotes, #5 of 2025-12-08
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation 11 upvotes, #18 of 2025-12-05
- Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark 32 upvotes, #9 of 2025-10-31
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation 24 upvotes, #7 of 2025-08-14
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning 12 upvotes, #21 of 2025-06-06
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT 39 upvotes, #3 of 2025-05-02
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency 27 upvotes, #9 of 2025-02-14
- EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM 21 upvotes, #7 of 2024-12-13
- MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines 33 upvotes, #3 of 2024-09-20
- MAVIS: Mathematical Visual Instruction Tuning 26 upvotes, #5 of 2024-07-12
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching 28 upvotes, #2 of 2024-04-05
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? 45 upvotes, #1 of 2024-03-22
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.