Fan Zhang
Fan Zhang on Hugging Face Daily Papers: 10 papers, 6 in the top 3 of their day, 451 upvotes.
- Emu3.5: Native Multimodal Models are World Learners 102 upvotes, #2 of 2025-10-31
- Uniform Discrete Diffusion with Metric Path for Video Generation 39 upvotes, #6 of 2025-10-29
- End-to-End Vision Tokenizer Tuning 20 upvotes, #8 of 2025-05-16
- Emu3: Next-Token Prediction is All You Need 75 upvotes, #1 of 2024-09-30
- Diffusion Feedback Helps CLIP See Better 33 upvotes, #8 of 2024-07-30
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception 17 upvotes, #10 of 2024-07-12
- EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters 31 upvotes, #3 of 2024-02-07
- Generative Multimodal Models are In-Context Learners 36 upvotes, #3 of 2023-12-21
- CapsFusion: Rethinking Image-Text Data at Scale 27 upvotes, #2 of 2023-11-01
- Generative Pretraining in Multimodality 23 upvotes, #3 of 2023-07-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.