Fan Zhang

Fan Zhang on Hugging Face Daily Papers: 10 papers, 6 in the top 3 of their day, 451 upvotes.

  1. Emu3.5: Native Multimodal Models are World Learners 102 upvotes, #2 of 2025-10-31
  2. Uniform Discrete Diffusion with Metric Path for Video Generation 39 upvotes, #6 of 2025-10-29
  3. End-to-End Vision Tokenizer Tuning 20 upvotes, #8 of 2025-05-16
  4. Emu3: Next-Token Prediction is All You Need 75 upvotes, #1 of 2024-09-30
  5. Diffusion Feedback Helps CLIP See Better 33 upvotes, #8 of 2024-07-30
  6. DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception 17 upvotes, #10 of 2024-07-12
  7. EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters 31 upvotes, #3 of 2024-02-07
  8. Generative Multimodal Models are In-Context Learners 36 upvotes, #3 of 2023-12-21
  9. CapsFusion: Rethinking Image-Text Data at Scale 27 upvotes, #2 of 2023-11-01
  10. Generative Pretraining in Multimodality 23 upvotes, #3 of 2023-07-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.