Xinlong Wang
Xinlong Wang on Hugging Face Daily Papers: 19 papers, 11 in the top 3 of their day, 803 upvotes.
- Emu3.5: Native Multimodal Models are World Learners 102 upvotes, #2 of 2025-10-31
- Uniform Discrete Diffusion with Metric Path for Video Generation 39 upvotes, #6 of 2025-10-29
- Unified Vision-Language-Action Model 21 upvotes, #9 of 2025-06-25
- OmniGen2: Exploration to Advanced Multimodal Generation 68 upvotes, #2 of 2025-06-24
- End-to-End Vision Tokenizer Tuning 20 upvotes, #8 of 2025-05-16
- EVEv2: Improved Baselines for Encoder-Free Vision-Language Models 11 upvotes, #14 of 2025-02-11
- Autoregressive Video Generation without Vector Quantization 13 upvotes, #10 of 2024-12-19
- You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale 11 upvotes, #9 of 2024-12-10
- Emu3: Next-Token Prediction is All You Need 75 upvotes, #1 of 2024-09-30
- Diffusion Feedback Helps CLIP See Better 33 upvotes, #8 of 2024-07-30
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception 17 upvotes, #10 of 2024-07-12
- Unveiling Encoder-Free Vision-Language Models 45 upvotes, #1 of 2024-07-08
- EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters 31 upvotes, #3 of 2024-02-07
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model 62 upvotes, #1 of 2024-01-18
- Generative Multimodal Models are In-Context Learners 36 upvotes, #3 of 2023-12-21
- CapsFusion: Rethinking Image-Text Data at Scale 27 upvotes, #2 of 2023-11-01
- JudgeLM: Fine-tuned Large Language Models are Scalable Judges 35 upvotes, #1 of 2023-10-27
- 3D-GPT: Procedural 3D Modeling with Large Language Models 61 upvotes, #1 of 2023-10-20
- Generative Pretraining in Multimodality 23 upvotes, #3 of 2023-07-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.