Xinggang Wang
Xinggang Wang on Hugging Face Daily Papers: 19 papers, 6 in the top 3 of their day, 625 upvotes.
- Mixture-of-Depths Attention 77 upvotes, #7 of 2026-03-17
- Towards Scalable Pre-training of Visual Tokenizers for Generation 91 upvotes, #3 of 2025-12-16
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models 17 upvotes, #6 of 2025-12-11
- PixelHacker: Image Inpainting with Structural and Semantic Consistency 40 upvotes, #1 of 2025-05-05
- GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding 18 upvotes, #16 of 2025-03-14
- OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models 16 upvotes, #14 of 2025-03-12
- AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning 19 upvotes, #12 of 2025-03-11
- RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning 36 upvotes, #3 of 2025-02-20
- Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models 34 upvotes, #6 of 2025-01-03
- DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving 14 upvotes, #9 of 2024-11-28
- ControlAR: Controllable Image Generation with Autoregressive Models 7 upvotes, #7 of 2024-10-09
- LKCell: Efficient Cell Nuclei Instance Segmentation with Large Convolution Kernels 10 upvotes, #9 of 2024-07-26
- Visual Text Generation in the Wild 7 upvotes, #11 of 2024-07-22
- EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model 6 upvotes, #8 of 2024-07-01
- GaussianDreamerPro: Text to Manipulable 3D Gaussians with Highly Enhanced Quality 9 upvotes, #5 of 2024-07-01
- YOLO-World: Real-Time Open-Vocabulary Object Detection 44 upvotes, #2 of 2024-01-31
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model 62 upvotes, #1 of 2024-01-18
- JudgeLM: Fine-tuned Large Language Models are Scalable Judges 35 upvotes, #1 of 2023-10-27
- GaussianDreamer: Fast Generation from Text to 3D Gaussian Splatting with Point Cloud Priors 17 upvotes, #4 of 2023-10-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.