Xihui Liu
Xihui Liu on Hugging Face Daily Papers: 50 papers, 11 in the top 3 of their day, 1,420 upvotes.
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation 13 upvotes, #15 of 2026-03-13
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance 123 upvotes, #1 of 2025-12-10
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images 13 upvotes, #24 of 2025-10-14
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation 3 upvotes, #35 of 2025-10-13
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation 27 upvotes, #8 of 2025-09-19
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark 39 upvotes, #8 of 2025-09-12
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation 26 upvotes, #6 of 2025-08-26
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation 17 upvotes, #7 of 2025-07-25
- OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding 37 upvotes, #5 of 2025-07-11
- StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling 42 upvotes, #5 of 2025-07-09
- OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion 53 upvotes, #3 of 2025-07-09
- FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation 5 upvotes, #15 of 2025-06-26
- GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning 26 upvotes, #7 of 2025-06-25
- DreamCube: 3D Panorama Generation via Multi-plane Synchronization 20 upvotes, #8 of 2025-06-23
- AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation 22 upvotes, #15 of 2025-06-04
- GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning 26 upvotes, #10 of 2025-05-23
- A Survey of Interactive Generative Video 43 upvotes, #2 of 2025-05-02
- Personalized Text-to-Image Generation with Auto-Regressive Models 18 upvotes, #11 of 2025-04-23
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation 47 upvotes, #2 of 2025-04-14
- HoloPart: Generative 3D Part Amodal Segmentation 25 upvotes, #7 of 2025-04-11
- Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 36 upvotes, #4 of 2025-04-02
- AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset 10 upvotes, #11 of 2025-03-27
- Position: Interactive Generative Video as Next-Generation Game Engine 59 upvotes, #3 of 2025-03-25
- RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints 39 upvotes, #5 of 2025-03-24
- Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation 34 upvotes, #6 of 2025-03-24
- GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing 45 upvotes, #4 of 2025-03-14
- GameFactory: Creating New Games with Generative Interactive Videos 60 upvotes, #1 of 2025-01-21
- Parallelized Autoregressive Visual Generation 47 upvotes, #1 of 2024-12-23
- Moto: Latent Motion Token as the Bridging Language for Robot Manipulation 20 upvotes, #7 of 2024-12-09
- GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration 17 upvotes, #8 of 2024-12-09
- MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 14 upvotes, #12 of 2024-12-05
- SAMPart3D: Segment Any Part in 3D Objects 25 upvotes, #2 of 2024-11-13
- WorldSimBench: Towards Video Generation Models as World Simulators 16 upvotes, #3 of 2024-10-24
- PUMA: Empowering Unified MLLM with Multi-granular Visual Generation 51 upvotes, #4 of 2024-10-22
- LVD-2M: A Long-take Video Dataset with Temporally Dense Captions 18 upvotes, #6 of 2024-10-16
- Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding 17 upvotes, #5 of 2024-10-07
- Loong: Generating Minute-level Long Videos with Autoregressive Language Models 35 upvotes, #4 of 2024-10-04
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness 32 upvotes, #3 of 2024-09-27
- DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion 11 upvotes, #5 of 2024-09-26
- T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation 22 upvotes, #4 of 2024-07-24
- Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images 8 upvotes, #8 of 2024-07-09
- 4Diffusion: Multi-view Video Diffusion Model for 4D Generation 8 upvotes, #5 of 2024-06-03
- DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis 8 upvotes, #11 of 2024-05-24
- TC4D: Trajectory-Conditioned Text-to-4D Generation 13 upvotes, #4 of 2024-03-27
- FiT: Flexible Vision Transformer for Diffusion Model 48 upvotes, #2 of 2024-02-20
- Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation 11 upvotes, #10 of 2024-01-30
- Point Transformer V3: Simpler, Faster, Stronger 21 upvotes, #6 of 2023-12-18
- DreamComposer: Controllable 3D Object Generation via Multi-View Conditions 8 upvotes, #12 of 2023-12-07
- HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion 16 upvotes, #6 of 2023-10-13
- T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation 7 upvotes, #9 of 2023-07-14
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.