Xihui Liu

Xihui Liu on Hugging Face Daily Papers: 50 papers, 11 in the top 3 of their day, 1,420 upvotes.

  1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation 13 upvotes, #15 of 2026-03-13
  2. Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance 123 upvotes, #1 of 2025-12-10
  3. CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images 13 upvotes, #24 of 2025-10-14
  4. Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation 3 upvotes, #35 of 2025-10-13
  5. Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation 27 upvotes, #8 of 2025-09-19
  6. FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark 39 upvotes, #8 of 2025-09-12
  7. T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation 26 upvotes, #6 of 2025-08-26
  8. TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation 17 upvotes, #7 of 2025-07-25
  9. OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding 37 upvotes, #5 of 2025-07-11
  10. StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling 42 upvotes, #5 of 2025-07-09
  11. OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion 53 upvotes, #3 of 2025-07-09
  12. FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation 5 upvotes, #15 of 2025-06-26
  13. GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning 26 upvotes, #7 of 2025-06-25
  14. DreamCube: 3D Panorama Generation via Multi-plane Synchronization 20 upvotes, #8 of 2025-06-23
  15. AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation 22 upvotes, #15 of 2025-06-04
  16. GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning 26 upvotes, #10 of 2025-05-23
  17. A Survey of Interactive Generative Video 43 upvotes, #2 of 2025-05-02
  18. Personalized Text-to-Image Generation with Auto-Regressive Models 18 upvotes, #11 of 2025-04-23
  19. GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation 47 upvotes, #2 of 2025-04-14
  20. HoloPart: Generative 3D Part Amodal Segmentation 25 upvotes, #7 of 2025-04-11
  21. Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 36 upvotes, #4 of 2025-04-02
  22. AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset 10 upvotes, #11 of 2025-03-27
  23. Position: Interactive Generative Video as Next-Generation Game Engine 59 upvotes, #3 of 2025-03-25
  24. RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints 39 upvotes, #5 of 2025-03-24
  25. Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation 34 upvotes, #6 of 2025-03-24
  26. GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing 45 upvotes, #4 of 2025-03-14
  27. GameFactory: Creating New Games with Generative Interactive Videos 60 upvotes, #1 of 2025-01-21
  28. Parallelized Autoregressive Visual Generation 47 upvotes, #1 of 2024-12-23
  29. Moto: Latent Motion Token as the Bridging Language for Robot Manipulation 20 upvotes, #7 of 2024-12-09
  30. GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration 17 upvotes, #8 of 2024-12-09
  31. MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 14 upvotes, #12 of 2024-12-05
  32. SAMPart3D: Segment Any Part in 3D Objects 25 upvotes, #2 of 2024-11-13
  33. WorldSimBench: Towards Video Generation Models as World Simulators 16 upvotes, #3 of 2024-10-24
  34. PUMA: Empowering Unified MLLM with Multi-granular Visual Generation 51 upvotes, #4 of 2024-10-22
  35. LVD-2M: A Long-take Video Dataset with Temporally Dense Captions 18 upvotes, #6 of 2024-10-16
  36. Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding 17 upvotes, #5 of 2024-10-07
  37. Loong: Generating Minute-level Long Videos with Autoregressive Language Models 35 upvotes, #4 of 2024-10-04
  38. LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness 32 upvotes, #3 of 2024-09-27
  39. DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion 11 upvotes, #5 of 2024-09-26
  40. T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation 22 upvotes, #4 of 2024-07-24
  41. Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images 8 upvotes, #8 of 2024-07-09
  42. 4Diffusion: Multi-view Video Diffusion Model for 4D Generation 8 upvotes, #5 of 2024-06-03
  43. DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis 8 upvotes, #11 of 2024-05-24
  44. TC4D: Trajectory-Conditioned Text-to-4D Generation 13 upvotes, #4 of 2024-03-27
  45. FiT: Flexible Vision Transformer for Diffusion Model 48 upvotes, #2 of 2024-02-20
  46. Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation 11 upvotes, #10 of 2024-01-30
  47. Point Transformer V3: Simpler, Faster, Stronger 21 upvotes, #6 of 2023-12-18
  48. DreamComposer: Controllable 3D Object Generation via Multi-View Conditions 8 upvotes, #12 of 2023-12-07
  49. HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion 16 upvotes, #6 of 2023-10-13
  50. T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation 7 upvotes, #9 of 2023-07-14

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.