Xintao Wang

Xintao Wang on Hugging Face Daily Papers: 46 papers, 14 in the top 3 of their day, 1,279 upvotes.

  1. ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling 154 upvotes, #1 of 2026-03-30
  2. Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control 24 upvotes, #11 of 2025-06-03
  3. ARIA: Training Language Agents with Intention-Driven Reward Aggregation 29 upvotes, #8 of 2025-06-03
  4. Flow-GRPO: Training Flow Matching Models via Online RL 68 upvotes, #2 of 2025-05-09
  5. A Survey of Interactive Generative Video 43 upvotes, #2 of 2025-05-02
  6. BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation 26 upvotes, #7 of 2025-04-23
  7. Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation 71 upvotes, #1 of 2025-04-02
  8. SketchVideo: Sketch-based Video Generation and Editing 21 upvotes, #9 of 2025-04-01
  9. Position: Interactive Generative Video as Next-Generation Game Engine 59 upvotes, #3 of 2025-03-25
  10. DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers 27 upvotes, #12 of 2025-03-21
  11. ReCamMaster: Camera-Controlled Generative Rendering from A Single Video 117 upvotes, #1 of 2025-03-17
  12. CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation 36 upvotes, #7 of 2025-02-13
  13. Improving Video Generation with Human Feedback 44 upvotes, #2 of 2025-01-24
  14. GameFactory: Creating New Games with Generative Interactive Videos 60 upvotes, #1 of 2025-01-21
  15. ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning 15 upvotes, #8 of 2025-01-13
  16. SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints 49 upvotes, #1 of 2024-12-12
  17. StyleMaster: Stylize Your Video with Artistic Generation and Translation 18 upvotes, #6 of 2024-12-12
  18. 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation 18 upvotes, #9 of 2024-12-11
  19. MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions 14 upvotes, #7 of 2024-07-10
  20. Image Conductor: Precision Control for Interactive Video Synthesis 8 upvotes, #15 of 2024-06-26
  21. VideoTetris: Towards Compositional Text-to-Video Generation 18 upvotes, #6 of 2024-06-07
  22. MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model 9 upvotes, #5 of 2024-05-31
  23. ReVideo: Remake a Video with Motion and Content Control 19 upvotes, #4 of 2024-05-24
  24. Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners 15 upvotes, #12 of 2024-02-28
  25. Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation 17 upvotes, #8 of 2024-02-19
  26. Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild 78 upvotes, #1 of 2024-01-25
  27. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models 14 upvotes, #7 of 2024-01-18
  28. Towards A Better Metric for Text-to-Video Generation 15 upvotes, #7 of 2024-01-17
  29. PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding 62 upvotes, #1 of 2023-12-08
  30. AnimateZero: Video Diffusion Models are Zero-Shot Image Animators 18 upvotes, #4 of 2023-12-08
  31. MagicStick: Controllable Video Editing via Control Handle Transformations 10 upvotes, #11 of 2023-12-07
  32. MotionCtrl: A Unified and Flexible Motion Controller for Video Generation 21 upvotes, #6 of 2023-12-07
  33. X-Adapter: Adding Universal Compatibility of Plugins for Upgraded Diffusion Model 26 upvotes, #5 of 2023-12-06
  34. VideoCrafter1: Open Diffusion Models for High-Quality Video Generation 16 upvotes, #3 of 2023-10-31
  35. CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models 10 upvotes, #5 of 2023-10-31
  36. FreeNoise: Tuning-Free Longer Video Diffusion Via Noise Rescheduling 10 upvotes, #4 of 2023-10-24
  37. EvalCrafter: Benchmarking and Evaluating Large Video Generation Models 17 upvotes, #5 of 2023-10-18
  38. StyleAdapter: A Single-Pass LoRA-Free Model for Stylized Image Generation 12 upvotes, #8 of 2023-09-06
  39. Planting a SEED of Vision in Large Language Model 12 upvotes, #7 of 2023-07-18
  40. Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation 11 upvotes, #6 of 2023-07-14
  41. DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models 35 upvotes, #3 of 2023-07-06
  42. DreamDiffusion: Generating High-Quality Images from Brain EEG Signals 32 upvotes, #3 of 2023-06-30
  43. Make-Your-Video: Customized Video Generation Using Textual and Structural Guidance 6 upvotes, #7 of 2023-06-02
  44. Inserting Anybody in Diffusion Models via Celeb Basis 3 upvotes, #13 of 2023-06-02
  45. Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models 6 upvotes, #4 of 2023-05-30
  46. TaleCrafter: Interactive Story Visualization with Multiple Characters 5 upvotes, #6 of 2023-05-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.