Ying Shan

Ying Shan on Hugging Face Daily Papers: 31 papers, 4 in the top 3 of their day, 765 upvotes.

  1. Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? 26 upvotes, #17 of 2025-05-28
  2. Cobra: Efficient Line Art COlorization with BRoAder References 27 upvotes, #4 of 2025-04-17
  3. AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction 57 upvotes, #4 of 2025-04-03
  4. GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors 28 upvotes, #7 of 2025-04-02
  5. Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 36 upvotes, #4 of 2025-04-02
  6. GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers 15 upvotes, #9 of 2025-03-27
  7. BlobCtrl: A Unified and Flexible Framework for Element-level Image Generation and Editing 24 upvotes, #8 of 2025-03-18
  8. VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control 21 upvotes, #10 of 2025-03-10
  9. TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models 17 upvotes, #12 of 2025-03-10
  10. VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models 13 upvotes, #7 of 2024-12-30
  11. DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation 9 upvotes, #10 of 2024-12-20
  12. BrushEdit: All-In-One Image Inpainting and Editing 33 upvotes, #3 of 2024-12-17
  13. ColorFlow: Retrieval-Augmented Image Sequence Colorization 26 upvotes, #5 of 2024-12-17
  14. FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction 7 upvotes, #19 of 2024-12-13
  15. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation 13 upvotes, #7 of 2024-12-10
  16. Moto: Latent Motion Token as the Bridging Language for Robot Manipulation 20 upvotes, #7 of 2024-12-09
  17. NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images 18 upvotes, #8 of 2024-12-05
  18. Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation 23 upvotes, #3 of 2024-09-09
  19. DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos 31 upvotes, #5 of 2024-09-04
  20. SEED-Story: Multimodal Long Story Generation with Large Language Model 18 upvotes, #7 of 2024-07-12
  21. SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation 16 upvotes, #7 of 2024-04-23
  22. Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation 17 upvotes, #8 of 2024-02-19
  23. Advances in 3D Generation: A Survey 19 upvotes, #6 of 2024-02-01
  24. YOLO-World: Real-Time Open-Vocabulary Object Detection 44 upvotes, #2 of 2024-01-31
  25. TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts 10 upvotes, #6 of 2024-01-29
  26. Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities 13 upvotes, #8 of 2024-01-26
  27. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models 14 upvotes, #7 of 2024-01-18
  28. Towards A Better Metric for Text-to-Video Generation 15 upvotes, #7 of 2024-01-17
  29. AnimateZero: Video Diffusion Models are Zero-Shot Image Animators 18 upvotes, #4 of 2023-12-08
  30. PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding 62 upvotes, #1 of 2023-12-08
  31. MotionCtrl: A Unified and Flexible Motion Controller for Video Generation 21 upvotes, #6 of 2023-12-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.