Gang Yu

Gang Yu on Hugging Face Daily Papers: 21 papers, 10 in the top 3 of their day, 966 upvotes.

  1. REASONEDIT: Towards Reasoning-Enhanced Image Editing Models 45 upvotes, #3 of 2025-12-01
  2. Step-Audio-R1 Technical Report 51 upvotes, #4 of 2025-11-21
  3. RegionE: Adaptive Region-Aware Generation for Efficient Image Editing 25 upvotes, #10 of 2025-10-30
  4. WithAnyone: Towards Controllable and ID Consistent Image Generation 76 upvotes, #3 of 2025-10-17
  5. OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation 38 upvotes, #5 of 2025-06-10
  6. Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers 27 upvotes, #13 of 2025-06-04
  7. ViStoryBench: Comprehensive Benchmark Suite for Story Visualization 31 upvotes, #7 of 2025-06-02
  8. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models 41 upvotes, #5 of 2025-05-23
  9. Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets 60 upvotes, #3 of 2025-05-13
  10. Step1X-Edit: A Practical Framework for General Image Editing 85 upvotes, #2 of 2025-04-25
  11. OmniSVG: A Unified Scalable Vector Graphics Generation Model 141 upvotes, #1 of 2025-04-09
  12. MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D 23 upvotes, #5 of 2024-11-05
  13. MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers 21 upvotes, #6 of 2024-06-18
  14. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment 37 upvotes, #2 of 2024-03-11
  15. MovieLLM: Enhancing Long Video Understanding with AI-Generated Movies 21 upvotes, #1 of 2024-03-05
  16. AppAgent: Multimodal Agents as Smartphone Users 54 upvotes, #1 of 2023-12-22
  17. M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts 18 upvotes, #6 of 2023-12-19
  18. FaceStudio: Put Your Face Everywhere in Seconds 32 upvotes, #2 of 2023-12-06
  19. Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation 12 upvotes, #6 of 2023-06-30
  20. MotionGPT: Human Motion as a Foreign Language 28 upvotes, #2 of 2023-06-27
  21. StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation 5 upvotes, #4 of 2023-05-31

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.