Mike Zheng Shou

Mike Zheng Shou on Hugging Face Daily Papers: 21 papers, 5 in the top 3 of their day, 670 upvotes.

  1. ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands 40 upvotes, #8 of 2026-01-14
  2. H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos 3 upvotes, #19 of 2025-12-12
  3. OmniPSD: Layered PSD Generation with Diffusion Transformer 45 upvotes, #2 of 2025-12-11
  4. Computer-Use Agents as Judges for Generative User Interface 50 upvotes, #5 of 2025-11-25
  5. LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale 32 upvotes, #6 of 2025-04-23
  6. Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model 11 upvotes, #10 of 2025-04-09
  7. Edit Transfer: Learning Image Editing via Vision In-Context Relations 24 upvotes, #8 of 2025-03-18
  8. VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning 13 upvotes, #13 of 2025-03-18
  9. TPDiff: Temporal Pyramid Video Diffusion Model 41 upvotes, #2 of 2025-03-13
  10. Automated Movie Generation via Multi-Agent CoT Planning 40 upvotes, #4 of 2025-03-11
  11. Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models 39 upvotes, #3 of 2025-03-04
  12. PhotoDoodle: Learning Artistic Image Editing from Few-Shot Pairwise Data 36 upvotes, #4 of 2025-02-24
  13. InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback 6 upvotes, #18 of 2025-02-24
  14. WorldGUI: Dynamic Testing for Comprehensive Desktop GUI Automation 24 upvotes, #8 of 2025-02-13
  15. ROICtrl: Boosting Instance Control for Visual Generation 77 upvotes, #1 of 2024-11-28
  16. Factorized Visual Tokenization and Generation 17 upvotes, #11 of 2024-11-26
  17. The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use 27 upvotes, #3 of 2024-11-18
  18. EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models 18 upvotes, #5 of 2024-10-14
  19. One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos 17 upvotes, #4 of 2024-10-02
  20. VideoLLM-online: Online Video Large Language Model for Streaming Video 19 upvotes, #8 of 2024-06-18
  21. VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.