Mike Zheng Shou
Mike Zheng Shou on Hugging Face Daily Papers: 21 papers, 5 in the top 3 of their day, 670 upvotes.
- ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands 40 upvotes, #8 of 2026-01-14
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos 3 upvotes, #19 of 2025-12-12
- OmniPSD: Layered PSD Generation with Diffusion Transformer 45 upvotes, #2 of 2025-12-11
- Computer-Use Agents as Judges for Generative User Interface 50 upvotes, #5 of 2025-11-25
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale 32 upvotes, #6 of 2025-04-23
- Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model 11 upvotes, #10 of 2025-04-09
- Edit Transfer: Learning Image Editing via Vision In-Context Relations 24 upvotes, #8 of 2025-03-18
- VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning 13 upvotes, #13 of 2025-03-18
- TPDiff: Temporal Pyramid Video Diffusion Model 41 upvotes, #2 of 2025-03-13
- Automated Movie Generation via Multi-Agent CoT Planning 40 upvotes, #4 of 2025-03-11
- Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models 39 upvotes, #3 of 2025-03-04
- PhotoDoodle: Learning Artistic Image Editing from Few-Shot Pairwise Data 36 upvotes, #4 of 2025-02-24
- InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback 6 upvotes, #18 of 2025-02-24
- WorldGUI: Dynamic Testing for Comprehensive Desktop GUI Automation 24 upvotes, #8 of 2025-02-13
- ROICtrl: Boosting Instance Control for Visual Generation 77 upvotes, #1 of 2024-11-28
- Factorized Visual Tokenization and Generation 17 upvotes, #11 of 2024-11-26
- The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use 27 upvotes, #3 of 2024-11-18
- EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models 18 upvotes, #5 of 2024-10-14
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos 17 upvotes, #4 of 2024-10-02
- VideoLLM-online: Online Video Large Language Model for Streaming Video 19 upvotes, #8 of 2024-06-18
- VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.