zhumuzhi

zhumuzhi on Hugging Face Daily Papers: 11 papers, 1 in the top 3 of their day, 330 upvotes.

  1. CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding 8 upvotes, #48 of 2026-10-01
  2. Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching 16 upvotes, #18 of 2026-06-04
  3. Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? 30 upvotes, #12 of 2026-06-02
  4. Exploring Spatial Intelligence from a Generative Perspective 21 upvotes, #7 of 2026-04-23
  5. OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering 25 upvotes, #16 of 2026-04-10
  6. LLaDA2.1: Speeding Up Text Diffusion via Token Editing 66 upvotes, #9 of 2026-02-10
  7. Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality 46 upvotes, #3 of 2025-12-10
  8. Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO 14 upvotes, #29 of 2025-05-28
  9. Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration 17 upvotes, #21 of 2025-05-27
  10. SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories 24 upvotes, #10 of 2025-03-12
  11. DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks 51 upvotes, #4 of 2025-02-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.