Min-Hung Chen

Min-Hung Chen on Hugging Face Daily Papers: 27 papers, 3 in the top 3 of their day, 914 upvotes.

  1. TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining 18 upvotes, #42 of 2026-09-29
  2. ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding 19 upvotes, #34 of 2026-09-09
  3. Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction 46 upvotes, #11 of 2026-09-04
  4. PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration 4 upvotes, #19 of 2026-08-24
  5. SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning 101 upvotes, #3 of 2026-06-12
  6. Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them 14 upvotes, #19 of 2026-06-08
  7. Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning 50 upvotes, #6 of 2026-01-15
  8. 3AM: Segment Anything with Geometric Consistency in Videos 33 upvotes, #10 of 2026-01-14
  9. GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization 191 upvotes, #1 of 2026-01-09
  10. 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation 42 upvotes, #6 of 2025-12-22
  11. Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in 7 upvotes, #23 of 2025-12-17
  12. BlurDM: A Blur Diffusion Model for Image Deblurring 2 upvotes, #24 of 2025-12-04
  13. VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models 4 upvotes, #26 of 2025-11-11
  14. DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning 15 upvotes, #15 of 2025-10-20
  15. TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control 7 upvotes, #20 of 2025-10-13
  16. Temporal Prompting Matters: Rethinking Referring Video Object Segmentation 2 upvotes, #37 of 2025-10-13
  17. LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models 1 upvotes, #33 of 2025-10-06
  18. V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts 3 upvotes, #26 of 2025-09-23
  19. MovieCORE: COgnitive REasoning in Movies 5 upvotes, #19 of 2025-08-27
  20. Autoregressive Universal Video Segmentation Model 26 upvotes, #9 of 2025-08-27
  21. LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos 55 upvotes, #2 of 2025-08-20
  22. ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning 35 upvotes, #4 of 2025-07-23
  23. V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multi-Modal Large Language Models 4 upvotes, #18 of 2025-02-17
  24. AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting 28 upvotes, #6 of 2025-02-10
  25. Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks 31 upvotes, #4 of 2025-01-15
  26. Hymba: A Hybrid-head Architecture for Small Language Models 37 upvotes, #4 of 2024-11-22
  27. EoRA: Training-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation 6 upvotes, #14 of 2024-10-29

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.