Min-Hung Chen
Min-Hung Chen on Hugging Face Daily Papers: 27 papers, 3 in the top 3 of their day, 914 upvotes.
- TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining 18 upvotes, #42 of 2026-09-29
- ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding 19 upvotes, #34 of 2026-09-09
- Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction 46 upvotes, #11 of 2026-09-04
- PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration 4 upvotes, #19 of 2026-08-24
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning 101 upvotes, #3 of 2026-06-12
- Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them 14 upvotes, #19 of 2026-06-08
- Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning 50 upvotes, #6 of 2026-01-15
- 3AM: Segment Anything with Geometric Consistency in Videos 33 upvotes, #10 of 2026-01-14
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization 191 upvotes, #1 of 2026-01-09
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation 42 upvotes, #6 of 2025-12-22
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in 7 upvotes, #23 of 2025-12-17
- BlurDM: A Blur Diffusion Model for Image Deblurring 2 upvotes, #24 of 2025-12-04
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models 4 upvotes, #26 of 2025-11-11
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning 15 upvotes, #15 of 2025-10-20
- TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control 7 upvotes, #20 of 2025-10-13
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation 2 upvotes, #37 of 2025-10-13
- LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models 1 upvotes, #33 of 2025-10-06
- V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts 3 upvotes, #26 of 2025-09-23
- MovieCORE: COgnitive REasoning in Movies 5 upvotes, #19 of 2025-08-27
- Autoregressive Universal Video Segmentation Model 26 upvotes, #9 of 2025-08-27
- LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos 55 upvotes, #2 of 2025-08-20
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning 35 upvotes, #4 of 2025-07-23
- V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multi-Modal Large Language Models 4 upvotes, #18 of 2025-02-17
- AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting 28 upvotes, #6 of 2025-02-10
- Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks 31 upvotes, #4 of 2025-01-15
- Hymba: A Hybrid-head Architecture for Small Language Models 37 upvotes, #4 of 2024-11-22
- EoRA: Training-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation 6 upvotes, #14 of 2024-10-29
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.