Hao Fei
Hao Fei on Hugging Face Daily Papers: 21 papers, 6 in the top 3 of their day, 1,101 upvotes.
- Omni-IO Skills: Harnessing Your Agent Omni-Native 292 upvotes, #5 of 2026-09-30
- UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City 106 upvotes, #3 of 2026-08-28
- V-RAE: Rethinking Video Latent Spaces for Generation 30 upvotes, #8 of 2026-08-19
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 87 upvotes, #5 of 2026-08-14
- Mental World Modeling 103 upvotes, #2 of 2026-08-03
- Towards One-to-Many Temporal Grounding 7 upvotes, #22 of 2026-06-05
- Audio-Visual Intelligence in Large Foundation Models 32 upvotes, #10 of 2026-05-08
- Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment 9 upvotes, #12 of 2026-04-28
- JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation 14 upvotes, #10 of 2026-02-26
- SAMTok: Representing Any Mask with Two Words 41 upvotes, #9 of 2026-01-23
- JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation 18 upvotes, #9 of 2026-01-01
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist 36 upvotes, #5 of 2025-11-14
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding 1 upvotes, #16 of 2025-09-16
- On Path to Multimodal Generalist: General-Level and General-Bench 72 upvotes, #1 of 2025-05-08
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models 21 upvotes, #8 of 2025-04-18
- JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization 49 upvotes, #6 of 2025-04-04
- Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation 71 upvotes, #1 of 2025-04-02
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey 28 upvotes, #6 of 2025-03-18
- Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent 22 upvotes, #8 of 2024-11-05
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding 50 upvotes, #1 of 2024-06-28
- NExT-GPT: Any-to-Any Multimodal LLM 79 upvotes, #2 of 2023-09-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.