Hao Fei

Hao Fei on Hugging Face Daily Papers: 21 papers, 6 in the top 3 of their day, 1,101 upvotes.

  1. Omni-IO Skills: Harnessing Your Agent Omni-Native 292 upvotes, #5 of 2026-09-30
  2. UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City 106 upvotes, #3 of 2026-08-28
  3. V-RAE: Rethinking Video Latent Spaces for Generation 30 upvotes, #8 of 2026-08-19
  4. OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 87 upvotes, #5 of 2026-08-14
  5. Mental World Modeling 103 upvotes, #2 of 2026-08-03
  6. Towards One-to-Many Temporal Grounding 7 upvotes, #22 of 2026-06-05
  7. Audio-Visual Intelligence in Large Foundation Models 32 upvotes, #10 of 2026-05-08
  8. Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment 9 upvotes, #12 of 2026-04-28
  9. JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation 14 upvotes, #10 of 2026-02-26
  10. SAMTok: Representing Any Mask with Two Words 41 upvotes, #9 of 2026-01-23
  11. JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation 18 upvotes, #9 of 2026-01-01
  12. UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist 36 upvotes, #5 of 2025-11-14
  13. Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding 1 upvotes, #16 of 2025-09-16
  14. On Path to Multimodal Generalist: General-Level and General-Bench 72 upvotes, #1 of 2025-05-08
  15. VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models 21 upvotes, #8 of 2025-04-18
  16. JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization 49 upvotes, #6 of 2025-04-04
  17. Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation 71 upvotes, #1 of 2025-04-02
  18. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey 28 upvotes, #6 of 2025-03-18
  19. Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent 22 upvotes, #8 of 2024-11-05
  20. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding 50 upvotes, #1 of 2024-06-28
  21. NExT-GPT: Any-to-Any Multimodal LLM 79 upvotes, #2 of 2023-09-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.