Shengqiong Wu

Shengqiong Wu on Hugging Face Daily Papers: 11 papers, 4 in the top 3 of their day, 498 upvotes.

  1. V-RAE: Rethinking Video Latent Spaces for Generation 30 upvotes, #8 of 2026-08-19
  2. Audio-Visual Intelligence in Large Foundation Models 32 upvotes, #10 of 2026-05-08
  3. UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist 36 upvotes, #5 of 2025-11-14
  4. Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding 1 upvotes, #16 of 2025-09-16
  5. On Path to Multimodal Generalist: General-Level and General-Bench 72 upvotes, #1 of 2025-05-08
  6. VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models 21 upvotes, #8 of 2025-04-18
  7. JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization 49 upvotes, #6 of 2025-04-04
  8. Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation 71 upvotes, #1 of 2025-04-02
  9. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey 28 upvotes, #6 of 2025-03-18
  10. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding 50 upvotes, #1 of 2024-06-28
  11. NExT-GPT: Any-to-Any Multimodal LLM 79 upvotes, #2 of 2023-09-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.