Wenxuan Huang

Wenxuan Huang on Hugging Face Daily Papers: 19 papers, 3 in the top 3 of their day, 839 upvotes.

  1. Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 50 upvotes, #5 of 2026-08-05
  2. JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents 124 upvotes, #3 of 2026-07-28
  3. VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation 27 upvotes, #12 of 2026-05-19
  4. SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation 10 upvotes, #25 of 2026-05-11
  5. Flow-OPD: On-Policy Distillation for Flow Matching Models 95 upvotes, #2 of 2026-05-11
  6. OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents 96 upvotes, #4 of 2026-05-07
  7. SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents 22 upvotes, #10 of 2026-04-21
  8. Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis 46 upvotes, #10 of 2026-04-01
  9. GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant 1 upvotes, #33 of 2026-03-04
  10. Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models 149 upvotes, #3 of 2026-02-03
  11. Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models 124 upvotes, #4 of 2026-02-03
  12. UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision 44 upvotes, #4 of 2026-01-07
  13. Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models 9 upvotes, #21 of 2025-11-04
  14. Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning? 6 upvotes, #24 of 2025-10-08
  15. Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models 9 upvotes, #24 of 2025-10-03
  16. IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video? 3 upvotes, #64 of 2025-09-30
  17. Interleaving Reasoning for Better Text-to-Image Generation 13 upvotes, #10 of 2025-09-09
  18. VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning 43 upvotes, #4 of 2025-04-11
  19. Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models 22 upvotes, #10 of 2025-03-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.