Wenxuan Huang
Wenxuan Huang on Hugging Face Daily Papers: 19 papers, 3 in the top 3 of their day, 839 upvotes.
- Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 50 upvotes, #5 of 2026-08-05
- JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents 124 upvotes, #3 of 2026-07-28
- VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation 27 upvotes, #12 of 2026-05-19
- SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation 10 upvotes, #25 of 2026-05-11
- Flow-OPD: On-Policy Distillation for Flow Matching Models 95 upvotes, #2 of 2026-05-11
- OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents 96 upvotes, #4 of 2026-05-07
- SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents 22 upvotes, #10 of 2026-04-21
- Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis 46 upvotes, #10 of 2026-04-01
- GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant 1 upvotes, #33 of 2026-03-04
- Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models 149 upvotes, #3 of 2026-02-03
- Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models 124 upvotes, #4 of 2026-02-03
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision 44 upvotes, #4 of 2026-01-07
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models 9 upvotes, #21 of 2025-11-04
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning? 6 upvotes, #24 of 2025-10-08
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models 9 upvotes, #24 of 2025-10-03
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video? 3 upvotes, #64 of 2025-09-30
- Interleaving Reasoning for Better Text-to-Image Generation 13 upvotes, #10 of 2025-09-09
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning 43 upvotes, #4 of 2025-04-11
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models 22 upvotes, #10 of 2025-03-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.