xuhaiyang
xuhaiyang on Hugging Face Daily Papers: 25 papers, 13 in the top 3 of their day, 1,135 upvotes.
- Qwen-AgentWorld: Language World Models for General Agents 144 upvotes, #1 of 2026-06-24
- ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents 27 upvotes, #14 of 2026-05-13
- Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents 46 upvotes, #2 of 2026-02-20
- AgentOCR: Reimagining Agent History via Optical Self-Compression 27 upvotes, #8 of 2026-01-12
- Qwen3-VL Technical Report 120 upvotes, #1 of 2025-12-04
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents 22 upvotes, #9 of 2025-10-29
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning 45 upvotes, #2 of 2025-09-16
- Mobile-Agent-v3: Foundamental Agents for GUI Automation 55 upvotes, #3 of 2025-08-22
- Perception-Aware Policy Optimization for Multimodal Reasoning 42 upvotes, #5 of 2025-07-10
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation 15 upvotes, #9 of 2025-06-11
- VLM-R^3: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought 7 upvotes, #33 of 2025-05-23
- Mobile-Agent-V: Learning Mobile Device Operation Through Video-Guided Multi-Agent Collaboration 11 upvotes, #13 of 2025-02-25
- Qwen2.5-VL Technical Report 146 upvotes, #1 of 2025-02-20
- Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks 26 upvotes, #9 of 2025-01-22
- SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization 20 upvotes, #2 of 2024-11-20
- mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding 22 upvotes, #6 of 2024-09-06
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 27 upvotes, #3 of 2024-08-12
- MIBench: Evaluating Multimodal Large Language Models over Multiple Images 6 upvotes, #18 of 2024-07-23
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration 26 upvotes, #3 of 2024-06-06
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding 24 upvotes, #1 of 2024-03-20
- Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception 20 upvotes, #7 of 2024-01-30
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration 22 upvotes, #2 of 2023-11-09
- ModelScope-Agent: Building Your Customizable Agent System with Open-source Large Language Models 22 upvotes, #3 of 2023-09-06
- mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding 16 upvotes, #3 of 2023-07-07
- Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks 2 upvotes, #9 of 2023-06-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.