xuhaiyang

xuhaiyang on Hugging Face Daily Papers: 25 papers, 13 in the top 3 of their day, 1,135 upvotes.

  1. Qwen-AgentWorld: Language World Models for General Agents 144 upvotes, #1 of 2026-06-24
  2. ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents 27 upvotes, #14 of 2026-05-13
  3. Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents 46 upvotes, #2 of 2026-02-20
  4. AgentOCR: Reimagining Agent History via Optical Self-Compression 27 upvotes, #8 of 2026-01-12
  5. Qwen3-VL Technical Report 120 upvotes, #1 of 2025-12-04
  6. OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents 22 upvotes, #9 of 2025-10-29
  7. UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning 45 upvotes, #2 of 2025-09-16
  8. Mobile-Agent-v3: Foundamental Agents for GUI Automation 55 upvotes, #3 of 2025-08-22
  9. Perception-Aware Policy Optimization for Multimodal Reasoning 42 upvotes, #5 of 2025-07-10
  10. Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation 15 upvotes, #9 of 2025-06-11
  11. VLM-R^3: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought 7 upvotes, #33 of 2025-05-23
  12. Mobile-Agent-V: Learning Mobile Device Operation Through Video-Guided Multi-Agent Collaboration 11 upvotes, #13 of 2025-02-25
  13. Qwen2.5-VL Technical Report 146 upvotes, #1 of 2025-02-20
  14. Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks 26 upvotes, #9 of 2025-01-22
  15. SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization 20 upvotes, #2 of 2024-11-20
  16. mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding 22 upvotes, #6 of 2024-09-06
  17. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 27 upvotes, #3 of 2024-08-12
  18. MIBench: Evaluating Multimodal Large Language Models over Multiple Images 6 upvotes, #18 of 2024-07-23
  19. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration 26 upvotes, #3 of 2024-06-06
  20. mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding 24 upvotes, #1 of 2024-03-20
  21. Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception 20 upvotes, #7 of 2024-01-30
  22. mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration 22 upvotes, #2 of 2023-11-09
  23. ModelScope-Agent: Building Your Customizable Agent System with Open-source Large Language Models 22 upvotes, #3 of 2023-09-06
  24. mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding 16 upvotes, #3 of 2023-07-07
  25. Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks 2 upvotes, #9 of 2023-06-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.