Weiyun Wang

Weiyun Wang on Hugging Face Daily Papers: 18 papers, 10 in the top 3 of their day, 1,620 upvotes.

  1. ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data 101 upvotes, #1 of 2025-09-19
  2. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 170 upvotes, #1 of 2025-08-26
  3. Intern-S1: A Scientific Multimodal Foundation Model 242 upvotes, #1 of 2025-08-22
  4. MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents 28 upvotes, #3 of 2025-07-28
  5. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models 14 upvotes, #6 of 2025-07-21
  6. AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning 37 upvotes, #6 of 2025-07-18
  7. VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models 71 upvotes, #1 of 2025-04-24
  8. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
  9. VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
  10. OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference 67 upvotes, #1 of 2025-02-26
  11. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
  12. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization 61 upvotes, #1 of 2024-11-22
  13. Seeing and Understanding: Bridging Vision with Chemical Knowledge Via ChemVLM 19 upvotes, #6 of 2024-08-14
  14. OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text 28 upvotes, #6 of 2024-06-17
  15. Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
  16. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
  17. The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World 13 upvotes, #8 of 2023-08-04
  18. InternChat: Solving Vision-Centric Tasks by Interacting with Chatbots Beyond Language 5 upvotes, #5 of 2023-05-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.