Xizhou Zhu

Xizhou Zhu on Hugging Face Daily Papers: 12 papers, 8 in the top 3 of their day, 667 upvotes.

  1. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models 14 upvotes, #6 of 2025-07-21
  2. VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models 71 upvotes, #1 of 2025-04-24
  3. Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy 47 upvotes, #2 of 2025-03-27
  4. VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
  5. Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding 7 upvotes, #8 of 2025-01-16
  6. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 35 upvotes, #3 of 2024-12-16
  7. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
  8. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization 61 upvotes, #1 of 2024-11-22
  9. MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models 56 upvotes, #1 of 2024-08-07
  10. Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
  11. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
  12. ControlLLM: Augment Language Models with Tools by Searching on Graphs 18 upvotes, #3 of 2023-10-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.