Xizhou Zhu
Xizhou Zhu on Hugging Face Daily Papers: 12 papers, 8 in the top 3 of their day, 667 upvotes.
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models 14 upvotes, #6 of 2025-07-21
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models 71 upvotes, #1 of 2025-04-24
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy 47 upvotes, #2 of 2025-03-27
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
- Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding 7 upvotes, #8 of 2025-01-16
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 35 upvotes, #3 of 2024-12-16
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
- Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization 61 upvotes, #1 of 2024-11-22
- MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models 56 upvotes, #1 of 2024-08-07
- Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
- ControlLLM: Augment Language Models with Tools by Searching on Graphs 18 upvotes, #3 of 2023-10-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.