shuai bai
shuai bai on Hugging Face Daily Papers: 10 papers, 7 in the top 3 of their day, 1,105 upvotes.
- Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments 138 upvotes, #2 of 2026-05-29
- Qwen3-VL Technical Report 120 upvotes, #1 of 2025-12-04
- Qwen2.5-Omni Technical Report 113 upvotes, #1 of 2025-03-27
- Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think 26 upvotes, #7 of 2025-02-28
- Qwen2.5-VL Technical Report 146 upvotes, #1 of 2025-02-20
- Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey 49 upvotes, #3 of 2024-12-30
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 63 upvotes, #2 of 2024-09-19
- Qwen2 Technical Report 142 upvotes, #1 of 2024-07-16
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models 23 upvotes, #4 of 2024-03-12
- Qwen Technical Report 39 upvotes, #4 of 2023-09-29
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.