Yongming Rao

Yongming Rao on Hugging Face Daily Papers: 19 papers, 5 in the top 3 of their day, 812 upvotes.

  1. Transferring the Intelligence of VLMs to Robotic Control 118 upvotes, #5 of 2026-09-22
  2. GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots 16 upvotes, #17 of 2026-06-30
  3. ViQ: Text-Aligned Visual Quantized Representations at Any Resolution 38 upvotes, #6 of 2026-06-26
  4. Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack 15 upvotes, #17 of 2026-06-15
  5. GEM: Generative Supervision Helps Embodied Intelligence 41 upvotes, #8 of 2026-05-28
  6. HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents 182 upvotes, #4 of 2026-04-10
  7. Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models 12 upvotes, #22 of 2026-03-24
  8. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training 90 upvotes, #1 of 2026-03-13
  9. GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization 89 upvotes, #2 of 2025-11-24
  10. Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs 49 upvotes, #4 of 2025-10-16
  11. X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again 37 upvotes, #2 of 2025-07-30
  12. SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs 16 upvotes, #17 of 2025-06-06
  13. Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment 21 upvotes, #7 of 2025-02-07
  14. Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 19 upvotes, #7 of 2024-11-22
  15. Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 22 upvotes, #5 of 2024-09-20
  16. Coarse Correspondence Elicit 3D Spacetime Understanding in Multimodal Language Model 20 upvotes, #4 of 2024-08-02
  17. Efficient Inference of Vision Instruction-Following Models with Elastic Cache 15 upvotes, #8 of 2024-07-26
  18. Generative Multimodal Models are In-Context Learners 36 upvotes, #3 of 2023-12-21
  19. Sherpa3D: Boosting High-Fidelity Text-to-3D Generation via Coarse 3D Prior 24 upvotes, #3 of 2023-12-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.