Lewei Lu

Lewei Lu on Hugging Face Daily Papers: 30 papers, 9 in the top 3 of their day, 1,646 upvotes.

  1. From Pixels to Words -- Towards Native One-Vision Models at Scale 72 upvotes, #4 of 2026-05-28
  2. SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 184 upvotes, #1 of 2026-05-13
  3. OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis 27 upvotes, #5 of 2026-04-23
  4. EVA: Efficient Reinforcement Learning for End-to-End Video Agent 42 upvotes, #4 of 2026-03-26
  5. SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning 33 upvotes, #6 of 2026-01-05
  6. Scaling Spatial Intelligence with Multimodal Foundation Models 41 upvotes, #6 of 2025-11-21
  7. From Pixels to Words -- Towards Native Vision-Language Primitives at Scale 64 upvotes, #5 of 2025-10-17
  8. InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue 28 upvotes, #9 of 2025-10-16
  9. CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving 24 upvotes, #12 of 2025-10-16
  10. NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints 19 upvotes, #22 of 2025-10-10
  11. Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
  12. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding 53 upvotes, #7 of 2025-09-03
  13. Has GPT-5 Achieved Spatial Intelligence? An Empirical Study 31 upvotes, #8 of 2025-08-19
  14. GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior 7 upvotes, #23 of 2025-06-10
  15. Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 32 upvotes, #10 of 2025-06-04
  16. Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM 3 upvotes, #38 of 2025-05-22
  17. VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models 71 upvotes, #1 of 2025-04-24
  18. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
  19. VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
  20. MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction 36 upvotes, #4 of 2025-02-24
  21. Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding 7 upvotes, #8 of 2025-01-16
  22. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 35 upvotes, #3 of 2024-12-16
  23. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
  24. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization 61 upvotes, #1 of 2024-11-22
  25. Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
  26. OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text 28 upvotes, #6 of 2024-06-17
  27. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
  28. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks 22 upvotes, #2 of 2023-12-26
  29. ControlLLM: Augment Language Models with Tools by Searching on Graphs 18 upvotes, #3 of 2023-10-30
  30. Ghost in the Minecraft: Generally Capable Agents for Open-World Enviroments via Large Language Models with Text-based Knowledge and Memory 4 upvotes, #8 of 2023-05-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.