Lewei Lu
Lewei Lu on Hugging Face Daily Papers: 30 papers, 9 in the top 3 of their day, 1,646 upvotes.
- From Pixels to Words -- Towards Native One-Vision Models at Scale 72 upvotes, #4 of 2026-05-28
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 184 upvotes, #1 of 2026-05-13
- OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis 27 upvotes, #5 of 2026-04-23
- EVA: Efficient Reinforcement Learning for End-to-End Video Agent 42 upvotes, #4 of 2026-03-26
- SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning 33 upvotes, #6 of 2026-01-05
- Scaling Spatial Intelligence with Multimodal Foundation Models 41 upvotes, #6 of 2025-11-21
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale 64 upvotes, #5 of 2025-10-17
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue 28 upvotes, #9 of 2025-10-16
- CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving 24 upvotes, #12 of 2025-10-16
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints 19 upvotes, #22 of 2025-10-10
- Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding 53 upvotes, #7 of 2025-09-03
- Has GPT-5 Achieved Spatial Intelligence? An Empirical Study 31 upvotes, #8 of 2025-08-19
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior 7 upvotes, #23 of 2025-06-10
- Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 32 upvotes, #10 of 2025-06-04
- Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM 3 upvotes, #38 of 2025-05-22
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models 71 upvotes, #1 of 2025-04-24
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
- MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction 36 upvotes, #4 of 2025-02-24
- Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding 7 upvotes, #8 of 2025-01-16
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 35 upvotes, #3 of 2024-12-16
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
- Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization 61 upvotes, #1 of 2024-11-22
- Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
- OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text 28 upvotes, #6 of 2024-06-17
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks 22 upvotes, #2 of 2023-12-26
- ControlLLM: Augment Language Models with Tools by Searching on Graphs 18 upvotes, #3 of 2023-10-30
- Ghost in the Minecraft: Generally Capable Agents for Open-World Enviroments via Large Language Models with Text-based Knowledge and Memory 4 upvotes, #8 of 2023-05-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.