Lei Li
Lei Li on Hugging Face Daily Papers: 21 papers, 7 in the top 3 of their day, 690 upvotes.
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining 144 upvotes, #1 of 2026-05-21
- AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval 3 upvotes, #16 of 2026-05-04
- Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows 40 upvotes, #6 of 2026-05-01
- Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents 114 upvotes, #2 of 2026-04-08
- TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions 27 upvotes, #8 of 2026-02-12
- MiMo-V2-Flash Technical Report 31 upvotes, #8 of 2026-01-07
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation 19 upvotes, #11 of 2025-12-22
- MiMo-VL Technical Report 70 upvotes, #1 of 2025-06-05
- MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining 74 upvotes, #2 of 2025-05-13
- TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos 10 upvotes, #12 of 2025-04-25
- LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding 6 upvotes, #20 of 2025-03-10
- Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey 49 upvotes, #3 of 2024-12-30
- VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models 10 upvotes, #8 of 2024-11-27
- Why Does the Effective Context Length of LLMs Fall Short? 15 upvotes, #9 of 2024-10-25
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 26 upvotes, #6 of 2024-10-15
- Temporal Reasoning Transfer from Text to Video 12 upvotes, #19 of 2024-10-10
- Jailbreaking as a Reward Misspecification Problem 12 upvotes, #9 of 2024-06-24
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 13 upvotes, #3 of 2024-06-03
- Weak-to-Strong Jailbreaking on Large Language Models 16 upvotes, #8 of 2024-01-31
- Silkie: Preference Distillation for Large Visual Language Models 10 upvotes, #9 of 2023-12-19
- M^3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning 10 upvotes, #1 of 2023-06-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.