Lei Li

Lei Li on Hugging Face Daily Papers: 21 papers, 7 in the top 3 of their day, 690 upvotes.

  1. Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining 144 upvotes, #1 of 2026-05-21
  2. AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval 3 upvotes, #16 of 2026-05-04
  3. Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows 40 upvotes, #6 of 2026-05-01
  4. Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents 114 upvotes, #2 of 2026-04-08
  5. TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions 27 upvotes, #8 of 2026-02-12
  6. MiMo-V2-Flash Technical Report 31 upvotes, #8 of 2026-01-07
  7. GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation 19 upvotes, #11 of 2025-12-22
  8. MiMo-VL Technical Report 70 upvotes, #1 of 2025-06-05
  9. MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining 74 upvotes, #2 of 2025-05-13
  10. TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos 10 upvotes, #12 of 2025-04-25
  11. LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding 6 upvotes, #20 of 2025-03-10
  12. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey 49 upvotes, #3 of 2024-12-30
  13. VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models 10 upvotes, #8 of 2024-11-27
  14. Why Does the Effective Context Length of LLMs Fall Short? 15 upvotes, #9 of 2024-10-25
  15. Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 26 upvotes, #6 of 2024-10-15
  16. Temporal Reasoning Transfer from Text to Video 12 upvotes, #19 of 2024-10-10
  17. Jailbreaking as a Reward Misspecification Problem 12 upvotes, #9 of 2024-06-24
  18. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 13 upvotes, #3 of 2024-06-03
  19. Weak-to-Strong Jailbreaking on Large Language Models 16 upvotes, #8 of 2024-01-31
  20. Silkie: Preference Distillation for Large Visual Language Models 10 upvotes, #9 of 2023-12-19
  21. M^3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning 10 upvotes, #1 of 2023-06-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.