Kevin Lin
Kevin Lin on Hugging Face Daily Papers: 30 papers, 14 in the top 3 of their day, 2,245 upvotes.
- Show-Harness: Just a VLM Agent Can Play Robots 160 upvotes, #2 of 2026-09-10
- Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories 110 upvotes, #3 of 2026-06-10
- Agents' Last Exam 345 upvotes, #1 of 2026-06-09
- AI for Auto-Research: Roadmap & User Guide 65 upvotes, #5 of 2026-05-19
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond 181 upvotes, #1 of 2026-04-27
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents 108 upvotes, #6 of 2026-04-10
- CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents 92 upvotes, #1 of 2026-03-26
- Code2World: A GUI World Model via Renderable Code Generation 189 upvotes, #2 of 2026-02-11
- FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection 16 upvotes, #12 of 2026-01-15
- ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands 40 upvotes, #8 of 2026-01-14
- Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans? 63 upvotes, #2 of 2025-12-17
- Computer-Use Agents as Judges for Generative User Interface 50 upvotes, #5 of 2025-11-25
- Grounding Computer Use Agents on Human Demonstrations 98 upvotes, #1 of 2025-11-11
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation 97 upvotes, #1 of 2025-11-05
- Paper2Video: Automatic Video Generation from Scientific Papers 99 upvotes, #1 of 2025-10-07
- Code2Video: A Code-centric Paradigm for Educational Video Generation 29 upvotes, #7 of 2025-10-02
- Reinforcement Learning in Vision: A Survey 27 upvotes, #11 of 2025-08-12
- Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers 91 upvotes, #2 of 2025-05-28
- Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models 11 upvotes, #25 of 2025-05-23
- VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning 13 upvotes, #13 of 2025-03-18
- VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary 6 upvotes, #13 of 2025-03-13
- ROICtrl: Boosting Instance Control for Visual Generation 77 upvotes, #1 of 2024-11-28
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent 68 upvotes, #1 of 2024-11-27
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation 47 upvotes, #3 of 2024-08-23
- VideoLLM-online: Online Video Large Language Model for Streaming Video 19 upvotes, #8 of 2024-06-18
- VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
- COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training 17 upvotes, #4 of 2024-01-02
- UniVTG: Towards Unified Video-Language Temporal Grounding 12 upvotes, #10 of 2023-08-01
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone 12 upvotes, #8 of 2023-07-12
- AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn 27 upvotes, #4 of 2023-06-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.