Lin Chen
Lin Chen on Hugging Face Daily Papers: 23 papers, 7 in the top 3 of their day, 1,225 upvotes.
- Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 50 upvotes, #5 of 2026-08-05
- VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System 70 upvotes, #6 of 2026-07-31
- AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios 16 upvotes, #22 of 2026-05-29
- ACC: Compiling Agent Trajectories for Long-Context Training 59 upvotes, #6 of 2026-05-22
- SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering 7 upvotes, #28 of 2026-05-21
- VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation 27 upvotes, #12 of 2026-05-19
- SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation 10 upvotes, #25 of 2026-05-11
- Flow-OPD: On-Policy Distillation for Flow Matching Models 95 upvotes, #2 of 2026-05-11
- SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents 22 upvotes, #10 of 2026-04-21
- Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models 149 upvotes, #3 of 2026-02-03
- Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models 124 upvotes, #4 of 2026-02-03
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision 44 upvotes, #4 of 2026-01-07
- DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action 21 upvotes, #12 of 2025-12-01
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models 9 upvotes, #24 of 2025-10-03
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning 10 upvotes, #29 of 2025-05-29
- Seed1.5-VL Technical Report 136 upvotes, #1 of 2025-05-13
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning 43 upvotes, #4 of 2025-04-11
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
- Open-Sora Plan: Open-Source Large Video Generation Model 30 upvotes, #3 of 2024-12-03
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models 11 upvotes, #6 of 2024-07-17
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions 61 upvotes, #1 of 2024-06-07
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.