Lin Chen

Lin Chen on Hugging Face Daily Papers: 23 papers, 7 in the top 3 of their day, 1,225 upvotes.

  1. Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 50 upvotes, #5 of 2026-08-05
  2. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System 70 upvotes, #6 of 2026-07-31
  3. AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios 16 upvotes, #22 of 2026-05-29
  4. ACC: Compiling Agent Trajectories for Long-Context Training 59 upvotes, #6 of 2026-05-22
  5. SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering 7 upvotes, #28 of 2026-05-21
  6. VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation 27 upvotes, #12 of 2026-05-19
  7. SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation 10 upvotes, #25 of 2026-05-11
  8. Flow-OPD: On-Policy Distillation for Flow Matching Models 95 upvotes, #2 of 2026-05-11
  9. SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents 22 upvotes, #10 of 2026-04-21
  10. Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models 149 upvotes, #3 of 2026-02-03
  11. Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models 124 upvotes, #4 of 2026-02-03
  12. UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision 44 upvotes, #4 of 2026-01-07
  13. DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action 21 upvotes, #12 of 2025-12-01
  14. Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models 9 upvotes, #24 of 2025-10-03
  15. VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning 10 upvotes, #29 of 2025-05-29
  16. Seed1.5-VL Technical Report 136 upvotes, #1 of 2025-05-13
  17. VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning 43 upvotes, #4 of 2025-04-11
  18. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
  19. Open-Sora Plan: Open-Source Large Video Generation Model 30 upvotes, #3 of 2024-12-03
  20. VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models 11 upvotes, #6 of 2024-07-17
  21. InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
  22. Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
  23. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions 61 upvotes, #1 of 2024-06-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.