Dahua Lin
Dahua Lin on Hugging Face Daily Papers: 64 papers, 27 in the top 3 of their day, 2,622 upvotes.
- Demystifing Video Reasoning 356 upvotes, #1 of 2026-03-18
- Visual-ERM: Reward Modeling for Visual Equivalence 21 upvotes, #8 of 2026-03-16
- From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space 14 upvotes, #12 of 2026-03-16
- Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving 44 upvotes, #2 of 2025-12-12
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control 24 upvotes, #11 of 2025-06-03
- Visual Agentic Reinforcement Fine-Tuning 31 upvotes, #6 of 2025-05-21
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
- MM-IFEngine: Towards Multimodal Instruction Following 31 upvotes, #6 of 2025-04-11
- GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography 21 upvotes, #6 of 2025-04-10
- HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance 11 upvotes, #10 of 2025-04-09
- LEGION: Learning to Ground and Explain for Synthetic Image Detection 19 upvotes, #8 of 2025-03-20
- Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM 41 upvotes, #4 of 2025-03-19
- Long Context Tuning for Video Generation 13 upvotes, #21 of 2025-03-14
- Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs 18 upvotes, #5 of 2025-03-05
- Visual-RFT: Visual Reinforcement Fine-Tuning 62 upvotes, #2 of 2025-03-04
- SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation 35 upvotes, #4 of 2025-02-20
- Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 55 upvotes, #3 of 2025-02-11
- VideoRoPE: What Makes for Good Video Rotary Position Embedding? 60 upvotes, #3 of 2025-02-10
- InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 39 upvotes, #6 of 2025-01-22
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction 32 upvotes, #4 of 2025-01-07
- BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 34 upvotes, #3 of 2025-01-07
- IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations 12 upvotes, #11 of 2024-12-17
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
- FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models 20 upvotes, #7 of 2024-12-11
- 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation 18 upvotes, #9 of 2024-12-11
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
- Imagine360: Immersive 360 Video Generation from Perspective Anchor 26 upvotes, #4 of 2024-12-05
- X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models 61 upvotes, #1 of 2024-12-03
- VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models 28 upvotes, #2 of 2024-11-21
- MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models 34 upvotes, #1 of 2024-10-24
- PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction 42 upvotes, #1 of 2024-10-23
- SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree 61 upvotes, #2 of 2024-10-22
- LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models 53 upvotes, #1 of 2024-10-15
- Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate 36 upvotes, #7 of 2024-10-10
- BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way 10 upvotes, #24 of 2024-10-10
- 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion 17 upvotes, #8 of 2024-09-20
- HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation 21 upvotes, #3 of 2024-07-25
- VEnhancer: Generative Space-Time Enhancement for Video Generation 8 upvotes, #7 of 2024-07-11
- Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images 8 upvotes, #8 of 2024-07-09
- ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models 1 upvotes, #17 of 2024-07-09
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
- MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding 27 upvotes, #7 of 2024-06-21
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
- OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI 14 upvotes, #10 of 2024-06-19
- MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs 52 upvotes, #1 of 2024-06-18
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions 61 upvotes, #1 of 2024-06-07
- Grounded 3D-LLM with Referent Tokens 7 upvotes, #4 of 2024-05-20
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
- Learning H-Infinity Locomotion Control 6 upvotes, #10 of 2024-04-23
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
- InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
- Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models 11 upvotes, #6 of 2024-03-20
- InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning 19 upvotes, #2 of 2024-02-12
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
- GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation 21 upvotes, #5 of 2024-01-09
- HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image 22 upvotes, #3 of 2023-12-08
- Alpha-CLIP: A CLIP Model Focusing on Wherever You Want 33 upvotes, #3 of 2023-12-07
- OneLLM: One Framework to Align All Modalities with Language 23 upvotes, #6 of 2023-12-06
- GPT4Point: A Unified Framework for Point-Language Understanding and Generation 9 upvotes, #18 of 2023-12-06
- Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering 11 upvotes, #13 of 2023-12-04
- VR-NeRF: High-Fidelity Virtualized Walkable Spaces 16 upvotes, #7 of 2023-11-07
- HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion 16 upvotes, #6 of 2023-10-13
- LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models 43 upvotes, #2 of 2023-09-27
- DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering 6 upvotes, #9 of 2023-07-20
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.