Dahua Lin

Dahua Lin on Hugging Face Daily Papers: 64 papers, 27 in the top 3 of their day, 2,622 upvotes.

  1. Demystifing Video Reasoning 356 upvotes, #1 of 2026-03-18
  2. Visual-ERM: Reward Modeling for Visual Equivalence 21 upvotes, #8 of 2026-03-16
  3. From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space 14 upvotes, #12 of 2026-03-16
  4. Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving 44 upvotes, #2 of 2025-12-12
  5. Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control 24 upvotes, #11 of 2025-06-03
  6. Visual Agentic Reinforcement Fine-Tuning 31 upvotes, #6 of 2025-05-21
  7. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
  8. MM-IFEngine: Towards Multimodal Instruction Following 31 upvotes, #6 of 2025-04-11
  9. GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography 21 upvotes, #6 of 2025-04-10
  10. HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance 11 upvotes, #10 of 2025-04-09
  11. LEGION: Learning to Ground and Explain for Synthetic Image Detection 19 upvotes, #8 of 2025-03-20
  12. Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM 41 upvotes, #4 of 2025-03-19
  13. Long Context Tuning for Video Generation 13 upvotes, #21 of 2025-03-14
  14. Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs 18 upvotes, #5 of 2025-03-05
  15. Visual-RFT: Visual Reinforcement Fine-Tuning 62 upvotes, #2 of 2025-03-04
  16. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation 35 upvotes, #4 of 2025-02-20
  17. Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 55 upvotes, #3 of 2025-02-11
  18. VideoRoPE: What Makes for Good Video Rotary Position Embedding? 60 upvotes, #3 of 2025-02-10
  19. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 39 upvotes, #6 of 2025-01-22
  20. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction 32 upvotes, #4 of 2025-01-07
  21. BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 34 upvotes, #3 of 2025-01-07
  22. IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations 12 upvotes, #11 of 2024-12-17
  23. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
  24. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models 20 upvotes, #7 of 2024-12-11
  25. 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation 18 upvotes, #9 of 2024-12-11
  26. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
  27. Imagine360: Immersive 360 Video Generation from Perspective Anchor 26 upvotes, #4 of 2024-12-05
  28. X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models 61 upvotes, #1 of 2024-12-03
  29. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models 28 upvotes, #2 of 2024-11-21
  30. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models 34 upvotes, #1 of 2024-10-24
  31. PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction 42 upvotes, #1 of 2024-10-23
  32. SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree 61 upvotes, #2 of 2024-10-22
  33. LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models 53 upvotes, #1 of 2024-10-15
  34. Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate 36 upvotes, #7 of 2024-10-10
  35. BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way 10 upvotes, #24 of 2024-10-10
  36. 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion 17 upvotes, #8 of 2024-09-20
  37. HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation 21 upvotes, #3 of 2024-07-25
  38. VEnhancer: Generative Space-Time Enhancement for Video Generation 8 upvotes, #7 of 2024-07-11
  39. Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images 8 upvotes, #8 of 2024-07-09
  40. ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models 1 upvotes, #17 of 2024-07-09
  41. InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
  42. MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding 27 upvotes, #7 of 2024-06-21
  43. Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
  44. OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI 14 upvotes, #10 of 2024-06-19
  45. MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs 52 upvotes, #1 of 2024-06-18
  46. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions 61 upvotes, #1 of 2024-06-07
  47. Grounded 3D-LLM with Referent Tokens 7 upvotes, #4 of 2024-05-20
  48. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
  49. Learning H-Infinity Locomotion Control 6 upvotes, #10 of 2024-04-23
  50. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
  51. InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
  52. Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models 11 upvotes, #6 of 2024-03-20
  53. InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning 19 upvotes, #2 of 2024-02-12
  54. InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
  55. GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation 21 upvotes, #5 of 2024-01-09
  56. HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image 22 upvotes, #3 of 2023-12-08
  57. Alpha-CLIP: A CLIP Model Focusing on Wherever You Want 33 upvotes, #3 of 2023-12-07
  58. OneLLM: One Framework to Align All Modalities with Language 23 upvotes, #6 of 2023-12-06
  59. GPT4Point: A Unified Framework for Point-Language Understanding and Generation 9 upvotes, #18 of 2023-12-06
  60. Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering 11 upvotes, #13 of 2023-12-04
  61. VR-NeRF: High-Fidelity Virtualized Walkable Spaces 16 upvotes, #7 of 2023-11-07
  62. HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion 16 upvotes, #6 of 2023-10-13
  63. LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models 43 upvotes, #2 of 2023-09-27
  64. DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering 6 upvotes, #9 of 2023-07-20

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.