Daily Papers of 2025-03-18
- DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation 84 upvotes, #1 of 2025-03-18
- Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills 60 upvotes, #2 of 2025-03-18
- Personalize Anything for Free with Diffusion Transformer 41 upvotes, #3 of 2025-03-18
- DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models 41 upvotes, #3 of 2025-03-18
- SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially? 39 upvotes, #5 of 2025-03-18
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey 28 upvotes, #6 of 2025-03-18
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization 26 upvotes, #7 of 2025-03-18
- Edit Transfer: Learning Image Editing via Vision In-Context Relations 24 upvotes, #8 of 2025-03-18
- BlobCtrl: A Unified and Flexible Framework for Element-level Image Generation and Editing 24 upvotes, #8 of 2025-03-18
- MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research 20 upvotes, #10 of 2025-03-18
- WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes 16 upvotes, #11 of 2025-03-18
- reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs 15 upvotes, #12 of 2025-03-18
- VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning 13 upvotes, #13 of 2025-03-18
- V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning 11 upvotes, #14 of 2025-03-18
- MTV-Inpaint: Multi-Task Long Video Inpainting 10 upvotes, #15 of 2025-03-18
- Long-Video Audio Synthesis with Multi-Agent Collaboration 9 upvotes, #16 of 2025-03-18
- Rewards Are Enough for Fast Photo-Realistic Text-to-image Generation 9 upvotes, #16 of 2025-03-18
- Free-form language-based robotic reasoning and grasping 9 upvotes, #16 of 2025-03-18
- Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions 7 upvotes, #19 of 2025-03-18
- Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework 5 upvotes, #20 of 2025-03-18
- Training Video Foundation Models with NVIDIA NeMo 5 upvotes, #20 of 2025-03-18
- Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models 4 upvotes, #22 of 2025-03-18
- Investigating Human-Aligned Large Language Model Uncertainty 4 upvotes, #22 of 2025-03-18
- GenStereo: Towards Open-World Generation of Stereo Images and Unsupervised Matching 4 upvotes, #22 of 2025-03-18
- WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation 3 upvotes, #25 of 2025-03-18
- Basic Category Usage in Vision Language Models 3 upvotes, #25 of 2025-03-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.