Daily Papers of 2025-03-26
- Long-Context Autoregressive Video Modeling with Next-Frame Prediction 70 upvotes, #1 of 2025-03-26
- Scaling Vision Pre-Training to 4K Resolution 37 upvotes, #2 of 2025-03-26
- Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing 32 upvotes, #3 of 2025-03-26
- CoMP: Continual Multimodal Pre-training for Vision Foundation Models 29 upvotes, #4 of 2025-03-26
- Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation 29 upvotes, #4 of 2025-03-26
- Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking 24 upvotes, #6 of 2025-03-26
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation 19 upvotes, #7 of 2025-03-26
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding 16 upvotes, #8 of 2025-03-26
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning 14 upvotes, #9 of 2025-03-26
- CoLLM: A Large Language Model for Composed Image Retrieval 11 upvotes, #10 of 2025-03-26
- Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models 9 upvotes, #11 of 2025-03-26
- WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation 9 upvotes, #11 of 2025-03-26
- DiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesis 8 upvotes, #13 of 2025-03-26
- FullDiT: Multi-Task Video Generative Foundation Model with Full Attention 8 upvotes, #13 of 2025-03-26
- FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement 7 upvotes, #15 of 2025-03-26
- PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos 7 upvotes, #15 of 2025-03-26
- Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation 6 upvotes, #17 of 2025-03-26
- LookAhead Tuning: Safer Language Models via Partial Answer Previews 5 upvotes, #18 of 2025-03-26
- When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only Training For Human-Centered Decision Making 4 upvotes, #19 of 2025-03-26
- Strong Baseline: Multi-UAV Tracking via YOLOv12 with BoT-SORT-ReID 4 upvotes, #19 of 2025-03-26
- Gumbel-Softmax Flow Matching with Straight-Through Guidance for Controllable Biological Sequence Generation 4 upvotes, #19 of 2025-03-26
- xKV: Cross-Layer SVD for KV-Cache Compression 4 upvotes, #19 of 2025-03-26
- FRESA:Feedforward Reconstruction of Personalized Skinned Avatars from Few Images 4 upvotes, #19 of 2025-03-26
- Efficient Model Development through Fine-tuning Transfer 4 upvotes, #19 of 2025-03-26
- Towards a Unified Copernicus Foundation Model for Earth Vision 3 upvotes, #25 of 2025-03-26
- OpenCity3D: What do Vision-Language Models know about Urban Environments? 3 upvotes, #25 of 2025-03-26
- LLaVAction: evaluating and training multi-modal large language models for action recognition 3 upvotes, #25 of 2025-03-26
- Any6D: Model-free 6D Pose Estimation of Novel Objects 2 upvotes, #28 of 2025-03-26
- Frequency Dynamic Convolution for Dense Image Prediction 2 upvotes, #28 of 2025-03-26
- Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling 2 upvotes, #28 of 2025-03-26
- Can Vision-Language Models Answer Face to Face Questions in the Real-World? 2 upvotes, #28 of 2025-03-26
- ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models 1 upvotes, #32 of 2025-03-26
- LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentation 1 upvotes, #32 of 2025-03-26
- Co-SemDepth: Fast Joint Semantic Segmentation and Depth Estimation on Aerial Images 0 upvotes, #34 of 2025-03-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.