Daily Papers of 2025-03-26

  1. Long-Context Autoregressive Video Modeling with Next-Frame Prediction 70 upvotes, #1 of 2025-03-26
  2. Scaling Vision Pre-Training to 4K Resolution 37 upvotes, #2 of 2025-03-26
  3. Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing 32 upvotes, #3 of 2025-03-26
  4. CoMP: Continual Multimodal Pre-training for Vision Foundation Models 29 upvotes, #4 of 2025-03-26
  5. Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation 29 upvotes, #4 of 2025-03-26
  6. Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking 24 upvotes, #6 of 2025-03-26
  7. Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation 19 upvotes, #7 of 2025-03-26
  8. MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding 16 upvotes, #8 of 2025-03-26
  9. ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning 14 upvotes, #9 of 2025-03-26
  10. CoLLM: A Large Language Model for Composed Image Retrieval 11 upvotes, #10 of 2025-03-26
  11. Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models 9 upvotes, #11 of 2025-03-26
  12. WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation 9 upvotes, #11 of 2025-03-26
  13. DiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesis 8 upvotes, #13 of 2025-03-26
  14. FullDiT: Multi-Task Video Generative Foundation Model with Full Attention 8 upvotes, #13 of 2025-03-26
  15. FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement 7 upvotes, #15 of 2025-03-26
  16. PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos 7 upvotes, #15 of 2025-03-26
  17. Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation 6 upvotes, #17 of 2025-03-26
  18. LookAhead Tuning: Safer Language Models via Partial Answer Previews 5 upvotes, #18 of 2025-03-26
  19. When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only Training For Human-Centered Decision Making 4 upvotes, #19 of 2025-03-26
  20. Strong Baseline: Multi-UAV Tracking via YOLOv12 with BoT-SORT-ReID 4 upvotes, #19 of 2025-03-26
  21. Gumbel-Softmax Flow Matching with Straight-Through Guidance for Controllable Biological Sequence Generation 4 upvotes, #19 of 2025-03-26
  22. xKV: Cross-Layer SVD for KV-Cache Compression 4 upvotes, #19 of 2025-03-26
  23. FRESA:Feedforward Reconstruction of Personalized Skinned Avatars from Few Images 4 upvotes, #19 of 2025-03-26
  24. Efficient Model Development through Fine-tuning Transfer 4 upvotes, #19 of 2025-03-26
  25. Towards a Unified Copernicus Foundation Model for Earth Vision 3 upvotes, #25 of 2025-03-26
  26. OpenCity3D: What do Vision-Language Models know about Urban Environments? 3 upvotes, #25 of 2025-03-26
  27. LLaVAction: evaluating and training multi-modal large language models for action recognition 3 upvotes, #25 of 2025-03-26
  28. Any6D: Model-free 6D Pose Estimation of Novel Objects 2 upvotes, #28 of 2025-03-26
  29. Frequency Dynamic Convolution for Dense Image Prediction 2 upvotes, #28 of 2025-03-26
  30. Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling 2 upvotes, #28 of 2025-03-26
  31. Can Vision-Language Models Answer Face to Face Questions in the Real-World? 2 upvotes, #28 of 2025-03-26
  32. ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models 1 upvotes, #32 of 2025-03-26
  33. LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentation 1 upvotes, #32 of 2025-03-26
  34. Co-SemDepth: Fast Joint Semantic Segmentation and Depth Estimation on Aerial Images 0 upvotes, #34 of 2025-03-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.