Daily Papers of 2026-03-25
- MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding 131 upvotes, #1 of 2026-03-25
- WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG 90 upvotes, #2 of 2026-03-25
- SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning 59 upvotes, #3 of 2026-03-25
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents 54 upvotes, #4 of 2026-03-25
- DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models 50 upvotes, #5 of 2026-03-25
- PEARL: Personalized Streaming Video Understanding Model 40 upvotes, #6 of 2026-03-25
- SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM 40 upvotes, #6 of 2026-03-25
- UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation 35 upvotes, #8 of 2026-03-25
- RealMaster: Lifting Rendered Scenes into Photorealistic Video 31 upvotes, #9 of 2026-03-25
- 2Xplat: Two Experts Are Better Than One Generalist 25 upvotes, #10 of 2026-03-25
- Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought 25 upvotes, #10 of 2026-03-25
- Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing 21 upvotes, #12 of 2026-03-25
- ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model 15 upvotes, #13 of 2026-03-25
- VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models 11 upvotes, #14 of 2026-03-25
- CanViT: Toward Active-Vision Foundation Models 11 upvotes, #14 of 2026-03-25
- AgentSLR: Automating Systematic Literature Reviews in Epidemiology with Agentic AI 10 upvotes, #16 of 2026-03-25
- Abstraction as a Memory-Efficient Inductive Bias for Continual Learning 8 upvotes, #17 of 2026-03-25
- MultiBind: A Benchmark for Attribute Misbinding in Multi-Subject Generation 7 upvotes, #18 of 2026-03-25
- Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs 7 upvotes, #18 of 2026-03-25
- Fair splits flip the leaderboard: CHANRG reveals limited generalization in RNA secondary-structure prediction 6 upvotes, #20 of 2026-03-25
- Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos 6 upvotes, #20 of 2026-03-25
- VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs 6 upvotes, #20 of 2026-03-25
- TrajLoom: Dense Future Trajectory Generation from Video 5 upvotes, #23 of 2026-03-25
- Regulating AI Agents 5 upvotes, #23 of 2026-03-25
- One View Is Enough! Monocular Training for In-the-Wild Novel View Generation 4 upvotes, #25 of 2026-03-25
- Can AI Agents Answer Your Data Questions? A Benchmark for Data Agents 3 upvotes, #26 of 2026-03-25
- Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models 3 upvotes, #26 of 2026-03-25
- Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models 3 upvotes, #26 of 2026-03-25
- STEM Agent: A Self-Adapting, Tool-Enabled, Extensible Architecture for Multi-Protocol AI Agent Systems 3 upvotes, #26 of 2026-03-25
- Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning 3 upvotes, #26 of 2026-03-25
- ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment 3 upvotes, #26 of 2026-03-25
- VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions 3 upvotes, #26 of 2026-03-25
- SHAMISA: SHAped Modeling of Implicit Structural Associations for Self-supervised No-Reference Image Quality Assessment 1 upvotes, #33 of 2026-03-25
- Session Risk Memory (SRM): Temporal Authorization for Deterministic Pre-Execution Safety Gates 1 upvotes, #33 of 2026-03-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.