Daily Papers of 2026-03-25

  1. MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding 131 upvotes, #1 of 2026-03-25
  2. WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG 90 upvotes, #2 of 2026-03-25
  3. SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning 59 upvotes, #3 of 2026-03-25
  4. From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents 54 upvotes, #4 of 2026-03-25
  5. DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models 50 upvotes, #5 of 2026-03-25
  6. PEARL: Personalized Streaming Video Understanding Model 40 upvotes, #6 of 2026-03-25
  7. SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM 40 upvotes, #6 of 2026-03-25
  8. UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation 35 upvotes, #8 of 2026-03-25
  9. RealMaster: Lifting Rendered Scenes into Photorealistic Video 31 upvotes, #9 of 2026-03-25
  10. 2Xplat: Two Experts Are Better Than One Generalist 25 upvotes, #10 of 2026-03-25
  11. Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought 25 upvotes, #10 of 2026-03-25
  12. Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing 21 upvotes, #12 of 2026-03-25
  13. ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model 15 upvotes, #13 of 2026-03-25
  14. VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models 11 upvotes, #14 of 2026-03-25
  15. CanViT: Toward Active-Vision Foundation Models 11 upvotes, #14 of 2026-03-25
  16. AgentSLR: Automating Systematic Literature Reviews in Epidemiology with Agentic AI 10 upvotes, #16 of 2026-03-25
  17. Abstraction as a Memory-Efficient Inductive Bias for Continual Learning 8 upvotes, #17 of 2026-03-25
  18. MultiBind: A Benchmark for Attribute Misbinding in Multi-Subject Generation 7 upvotes, #18 of 2026-03-25
  19. Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs 7 upvotes, #18 of 2026-03-25
  20. Fair splits flip the leaderboard: CHANRG reveals limited generalization in RNA secondary-structure prediction 6 upvotes, #20 of 2026-03-25
  21. Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos 6 upvotes, #20 of 2026-03-25
  22. VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs 6 upvotes, #20 of 2026-03-25
  23. TrajLoom: Dense Future Trajectory Generation from Video 5 upvotes, #23 of 2026-03-25
  24. Regulating AI Agents 5 upvotes, #23 of 2026-03-25
  25. One View Is Enough! Monocular Training for In-the-Wild Novel View Generation 4 upvotes, #25 of 2026-03-25
  26. Can AI Agents Answer Your Data Questions? A Benchmark for Data Agents 3 upvotes, #26 of 2026-03-25
  27. Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models 3 upvotes, #26 of 2026-03-25
  28. Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models 3 upvotes, #26 of 2026-03-25
  29. STEM Agent: A Self-Adapting, Tool-Enabled, Extensible Architecture for Multi-Protocol AI Agent Systems 3 upvotes, #26 of 2026-03-25
  30. Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning 3 upvotes, #26 of 2026-03-25
  31. ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment 3 upvotes, #26 of 2026-03-25
  32. VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions 3 upvotes, #26 of 2026-03-25
  33. SHAMISA: SHAped Modeling of Implicit Structural Associations for Self-supervised No-Reference Image Quality Assessment 1 upvotes, #33 of 2026-03-25
  34. Session Risk Memory (SRM): Temporal Authorization for Deterministic Pre-Execution Safety Gates 1 upvotes, #33 of 2026-03-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.