Daily Papers of 2026-05-25

  1. SkillOpt: Executive Strategy for Self-Evolving Agent Skills 212 upvotes, #1 of 2026-05-25
  2. Rethinking Cross-Layer Information Routing in Diffusion Transformers 109 upvotes, #2 of 2026-05-25
  3. Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models 107 upvotes, #3 of 2026-05-25
  4. SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research 58 upvotes, #4 of 2026-05-25
  5. StepAudio 2.5 Technical Report 49 upvotes, #5 of 2026-05-25
  6. PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion 45 upvotes, #6 of 2026-05-25
  7. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding 33 upvotes, #7 of 2026-05-25
  8. Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration 31 upvotes, #8 of 2026-05-25
  9. From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills 29 upvotes, #9 of 2026-05-25
  10. PhotoFlow: Agentic 3D Virtual Photography Missions 26 upvotes, #10 of 2026-05-25
  11. VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis 24 upvotes, #11 of 2026-05-25
  12. Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback 19 upvotes, #12 of 2026-05-25
  13. RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution 18 upvotes, #13 of 2026-05-25
  14. SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models 17 upvotes, #14 of 2026-05-25
  15. GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction 13 upvotes, #15 of 2026-05-25
  16. ETCHR: Editing To Clarify and Harness Reasoning 13 upvotes, #15 of 2026-05-25
  17. LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws 13 upvotes, #15 of 2026-05-25
  18. HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents 11 upvotes, #18 of 2026-05-25
  19. From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models 10 upvotes, #19 of 2026-05-25
  20. Geo-Align: Video Generation Alignment via Metric Geometry Reward 10 upvotes, #19 of 2026-05-25
  21. LatentUMM: Dual Latent Alignment for Unified Multimodal Models 8 upvotes, #21 of 2026-05-25
  22. Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR 8 upvotes, #21 of 2026-05-25
  23. The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation 8 upvotes, #21 of 2026-05-25
  24. Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers 8 upvotes, #21 of 2026-05-25
  25. The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm 7 upvotes, #25 of 2026-05-25
  26. Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning 6 upvotes, #26 of 2026-05-25
  27. Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs 5 upvotes, #27 of 2026-05-25
  28. VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation 4 upvotes, #28 of 2026-05-25
  29. Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction 3 upvotes, #29 of 2026-05-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.