Daily Papers of 2026-05-25
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills 212 upvotes, #1 of 2026-05-25
- Rethinking Cross-Layer Information Routing in Diffusion Transformers 109 upvotes, #2 of 2026-05-25
- Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models 107 upvotes, #3 of 2026-05-25
- SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research 58 upvotes, #4 of 2026-05-25
- StepAudio 2.5 Technical Report 49 upvotes, #5 of 2026-05-25
- PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion 45 upvotes, #6 of 2026-05-25
- See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding 33 upvotes, #7 of 2026-05-25
- Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration 31 upvotes, #8 of 2026-05-25
- From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills 29 upvotes, #9 of 2026-05-25
- PhotoFlow: Agentic 3D Virtual Photography Missions 26 upvotes, #10 of 2026-05-25
- VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis 24 upvotes, #11 of 2026-05-25
- Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback 19 upvotes, #12 of 2026-05-25
- RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution 18 upvotes, #13 of 2026-05-25
- SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models 17 upvotes, #14 of 2026-05-25
- GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction 13 upvotes, #15 of 2026-05-25
- ETCHR: Editing To Clarify and Harness Reasoning 13 upvotes, #15 of 2026-05-25
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws 13 upvotes, #15 of 2026-05-25
- HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents 11 upvotes, #18 of 2026-05-25
- From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models 10 upvotes, #19 of 2026-05-25
- Geo-Align: Video Generation Alignment via Metric Geometry Reward 10 upvotes, #19 of 2026-05-25
- LatentUMM: Dual Latent Alignment for Unified Multimodal Models 8 upvotes, #21 of 2026-05-25
- Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR 8 upvotes, #21 of 2026-05-25
- The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation 8 upvotes, #21 of 2026-05-25
- Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers 8 upvotes, #21 of 2026-05-25
- The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm 7 upvotes, #25 of 2026-05-25
- Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning 6 upvotes, #26 of 2026-05-25
- Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs 5 upvotes, #27 of 2026-05-25
- VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation 4 upvotes, #28 of 2026-05-25
- Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction 3 upvotes, #29 of 2026-05-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.