Daily Papers of 2026-10-01
- The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation 521 upvotes, #1 of 2026-10-01
- False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents 469 upvotes, #2 of 2026-10-01
- LoopVL: Recurrent Visual Intelligence 460 upvotes, #3 of 2026-10-01
- UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement 280 upvotes, #4 of 2026-10-01
- Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence 210 upvotes, #5 of 2026-10-01
- AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks 133 upvotes, #6 of 2026-10-01
- Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents 113 upvotes, #7 of 2026-10-01
- EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery 107 upvotes, #8 of 2026-10-01
- WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents 99 upvotes, #9 of 2026-10-01
- RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement 84 upvotes, #10 of 2026-10-01
- Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI 82 upvotes, #11 of 2026-10-01
- EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making 73 upvotes, #12 of 2026-10-01
- Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? 65 upvotes, #13 of 2026-10-01
- Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering 62 upvotes, #14 of 2026-10-01
- OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software 59 upvotes, #15 of 2026-10-01
- More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models 57 upvotes, #16 of 2026-10-01
- AIM: Agentic Idea Management for Automated Research 50 upvotes, #17 of 2026-10-01
- Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training 50 upvotes, #17 of 2026-10-01
- LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models 47 upvotes, #19 of 2026-10-01
- Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models 47 upvotes, #19 of 2026-10-01
- DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence 40 upvotes, #21 of 2026-10-01
- DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes 39 upvotes, #22 of 2026-10-01
- It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them 38 upvotes, #23 of 2026-10-01
- TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion 35 upvotes, #24 of 2026-10-01
- Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation 35 upvotes, #24 of 2026-10-01
- ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing 32 upvotes, #26 of 2026-10-01
- Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies 30 upvotes, #27 of 2026-10-01
- PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation 28 upvotes, #28 of 2026-10-01
- BiasReducer: Adaptive Bias Mitigation for Reward Models 24 upvotes, #29 of 2026-10-01
- PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents 21 upvotes, #30 of 2026-10-01
- Scaling Laws for Looped Mixture of Experts 21 upvotes, #30 of 2026-10-01
- The Low-Rank Structure of VLA Reinforcement Learning 19 upvotes, #32 of 2026-10-01
- CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering 18 upvotes, #33 of 2026-10-01
- DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation 17 upvotes, #34 of 2026-10-01
- Rubric Rewards from Item Response Theory 16 upvotes, #35 of 2026-10-01
- MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution 16 upvotes, #35 of 2026-10-01
- Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model 15 upvotes, #37 of 2026-10-01
- WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms 14 upvotes, #38 of 2026-10-01
- A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications? 14 upvotes, #38 of 2026-10-01
- RoboCoach: World Models as Active Coaches for Compositional Robot Skills 14 upvotes, #38 of 2026-10-01
- I Have a Stream: Making Self-Supervised Learning Work on Continuous Video 14 upvotes, #38 of 2026-10-01
- Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning 12 upvotes, #42 of 2026-10-01
- The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends 12 upvotes, #42 of 2026-10-01
- Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior 12 upvotes, #42 of 2026-10-01
- SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing 11 upvotes, #45 of 2026-10-01
- AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation 10 upvotes, #46 of 2026-10-01
- CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding 10 upvotes, #46 of 2026-10-01
- PatchHolmes: Agentic Patch Retrieval via Listwise Selection 9 upvotes, #48 of 2026-10-01
- SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation 7 upvotes, #49 of 2026-10-01
- Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning 7 upvotes, #49 of 2026-10-01
- Safety of Latent Communication in Multi-Agent Systems 7 upvotes, #49 of 2026-10-01
- SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale 6 upvotes, #52 of 2026-10-01
- Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD 5 upvotes, #53 of 2026-10-01
- NavHarness: Towards Lifelong Embodied Navigation 5 upvotes, #53 of 2026-10-01
- Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering 5 upvotes, #53 of 2026-10-01
- Training LLM Judges from Language Feedback via Position-Selective Self-Distillation 5 upvotes, #53 of 2026-10-01
- Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds 5 upvotes, #53 of 2026-10-01
- Game-Guided Skill Discovery through Self-Play for Playable Agent Control 5 upvotes, #53 of 2026-10-01
- Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text 5 upvotes, #53 of 2026-10-01
- Understanding Multimodality in Generative Behavioral Cloning 4 upvotes, #60 of 2026-10-01
- See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs 4 upvotes, #60 of 2026-10-01
- CheatBench: Measuring Reward Gaming in AI Agents 4 upvotes, #60 of 2026-10-01
- Mitigating the Length-Scaling Tax with Online Distillation 4 upvotes, #60 of 2026-10-01
- DAGent: Evaluate-then-Grow Planning for Deep Research Agents 4 upvotes, #60 of 2026-10-01
- How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text 4 upvotes, #60 of 2026-10-01
- Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces 3 upvotes, #66 of 2026-10-01
- The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence 3 upvotes, #66 of 2026-10-01
- BIABench: Evaluating AI agents on real-world bioimage analysis tasks 3 upvotes, #66 of 2026-10-01
- Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation 3 upvotes, #66 of 2026-10-01
- Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance 3 upvotes, #66 of 2026-10-01
- SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models 3 upvotes, #66 of 2026-10-01
- Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost 3 upvotes, #66 of 2026-10-01
- EviRover: Reinforcing Agentic Perception Beyond a Glance 3 upvotes, #66 of 2026-10-01
- Decision-Oriented Recommendation Reranking: An Empirical Study of Jev 3 upvotes, #66 of 2026-10-01
- Aligning One-Step Generative Models with Reward-Weighted Transport Distillation 2 upvotes, #75 of 2026-10-01
- SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs 2 upvotes, #75 of 2026-10-01
- Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change 2 upvotes, #75 of 2026-10-01
- Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems 2 upvotes, #75 of 2026-10-01
- ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning 2 upvotes, #75 of 2026-10-01
- The Geometry of Inference in Transformer Residual Streams 2 upvotes, #75 of 2026-10-01
- Retrieval Capacity of Self-Attention Under Competition 2 upvotes, #75 of 2026-10-01
- Soft Spatial Reasoning 2 upvotes, #75 of 2026-10-01
- Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers 2 upvotes, #75 of 2026-10-01
- Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings 2 upvotes, #75 of 2026-10-01
- LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception 2 upvotes, #75 of 2026-10-01
- MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories 2 upvotes, #75 of 2026-10-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.