Daily Papers of 2026-10-01

  1. The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation 521 upvotes, #1 of 2026-10-01
  2. False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents 469 upvotes, #2 of 2026-10-01
  3. LoopVL: Recurrent Visual Intelligence 460 upvotes, #3 of 2026-10-01
  4. UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement 280 upvotes, #4 of 2026-10-01
  5. Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence 210 upvotes, #5 of 2026-10-01
  6. AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks 133 upvotes, #6 of 2026-10-01
  7. Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents 113 upvotes, #7 of 2026-10-01
  8. EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery 107 upvotes, #8 of 2026-10-01
  9. WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents 99 upvotes, #9 of 2026-10-01
  10. RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement 84 upvotes, #10 of 2026-10-01
  11. Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI 82 upvotes, #11 of 2026-10-01
  12. EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making 73 upvotes, #12 of 2026-10-01
  13. Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? 65 upvotes, #13 of 2026-10-01
  14. Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering 62 upvotes, #14 of 2026-10-01
  15. OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software 59 upvotes, #15 of 2026-10-01
  16. More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models 57 upvotes, #16 of 2026-10-01
  17. AIM: Agentic Idea Management for Automated Research 50 upvotes, #17 of 2026-10-01
  18. Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training 50 upvotes, #17 of 2026-10-01
  19. LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models 47 upvotes, #19 of 2026-10-01
  20. Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models 47 upvotes, #19 of 2026-10-01
  21. DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence 40 upvotes, #21 of 2026-10-01
  22. DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes 39 upvotes, #22 of 2026-10-01
  23. It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them 38 upvotes, #23 of 2026-10-01
  24. TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion 35 upvotes, #24 of 2026-10-01
  25. Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation 35 upvotes, #24 of 2026-10-01
  26. ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing 32 upvotes, #26 of 2026-10-01
  27. Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies 30 upvotes, #27 of 2026-10-01
  28. PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation 28 upvotes, #28 of 2026-10-01
  29. BiasReducer: Adaptive Bias Mitigation for Reward Models 24 upvotes, #29 of 2026-10-01
  30. PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents 21 upvotes, #30 of 2026-10-01
  31. Scaling Laws for Looped Mixture of Experts 21 upvotes, #30 of 2026-10-01
  32. The Low-Rank Structure of VLA Reinforcement Learning 19 upvotes, #32 of 2026-10-01
  33. CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering 18 upvotes, #33 of 2026-10-01
  34. DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation 17 upvotes, #34 of 2026-10-01
  35. Rubric Rewards from Item Response Theory 16 upvotes, #35 of 2026-10-01
  36. MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution 16 upvotes, #35 of 2026-10-01
  37. Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model 15 upvotes, #37 of 2026-10-01
  38. WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms 14 upvotes, #38 of 2026-10-01
  39. A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications? 14 upvotes, #38 of 2026-10-01
  40. RoboCoach: World Models as Active Coaches for Compositional Robot Skills 14 upvotes, #38 of 2026-10-01
  41. I Have a Stream: Making Self-Supervised Learning Work on Continuous Video 14 upvotes, #38 of 2026-10-01
  42. Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning 12 upvotes, #42 of 2026-10-01
  43. The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends 12 upvotes, #42 of 2026-10-01
  44. Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior 12 upvotes, #42 of 2026-10-01
  45. SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing 11 upvotes, #45 of 2026-10-01
  46. AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation 10 upvotes, #46 of 2026-10-01
  47. CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding 10 upvotes, #46 of 2026-10-01
  48. PatchHolmes: Agentic Patch Retrieval via Listwise Selection 9 upvotes, #48 of 2026-10-01
  49. SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation 7 upvotes, #49 of 2026-10-01
  50. Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning 7 upvotes, #49 of 2026-10-01
  51. Safety of Latent Communication in Multi-Agent Systems 7 upvotes, #49 of 2026-10-01
  52. SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale 6 upvotes, #52 of 2026-10-01
  53. Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD 5 upvotes, #53 of 2026-10-01
  54. NavHarness: Towards Lifelong Embodied Navigation 5 upvotes, #53 of 2026-10-01
  55. Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering 5 upvotes, #53 of 2026-10-01
  56. Training LLM Judges from Language Feedback via Position-Selective Self-Distillation 5 upvotes, #53 of 2026-10-01
  57. Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds 5 upvotes, #53 of 2026-10-01
  58. Game-Guided Skill Discovery through Self-Play for Playable Agent Control 5 upvotes, #53 of 2026-10-01
  59. Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text 5 upvotes, #53 of 2026-10-01
  60. Understanding Multimodality in Generative Behavioral Cloning 4 upvotes, #60 of 2026-10-01
  61. See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs 4 upvotes, #60 of 2026-10-01
  62. CheatBench: Measuring Reward Gaming in AI Agents 4 upvotes, #60 of 2026-10-01
  63. Mitigating the Length-Scaling Tax with Online Distillation 4 upvotes, #60 of 2026-10-01
  64. DAGent: Evaluate-then-Grow Planning for Deep Research Agents 4 upvotes, #60 of 2026-10-01
  65. How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text 4 upvotes, #60 of 2026-10-01
  66. Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces 3 upvotes, #66 of 2026-10-01
  67. The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence 3 upvotes, #66 of 2026-10-01
  68. BIABench: Evaluating AI agents on real-world bioimage analysis tasks 3 upvotes, #66 of 2026-10-01
  69. Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation 3 upvotes, #66 of 2026-10-01
  70. Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance 3 upvotes, #66 of 2026-10-01
  71. SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models 3 upvotes, #66 of 2026-10-01
  72. Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost 3 upvotes, #66 of 2026-10-01
  73. EviRover: Reinforcing Agentic Perception Beyond a Glance 3 upvotes, #66 of 2026-10-01
  74. Decision-Oriented Recommendation Reranking: An Empirical Study of Jev 3 upvotes, #66 of 2026-10-01
  75. Aligning One-Step Generative Models with Reward-Weighted Transport Distillation 2 upvotes, #75 of 2026-10-01
  76. SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs 2 upvotes, #75 of 2026-10-01
  77. Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change 2 upvotes, #75 of 2026-10-01
  78. Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems 2 upvotes, #75 of 2026-10-01
  79. ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning 2 upvotes, #75 of 2026-10-01
  80. The Geometry of Inference in Transformer Residual Streams 2 upvotes, #75 of 2026-10-01
  81. Retrieval Capacity of Self-Attention Under Competition 2 upvotes, #75 of 2026-10-01
  82. Soft Spatial Reasoning 2 upvotes, #75 of 2026-10-01
  83. Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers 2 upvotes, #75 of 2026-10-01
  84. Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings 2 upvotes, #75 of 2026-10-01
  85. LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception 2 upvotes, #75 of 2026-10-01
  86. MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories 2 upvotes, #75 of 2026-10-01

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.