Daily Papers of 2026-09-29

  1. Post-Training Leaves Behavioral Shadows on Unrelated Decisions 271 upvotes, #1 of 2026-09-29
  2. YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality 244 upvotes, #2 of 2026-09-29
  3. VisionHOPE: Visual Backbones as Self-Modifying Learning Systems 217 upvotes, #3 of 2026-09-29
  4. Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation 180 upvotes, #4 of 2026-09-29
  5. Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL 129 upvotes, #5 of 2026-09-29
  6. Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence 107 upvotes, #6 of 2026-09-29
  7. Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue 89 upvotes, #7 of 2026-09-29
  8. MassAlloc Attention: Let Attention Allocate Its Own Compute 75 upvotes, #8 of 2026-09-29
  9. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces 72 upvotes, #9 of 2026-09-29
  10. CoWindow Attention: Full Causal Coverage Is a Collective Property 67 upvotes, #10 of 2026-09-29
  11. How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining 65 upvotes, #11 of 2026-09-29
  12. Improving Test-Time Scaling with Adaptive Looped Transformers 58 upvotes, #12 of 2026-09-29
  13. Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning 49 upvotes, #13 of 2026-09-29
  14. Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning 45 upvotes, #14 of 2026-09-29
  15. QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents 43 upvotes, #15 of 2026-09-29
  16. CompoWorld: Compositional Environment Scaling for General Agents 42 upvotes, #16 of 2026-09-29
  17. Recursive Harness Distillation across Agents for Robot Manipulation 40 upvotes, #17 of 2026-09-29
  18. Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge 40 upvotes, #17 of 2026-09-29
  19. Learning from Teacher Continuations at Student States 40 upvotes, #17 of 2026-09-29
  20. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks 39 upvotes, #20 of 2026-09-29
  21. Selecting Diverse SFT Traces Improves Post-RL Generalization 38 upvotes, #21 of 2026-09-29
  22. Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents 35 upvotes, #22 of 2026-09-29
  23. Diffusion Reward Models 35 upvotes, #22 of 2026-09-29
  24. In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion 34 upvotes, #24 of 2026-09-29
  25. FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching 33 upvotes, #25 of 2026-09-29
  26. DepthBench: Measuring How Residual Connections Enable More Computational Depth 32 upvotes, #26 of 2026-09-29
  27. Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It 31 upvotes, #27 of 2026-09-29
  28. Learning to Learn from Context: Synthetic Training from Perturbed Public Documents 31 upvotes, #27 of 2026-09-29
  29. WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon 31 upvotes, #27 of 2026-09-29
  30. Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models 30 upvotes, #30 of 2026-09-29
  31. SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning 26 upvotes, #31 of 2026-09-29
  32. RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents 25 upvotes, #32 of 2026-09-29
  33. An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning 25 upvotes, #32 of 2026-09-29
  34. Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision 23 upvotes, #34 of 2026-09-29
  35. Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers 22 upvotes, #35 of 2026-09-29
  36. REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening 20 upvotes, #36 of 2026-09-29
  37. Precise Editing and Flexible Referencing for Interactable Worlds 20 upvotes, #36 of 2026-09-29
  38. AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research 19 upvotes, #38 of 2026-09-29
  39. ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis 19 upvotes, #38 of 2026-09-29
  40. Nereus: Adaptive Parallelism for LLM Post-Training 19 upvotes, #38 of 2026-09-29
  41. VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction 19 upvotes, #38 of 2026-09-29
  42. SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation 18 upvotes, #42 of 2026-09-29
  43. TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining 18 upvotes, #42 of 2026-09-29
  44. Structured Residual Connectivity Matters for Diffusion Transformers 17 upvotes, #44 of 2026-09-29
  45. WideSWE: Can Coding Agents Coordinate Changes Across Repositories? 17 upvotes, #44 of 2026-09-29
  46. Imprint Reader: From Weight-Update Readout to Behavioral Intervention 16 upvotes, #46 of 2026-09-29
  47. SolveEdit: Benchmarking Visual Problem Solving in Generative Models 16 upvotes, #46 of 2026-09-29
  48. Program-Verified Self-Evolution for Vision-Language Models 14 upvotes, #48 of 2026-09-29
  49. Residual Transferability in Neural Image Watermarking 12 upvotes, #49 of 2026-09-29
  50. InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video 9 upvotes, #50 of 2026-09-29
  51. WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models 8 upvotes, #51 of 2026-09-29
  52. Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction 8 upvotes, #51 of 2026-09-29
  53. BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation 8 upvotes, #51 of 2026-09-29
  54. GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space 8 upvotes, #51 of 2026-09-29
  55. Relic: From Multi-Agent Collaboration to Persistent Organizational Capability 7 upvotes, #55 of 2026-09-29
  56. Draft-KV: Learning Useful Latent Communication Between Language Models 7 upvotes, #55 of 2026-09-29
  57. ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport 7 upvotes, #55 of 2026-09-29
  58. RenderRank: Learning to Rerank Text with Compressed Visual Tokens 7 upvotes, #55 of 2026-09-29
  59. Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry 6 upvotes, #59 of 2026-09-29
  60. SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis 6 upvotes, #59 of 2026-09-29
  61. Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation 6 upvotes, #59 of 2026-09-29
  62. Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs 5 upvotes, #62 of 2026-09-29
  63. KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation 5 upvotes, #62 of 2026-09-29
  64. DroneWAM: Efficient World Action Model for Drone Visual Navigation 5 upvotes, #62 of 2026-09-29
  65. SMAT: Simple and Efficient Merge-Aware Training 5 upvotes, #62 of 2026-09-29
  66. Measuring Collapse and Correction in Homogeneous-Panel LLM Debate 5 upvotes, #62 of 2026-09-29
  67. Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization 4 upvotes, #67 of 2026-09-29
  68. G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation 4 upvotes, #67 of 2026-09-29
  69. VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis 4 upvotes, #67 of 2026-09-29
  70. Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents 4 upvotes, #67 of 2026-09-29
  71. Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration 4 upvotes, #67 of 2026-09-29
  72. PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction 4 upvotes, #67 of 2026-09-29
  73. KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems 4 upvotes, #67 of 2026-09-29
  74. When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety 4 upvotes, #67 of 2026-09-29
  75. FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models 4 upvotes, #67 of 2026-09-29
  76. Reinforcing Agentic Creativity in Scientific Ideation with Night Science 4 upvotes, #67 of 2026-09-29
  77. TokenCast: Forecasting Token Consumption During LLM Agent Execution 4 upvotes, #67 of 2026-09-29
  78. WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing 3 upvotes, #78 of 2026-09-29
  79. Rolling-WAM: World Action Models with Rolling Imagination 3 upvotes, #78 of 2026-09-29
  80. Self-Play Search Distillation for Large Language Model Reasoning 3 upvotes, #78 of 2026-09-29
  81. NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech 3 upvotes, #78 of 2026-09-29
  82. Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions 3 upvotes, #78 of 2026-09-29
  83. Adaptive Consistency Graph for Long-Horizon Agents 3 upvotes, #78 of 2026-09-29
  84. Allspark: Weak to Strong Transfer via Alternating Chain of Thought 3 upvotes, #78 of 2026-09-29
  85. DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation 3 upvotes, #78 of 2026-09-29
  86. What masking geometry works best for EEG foundation models? 3 upvotes, #78 of 2026-09-29
  87. Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training 3 upvotes, #78 of 2026-09-29
  88. Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models 3 upvotes, #78 of 2026-09-29
  89. Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation 3 upvotes, #78 of 2026-09-29
  90. On-Policy Self-Distillation for Multi-Turn Image Editing 3 upvotes, #78 of 2026-09-29
  91. EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold 3 upvotes, #78 of 2026-09-29
  92. Distillation Defenses Easily Break After Reinforcement Learning 3 upvotes, #78 of 2026-09-29
  93. Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation 2 upvotes, #93 of 2026-09-29
  94. How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure 2 upvotes, #93 of 2026-09-29
  95. NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization 2 upvotes, #93 of 2026-09-29
  96. SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback 2 upvotes, #93 of 2026-09-29
  97. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models 2 upvotes, #93 of 2026-09-29
  98. ControlScope: Workflow Revision and Reliability in LLM Agents 2 upvotes, #93 of 2026-09-29
  99. AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors 2 upvotes, #93 of 2026-09-29
  100. Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning 2 upvotes, #93 of 2026-09-29
  101. LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning 2 upvotes, #93 of 2026-09-29
  102. How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline 1 upvotes, #102 of 2026-09-29
  103. PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift 1 upvotes, #102 of 2026-09-29
  104. When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions 1 upvotes, #102 of 2026-09-29
  105. Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models? 1 upvotes, #102 of 2026-09-29
  106. Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction 1 upvotes, #102 of 2026-09-29
  107. AdaGuard: An Adaptive Guard Model with User-defined Policies 1 upvotes, #102 of 2026-09-29
  108. CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles 1 upvotes, #102 of 2026-09-29
  109. Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks 0 upvotes, #109 of 2026-09-29

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.