Daily Papers of 2026-10-02

  1. OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction 196 upvotes, #1 of 2026-10-02
  2. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics 176 upvotes, #2 of 2026-10-02
  3. GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis 144 upvotes, #3 of 2026-10-02
  4. Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL 124 upvotes, #4 of 2026-10-02
  5. Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts 90 upvotes, #5 of 2026-10-02
  6. Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States 83 upvotes, #6 of 2026-10-02
  7. Sharpening Tax in Post-Training 83 upvotes, #6 of 2026-10-02
  8. Hierarchical Continuous Diffusion Language Models 80 upvotes, #8 of 2026-10-02
  9. Agent Priors-guided Policy Learning 77 upvotes, #9 of 2026-10-02
  10. World Observer: Joint Actor-Observer Generation for Persistent World Modeling 76 upvotes, #10 of 2026-10-02
  11. Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It 72 upvotes, #11 of 2026-10-02
  12. A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review 71 upvotes, #12 of 2026-10-02
  13. Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs 67 upvotes, #13 of 2026-10-02
  14. X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization 65 upvotes, #14 of 2026-10-02
  15. ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization 64 upvotes, #15 of 2026-10-02
  16. ROWBench: Do Video Models Render What the Program Specifies? 64 upvotes, #15 of 2026-10-02
  17. E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models 60 upvotes, #17 of 2026-10-02
  18. EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos 57 upvotes, #18 of 2026-10-02
  19. Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL 57 upvotes, #18 of 2026-10-02
  20. Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation 56 upvotes, #20 of 2026-10-02
  21. Video Generation Models: A Survey of Post-Training and Alignment 55 upvotes, #21 of 2026-10-02
  22. AutoGUIWorld: Image Generators as Visual World Models for GUI Agent 55 upvotes, #21 of 2026-10-02
  23. InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation 55 upvotes, #21 of 2026-10-02
  24. Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding 51 upvotes, #24 of 2026-10-02
  25. Scaling and Distilling Text Embeddings for Better Diffusibility 50 upvotes, #25 of 2026-10-02
  26. Persona Dosing: Calibrated Activation Steering for Graded Trait Control 48 upvotes, #26 of 2026-10-02
  27. Beyond the Current Scene: Event-Referential Grasping with Active View Selection 46 upvotes, #27 of 2026-10-02
  28. Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens 45 upvotes, #28 of 2026-10-02
  29. Decoding Looped Transformers Better for (Almost) Free 42 upvotes, #29 of 2026-10-02
  30. Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans 41 upvotes, #30 of 2026-10-02
  31. AutoDataBench: A Data-centric Testbed for Accelerating Auto Research 37 upvotes, #31 of 2026-10-02
  32. SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation 35 upvotes, #32 of 2026-10-02
  33. 4Director: Controlling Video World Models with Rigid 3D Geometry 34 upvotes, #33 of 2026-10-02
  34. RPTune: Learned Context Curation for LLM Catalog Search 31 upvotes, #34 of 2026-10-02
  35. CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning 30 upvotes, #35 of 2026-10-02
  36. Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces 27 upvotes, #36 of 2026-10-02
  37. Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows 26 upvotes, #37 of 2026-10-02
  38. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation 25 upvotes, #38 of 2026-10-02
  39. PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop 23 upvotes, #39 of 2026-10-02
  40. Smaller Models, Better Rejects: Preference Distillation Scaling 21 upvotes, #40 of 2026-10-02
  41. Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation 20 upvotes, #41 of 2026-10-02
  42. PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion 19 upvotes, #42 of 2026-10-02
  43. When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents 17 upvotes, #43 of 2026-10-02
  44. Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing 17 upvotes, #43 of 2026-10-02
  45. Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes 17 upvotes, #43 of 2026-10-02
  46. Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing 16 upvotes, #46 of 2026-10-02
  47. Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation 16 upvotes, #46 of 2026-10-02
  48. Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation 16 upvotes, #46 of 2026-10-02
  49. DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration 15 upvotes, #49 of 2026-10-02
  50. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning 15 upvotes, #49 of 2026-10-02
  51. OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories 14 upvotes, #51 of 2026-10-02
  52. AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines 13 upvotes, #52 of 2026-10-02
  53. RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers 12 upvotes, #53 of 2026-10-02
  54. Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation 12 upvotes, #53 of 2026-10-02
  55. Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents 11 upvotes, #55 of 2026-10-02
  56. Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models 11 upvotes, #55 of 2026-10-02
  57. Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models 11 upvotes, #55 of 2026-10-02
  58. LOCI: Spatial Linear Memory for Streaming World Models 11 upvotes, #55 of 2026-10-02
  59. Memorizon: Training World Models Beyond Their Context Window 11 upvotes, #55 of 2026-10-02
  60. SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation 11 upvotes, #55 of 2026-10-02
  61. Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs 11 upvotes, #55 of 2026-10-02
  62. VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation 11 upvotes, #55 of 2026-10-02
  63. ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research 11 upvotes, #55 of 2026-10-02
  64. Rules to Tools: Executable Checks for LLM Agents in Scientific Computing 10 upvotes, #64 of 2026-10-02
  65. Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling 9 upvotes, #65 of 2026-10-02
  66. Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions 9 upvotes, #65 of 2026-10-02
  67. Does Native 3D Texture Generation Necessarily Require 3D Assets for Training? 8 upvotes, #67 of 2026-10-02
  68. FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing 8 upvotes, #67 of 2026-10-02
  69. KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards 8 upvotes, #67 of 2026-10-02
  70. Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy 7 upvotes, #70 of 2026-10-02
  71. MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization 7 upvotes, #70 of 2026-10-02
  72. Personalized Image Generation with Reasoning and Reflection 7 upvotes, #70 of 2026-10-02
  73. When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs 6 upvotes, #73 of 2026-10-02
  74. Controlled Decoding Attacks on Black-Box LLMs 6 upvotes, #73 of 2026-10-02
  75. DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation 6 upvotes, #73 of 2026-10-02
  76. JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces 6 upvotes, #73 of 2026-10-02
  77. It Takes Workflows to Evolve Better Workflows 6 upvotes, #73 of 2026-10-02
  78. Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration 5 upvotes, #78 of 2026-10-02
  79. OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport 5 upvotes, #78 of 2026-10-02
  80. Honeycomb: Constant-Size Scene Memory Representation for Video World Models 5 upvotes, #78 of 2026-10-02
  81. Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization 5 upvotes, #78 of 2026-10-02
  82. Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs 5 upvotes, #78 of 2026-10-02
  83. Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models 5 upvotes, #78 of 2026-10-02
  84. Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models 5 upvotes, #78 of 2026-10-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.