Daily Papers of 2026-09-30
- Raven: The Harness of Harnesses for Composable Agentic Intelligence 563 upvotes, #1 of 2026-09-30
- MaLiang-Harness: A Programmable Path to Image and Video Generation 410 upvotes, #2 of 2026-09-30
- In-Context Learning for Robots: Methods and Applications 387 upvotes, #3 of 2026-09-30
- Scaling Properties of Same-Family On-Policy Distillation 321 upvotes, #4 of 2026-09-30
- Omni-IO Skills: Harnessing Your Agent Omni-Native 292 upvotes, #5 of 2026-09-30
- PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation 166 upvotes, #6 of 2026-09-30
- VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models 153 upvotes, #7 of 2026-09-30
- What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling 136 upvotes, #8 of 2026-09-30
- LEGO-Anything: Coding Agents for 3D Scene Reconstruction 135 upvotes, #9 of 2026-09-30
- Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies 113 upvotes, #10 of 2026-09-30
- Think Before You Score: Thinking Reward Model for Visual Generation 101 upvotes, #11 of 2026-09-30
- Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression 100 upvotes, #12 of 2026-09-30
- Follow the Entities: A Corpus Map for Agentic Search 98 upvotes, #13 of 2026-09-30
- SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation 94 upvotes, #14 of 2026-09-30
- LLMs are General Asynchronous Agents 82 upvotes, #15 of 2026-09-30
- EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? 70 upvotes, #16 of 2026-09-30
- Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations 67 upvotes, #17 of 2026-09-30
- LongCat-DeepResearch Technical Report 66 upvotes, #18 of 2026-09-30
- Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR 62 upvotes, #19 of 2026-09-30
- OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? 55 upvotes, #20 of 2026-09-30
- WorldAttention: An Efficient Attention Architecture for Interactive Video World Models 49 upvotes, #21 of 2026-09-30
- ROSS: Relearning from Self-Generated Rollouts through Selective Supervision 49 upvotes, #21 of 2026-09-30
- EVO-WAM: Evolving World Action Models through Video-Action Verification 48 upvotes, #23 of 2026-09-30
- HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents 47 upvotes, #24 of 2026-09-30
- APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants 46 upvotes, #25 of 2026-09-30
- Marathoner: Ultra-Long-Horizon Autonomous Intelligence 42 upvotes, #26 of 2026-09-30
- SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video 41 upvotes, #27 of 2026-09-30
- LongLive-Plug: Once-for-All Distillation for Video Generation 41 upvotes, #27 of 2026-09-30
- Context Language Models 40 upvotes, #29 of 2026-09-30
- Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents 38 upvotes, #30 of 2026-09-30
- ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces 35 upvotes, #31 of 2026-09-30
- EasyPPO: Stabilizing the Critic Is Key 35 upvotes, #31 of 2026-09-30
- FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution 33 upvotes, #33 of 2026-09-30
- Anisotropic Representations Improve Planning in JEPA World Models 32 upvotes, #34 of 2026-09-30
- WorldLine: Action-Driven Visual Simulation for Robotic Manipulation 32 upvotes, #34 of 2026-09-30
- Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents 30 upvotes, #36 of 2026-09-30
- When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections 26 upvotes, #37 of 2026-09-30
- Chinese-Jev: Bringing System One Model to Chinese-Language Tasks 25 upvotes, #38 of 2026-09-30
- HiRAE: Hierarchical Representation Autoencoding with Residual Budgets 25 upvotes, #38 of 2026-09-30
- CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments 24 upvotes, #40 of 2026-09-30
- Reasoning with Image Generation 23 upvotes, #41 of 2026-09-30
- On the Off-Policy Teacher in On-Policy Distillation 22 upvotes, #42 of 2026-09-30
- AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation 21 upvotes, #43 of 2026-09-30
- Adversarial Training for Pixel Diffusion 21 upvotes, #43 of 2026-09-30
- TabFM: A Zero-Shot Foundation Model for Tabular Data 20 upvotes, #45 of 2026-09-30
- EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation 20 upvotes, #45 of 2026-09-30
- Scheduling Recursive Reasoning in Looped Transformers 17 upvotes, #47 of 2026-09-30
- Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing 17 upvotes, #47 of 2026-09-30
- Language Models Are "Insecure" Reporters 16 upvotes, #49 of 2026-09-30
- Selecting The Most Informative Tokens in Natural Language Autoencoders 16 upvotes, #49 of 2026-09-30
- Can Agents Design Libraries for Agents? 11 upvotes, #51 of 2026-09-30
- TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models 11 upvotes, #51 of 2026-09-30
- VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents 11 upvotes, #51 of 2026-09-30
- StoryEngine: A State-Grounded Agentic Framework for Video Storytelling 10 upvotes, #54 of 2026-09-30
- Fractional State Space Transition for Long Sequence Modeling 10 upvotes, #54 of 2026-09-30
- What Makes Recurrence Effective in Looped Language Models? 10 upvotes, #54 of 2026-09-30
- Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards 9 upvotes, #57 of 2026-09-30
- TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs 8 upvotes, #58 of 2026-09-30
- FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets 8 upvotes, #58 of 2026-09-30
- EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory 8 upvotes, #58 of 2026-09-30
- ALICE: In-context, Zero-shot, Mutual Information Estimation 7 upvotes, #61 of 2026-09-30
- Beyond Selection: Token Parameterization for Extreme Visual Token Compression 7 upvotes, #61 of 2026-09-30
- Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration 7 upvotes, #61 of 2026-09-30
- Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection 7 upvotes, #61 of 2026-09-30
- Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss 6 upvotes, #65 of 2026-09-30
- AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop? 6 upvotes, #65 of 2026-09-30
- Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents 6 upvotes, #65 of 2026-09-30
- PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents 6 upvotes, #65 of 2026-09-30
- Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics 6 upvotes, #65 of 2026-09-30
- Real2Gym: Building Gyms from Videos, Bringing Skills to Robots 6 upvotes, #65 of 2026-09-30
- WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation 6 upvotes, #65 of 2026-09-30
- Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE 6 upvotes, #65 of 2026-09-30
- LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation 6 upvotes, #65 of 2026-09-30
- Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning 5 upvotes, #74 of 2026-09-30
- CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs 5 upvotes, #74 of 2026-09-30
- Pretraining Transformers with Quantized Softmax in Attention 5 upvotes, #74 of 2026-09-30
- Org-Agent: Beyond Personal Assistants Towards Organizational Agents 5 upvotes, #74 of 2026-09-30
- Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement 5 upvotes, #74 of 2026-09-30
- PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation 5 upvotes, #74 of 2026-09-30
- Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders 5 upvotes, #74 of 2026-09-30
- Principled Thoughts for Latent Recursive LLM Systems 5 upvotes, #74 of 2026-09-30
- StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks 5 upvotes, #74 of 2026-09-30
- Improved Distributional Diffusion Models 5 upvotes, #74 of 2026-09-30
- PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers 4 upvotes, #84 of 2026-09-30
- AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models 4 upvotes, #84 of 2026-09-30
- PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond? 4 upvotes, #84 of 2026-09-30
- One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices 4 upvotes, #84 of 2026-09-30
- Hyperspherical Semantic Trajectory Analysis: Mapping Technological Diffusion across Academic Preprints, Patent Signals, and Compute Scaling 4 upvotes, #84 of 2026-09-30
- Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion 4 upvotes, #84 of 2026-09-30
- SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents 3 upvotes, #90 of 2026-09-30
- AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents 2 upvotes, #91 of 2026-09-30
- Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation 0 upvotes, #92 of 2026-09-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.