Daily Papers of 2026-09-18
- DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression 180 upvotes, #1 of 2026-09-18
- SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness 130 upvotes, #2 of 2026-09-18
- Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model 113 upvotes, #3 of 2026-09-18
- When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation 107 upvotes, #4 of 2026-09-18
- An Empirical Study of Harness Design for Coding Agents 86 upvotes, #5 of 2026-09-18
- JEPA-Anything: Learning Predictive Models across Different Worlds 70 upvotes, #6 of 2026-09-18
- Verifiable Social Reasoning for LLM Assistants 65 upvotes, #7 of 2026-09-18
- Self-Evolving Search Index 60 upvotes, #8 of 2026-09-18
- RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning 56 upvotes, #9 of 2026-09-18
- When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models 52 upvotes, #10 of 2026-09-18
- Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation 50 upvotes, #11 of 2026-09-18
- RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation 47 upvotes, #12 of 2026-09-18
- WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing 47 upvotes, #12 of 2026-09-18
- Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents 45 upvotes, #14 of 2026-09-18
- UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation 43 upvotes, #15 of 2026-09-18
- Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL 42 upvotes, #16 of 2026-09-18
- VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control 41 upvotes, #17 of 2026-09-18
- Region-Level Policy Optimization for Fine-grained MLLM Perception 41 upvotes, #17 of 2026-09-18
- PACT: Can Enterprise AI Assistants Be Trusted Under Pressure? 40 upvotes, #19 of 2026-09-18
- FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations 38 upvotes, #20 of 2026-09-18
- Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts 37 upvotes, #21 of 2026-09-18
- Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling 37 upvotes, #21 of 2026-09-18
- What Does Privileged Information Add to On-Policy Self-Distillation? 36 upvotes, #23 of 2026-09-18
- VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering 31 upvotes, #24 of 2026-09-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.