Daily Papers of 2026-09-18

  1. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression 180 upvotes, #1 of 2026-09-18
  2. SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness 130 upvotes, #2 of 2026-09-18
  3. Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model 113 upvotes, #3 of 2026-09-18
  4. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation 107 upvotes, #4 of 2026-09-18
  5. An Empirical Study of Harness Design for Coding Agents 86 upvotes, #5 of 2026-09-18
  6. JEPA-Anything: Learning Predictive Models across Different Worlds 70 upvotes, #6 of 2026-09-18
  7. Verifiable Social Reasoning for LLM Assistants 65 upvotes, #7 of 2026-09-18
  8. Self-Evolving Search Index 60 upvotes, #8 of 2026-09-18
  9. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning 56 upvotes, #9 of 2026-09-18
  10. When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models 52 upvotes, #10 of 2026-09-18
  11. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation 50 upvotes, #11 of 2026-09-18
  12. RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation 47 upvotes, #12 of 2026-09-18
  13. WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing 47 upvotes, #12 of 2026-09-18
  14. Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents 45 upvotes, #14 of 2026-09-18
  15. UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation 43 upvotes, #15 of 2026-09-18
  16. Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL 42 upvotes, #16 of 2026-09-18
  17. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control 41 upvotes, #17 of 2026-09-18
  18. Region-Level Policy Optimization for Fine-grained MLLM Perception 41 upvotes, #17 of 2026-09-18
  19. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure? 40 upvotes, #19 of 2026-09-18
  20. FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations 38 upvotes, #20 of 2026-09-18
  21. Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts 37 upvotes, #21 of 2026-09-18
  22. Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling 37 upvotes, #21 of 2026-09-18
  23. What Does Privileged Information Add to On-Policy Self-Distillation? 36 upvotes, #23 of 2026-09-18
  24. VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering 31 upvotes, #24 of 2026-09-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.