Daily Papers of 2026-09-10

  1. Scaling Automatic Research Agents via World Models 459 upvotes, #1 of 2026-09-10
  2. Show-Harness: Just a VLM Agent Can Play Robots 160 upvotes, #2 of 2026-09-10
  3. Programmable World Model 109 upvotes, #3 of 2026-09-10
  4. AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems 73 upvotes, #4 of 2026-09-10
  5. T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks 59 upvotes, #5 of 2026-09-10
  6. WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data 29 upvotes, #6 of 2026-09-10
  7. PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents 24 upvotes, #7 of 2026-09-10
  8. Revisiting Complete Reasoning Traces for Post-Training 23 upvotes, #8 of 2026-09-10
  9. Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation 22 upvotes, #9 of 2026-09-10
  10. Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? 21 upvotes, #10 of 2026-09-10
  11. SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? 20 upvotes, #11 of 2026-09-10
  12. Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents 20 upvotes, #11 of 2026-09-10
  13. SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators 16 upvotes, #13 of 2026-09-10
  14. Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs 15 upvotes, #14 of 2026-09-10
  15. SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents 14 upvotes, #15 of 2026-09-10
  16. DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents 13 upvotes, #16 of 2026-09-10
  17. OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution 12 upvotes, #17 of 2026-09-10
  18. Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR 12 upvotes, #17 of 2026-09-10
  19. SchemeArena: Factorized Stress Testing of Scheming in LLM Agents 11 upvotes, #19 of 2026-09-10
  20. A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware 10 upvotes, #20 of 2026-09-10
  21. Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models 10 upvotes, #20 of 2026-09-10
  22. Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails 10 upvotes, #20 of 2026-09-10
  23. PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving 9 upvotes, #23 of 2026-09-10
  24. AgenticGen: Reward-Guided Agentic Video Generation for Advertising 9 upvotes, #23 of 2026-09-10
  25. StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean 9 upvotes, #23 of 2026-09-10
  26. Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning 7 upvotes, #26 of 2026-09-10
  27. DF26: We Cannot Tell Fake From Real Anymore 6 upvotes, #27 of 2026-09-10
  28. RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems 5 upvotes, #28 of 2026-09-10
  29. Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States 5 upvotes, #28 of 2026-09-10
  30. The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding 5 upvotes, #28 of 2026-09-10
  31. The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements 4 upvotes, #31 of 2026-09-10
  32. From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution 4 upvotes, #31 of 2026-09-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.