Daily Papers of 2026-09-10
- Scaling Automatic Research Agents via World Models 459 upvotes, #1 of 2026-09-10
- Show-Harness: Just a VLM Agent Can Play Robots 160 upvotes, #2 of 2026-09-10
- Programmable World Model 109 upvotes, #3 of 2026-09-10
- AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems 73 upvotes, #4 of 2026-09-10
- T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks 59 upvotes, #5 of 2026-09-10
- WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data 29 upvotes, #6 of 2026-09-10
- PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents 24 upvotes, #7 of 2026-09-10
- Revisiting Complete Reasoning Traces for Post-Training 23 upvotes, #8 of 2026-09-10
- Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation 22 upvotes, #9 of 2026-09-10
- Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? 21 upvotes, #10 of 2026-09-10
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? 20 upvotes, #11 of 2026-09-10
- Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents 20 upvotes, #11 of 2026-09-10
- SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators 16 upvotes, #13 of 2026-09-10
- Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs 15 upvotes, #14 of 2026-09-10
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents 14 upvotes, #15 of 2026-09-10
- DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents 13 upvotes, #16 of 2026-09-10
- OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution 12 upvotes, #17 of 2026-09-10
- Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR 12 upvotes, #17 of 2026-09-10
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents 11 upvotes, #19 of 2026-09-10
- A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware 10 upvotes, #20 of 2026-09-10
- Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models 10 upvotes, #20 of 2026-09-10
- Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails 10 upvotes, #20 of 2026-09-10
- PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving 9 upvotes, #23 of 2026-09-10
- AgenticGen: Reward-Guided Agentic Video Generation for Advertising 9 upvotes, #23 of 2026-09-10
- StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean 9 upvotes, #23 of 2026-09-10
- Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning 7 upvotes, #26 of 2026-09-10
- DF26: We Cannot Tell Fake From Real Anymore 6 upvotes, #27 of 2026-09-10
- RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems 5 upvotes, #28 of 2026-09-10
- Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States 5 upvotes, #28 of 2026-09-10
- The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding 5 upvotes, #28 of 2026-09-10
- The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements 4 upvotes, #31 of 2026-09-10
- From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution 4 upvotes, #31 of 2026-09-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.