Daily Papers of 2026-08-18

  1. StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling 439 upvotes, #1 of 2026-08-18
  2. Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization 148 upvotes, #2 of 2026-08-18
  3. HarnessEval-W: Agentifying the Evaluation of Visual Worlds 123 upvotes, #3 of 2026-08-18
  4. Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search 62 upvotes, #4 of 2026-08-18
  5. VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? 59 upvotes, #5 of 2026-08-18
  6. MOSS-VL Technical Report 47 upvotes, #6 of 2026-08-18
  7. UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations 47 upvotes, #6 of 2026-08-18
  8. ClawGym II: Exploring Black-Box RL on Agent Harness 39 upvotes, #8 of 2026-08-18
  9. An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models 33 upvotes, #9 of 2026-08-18
  10. How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks 30 upvotes, #10 of 2026-08-18
  11. PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments 28 upvotes, #11 of 2026-08-18
  12. Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI 28 upvotes, #11 of 2026-08-18
  13. Agentic Transaction: Towards ACID-Compliant Agent Systems 27 upvotes, #13 of 2026-08-18
  14. GenRouter: Unified Workflow Routing for Agentic Image Generation 25 upvotes, #14 of 2026-08-18
  15. HarmProfile: Characterizing Harmful Distributions in Frontier LLMs 20 upvotes, #15 of 2026-08-18
  16. MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling 19 upvotes, #16 of 2026-08-18
  17. Improving the matrix multiplication exponent with modern optimization and AlphaEvolve 18 upvotes, #17 of 2026-08-18
  18. Understanding Cognition-Induced Risks in Agentic AI Systems 17 upvotes, #18 of 2026-08-18
  19. GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks 13 upvotes, #19 of 2026-08-18
  20. Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs 13 upvotes, #19 of 2026-08-18
  21. TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation 13 upvotes, #19 of 2026-08-18
  22. R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets 12 upvotes, #22 of 2026-08-18
  23. VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding 10 upvotes, #23 of 2026-08-18
  24. AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model 9 upvotes, #24 of 2026-08-18
  25. StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding 9 upvotes, #24 of 2026-08-18
  26. Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning 8 upvotes, #26 of 2026-08-18
  27. ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering 6 upvotes, #27 of 2026-08-18
  28. NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents 6 upvotes, #27 of 2026-08-18
  29. ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval 6 upvotes, #27 of 2026-08-18
  30. DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs 4 upvotes, #30 of 2026-08-18
  31. Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form 4 upvotes, #30 of 2026-08-18
  32. Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency 4 upvotes, #30 of 2026-08-18
  33. Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents 3 upvotes, #33 of 2026-08-18
  34. When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse 2 upvotes, #34 of 2026-08-18
  35. Accuracy and Order Sensitivity Diverge Under Label-Free Strategies 2 upvotes, #34 of 2026-08-18
  36. Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays 2 upvotes, #34 of 2026-08-18
  37. Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems 2 upvotes, #34 of 2026-08-18
  38. WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations 2 upvotes, #34 of 2026-08-18
  39. Drive, Pack, Fly: The Travelling Thief Problem with Drone 2 upvotes, #34 of 2026-08-18
  40. Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift 1 upvotes, #40 of 2026-08-18
  41. A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models 1 upvotes, #40 of 2026-08-18
  42. HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation 1 upvotes, #40 of 2026-08-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.