Daily Papers of 2026-08-13

  1. Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill 283 upvotes, #1 of 2026-08-13
  2. OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution 260 upvotes, #2 of 2026-08-13
  3. AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses 111 upvotes, #3 of 2026-08-13
  4. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence 84 upvotes, #4 of 2026-08-13
  5. SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries 75 upvotes, #5 of 2026-08-13
  6. AVA-Encoder: Towards Agent-Native Video Representation Learning 40 upvotes, #6 of 2026-08-13
  7. Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives 29 upvotes, #7 of 2026-08-13
  8. StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization 26 upvotes, #8 of 2026-08-13
  9. Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models 15 upvotes, #9 of 2026-08-13
  10. AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research 13 upvotes, #10 of 2026-08-13
  11. Self-Evolving Embodied Agents via Skill-Harness Evolution 13 upvotes, #10 of 2026-08-13
  12. Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning 11 upvotes, #12 of 2026-08-13
  13. ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents 10 upvotes, #13 of 2026-08-13
  14. From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection 9 upvotes, #14 of 2026-08-13
  15. The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images 8 upvotes, #15 of 2026-08-13
  16. Parameter Exploration for RLVR via Variational Learning 6 upvotes, #16 of 2026-08-13
  17. Persistent Recursive Worlds Enable Autonomous Software Evolution 6 upvotes, #16 of 2026-08-13
  18. Simplex Relaxation for Discrete Diffusion 6 upvotes, #16 of 2026-08-13
  19. Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop 6 upvotes, #16 of 2026-08-13
  20. MBA: Multimodal Benchmark and Agents for Real-World Business Ideation 6 upvotes, #16 of 2026-08-13
  21. Gaze Target Estimation Anywhere with Concepts 5 upvotes, #21 of 2026-08-13
  22. Agent Safety Should Be a Runtime Contract 4 upvotes, #22 of 2026-08-13
  23. NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs 3 upvotes, #23 of 2026-08-13
  24. ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization 3 upvotes, #23 of 2026-08-13
  25. AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models 2 upvotes, #25 of 2026-08-13
  26. Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands 2 upvotes, #25 of 2026-08-13
  27. Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control 2 upvotes, #25 of 2026-08-13
  28. From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options 2 upvotes, #25 of 2026-08-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.