Daily Papers of 2026-06-09

  1. Agents' Last Exam 345 upvotes, #1 of 2026-06-09
  2. SWE-Explore: Benchmarking How Coding Agents Explore Repositories 112 upvotes, #2 of 2026-06-09
  3. On the Geometry of On-Policy Distillation 71 upvotes, #3 of 2026-06-09
  4. Latent Spatial Memory for Video World Models 66 upvotes, #4 of 2026-06-09
  5. LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents 63 upvotes, #5 of 2026-06-09
  6. FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention 61 upvotes, #6 of 2026-06-09
  7. CoVEBench: Can Video Editing Models Handle Complex Instructions? 48 upvotes, #7 of 2026-06-09
  8. SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 42 upvotes, #8 of 2026-06-09
  9. Human Psychometric Questionnaires Mischaracterize LLM Behavior 35 upvotes, #9 of 2026-06-09
  10. Echo-Memory: A Controlled Study of Memory in Action World Models 32 upvotes, #10 of 2026-06-09
  11. End-to-End Context Compression at Scale 26 upvotes, #11 of 2026-06-09
  12. Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning 24 upvotes, #12 of 2026-06-09
  13. WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models 22 upvotes, #13 of 2026-06-09
  14. A Geometric Account of Activation Steering through Angle-Norm Decomposition 22 upvotes, #13 of 2026-06-09
  15. OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics 19 upvotes, #15 of 2026-06-09
  16. SwiftVR: Real-Time One-Step Generative Video Restoration 16 upvotes, #16 of 2026-06-09
  17. Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders 15 upvotes, #17 of 2026-06-09
  18. DEI: Diversity in Evolutionary Inference for Quality-Diversity Search 14 upvotes, #18 of 2026-06-09
  19. Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses 14 upvotes, #18 of 2026-06-09
  20. OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning 14 upvotes, #18 of 2026-06-09
  21. AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing 14 upvotes, #18 of 2026-06-09
  22. Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill 13 upvotes, #22 of 2026-06-09
  23. SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating 12 upvotes, #23 of 2026-06-09
  24. Text-to-Image Models Need Less from Text Encoders Than You Think 11 upvotes, #24 of 2026-06-09
  25. Why Muon Outperforms Adam: A Curvature Perspective 10 upvotes, #25 of 2026-06-09
  26. Light-WAM: Efficient World Action Models with State-Fusion Action Decoding 10 upvotes, #25 of 2026-06-09
  27. Answer Presence Drives RAG Rewriting Gains 8 upvotes, #27 of 2026-06-09
  28. Liberating LLM Capabilities in Full-Duplex Speech Models 8 upvotes, #27 of 2026-06-09
  29. Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short 8 upvotes, #27 of 2026-06-09
  30. Trajectory-Refined Distillation 7 upvotes, #30 of 2026-06-09
  31. Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher 6 upvotes, #31 of 2026-06-09
  32. Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory 6 upvotes, #31 of 2026-06-09
  33. DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning 6 upvotes, #31 of 2026-06-09
  34. VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation 5 upvotes, #34 of 2026-06-09
  35. AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents 4 upvotes, #35 of 2026-06-09
  36. Robotic Policy Adaptation via Weight-Space Meta-Learning 4 upvotes, #35 of 2026-06-09
  37. Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text 4 upvotes, #35 of 2026-06-09
  38. Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting 4 upvotes, #35 of 2026-06-09
  39. SDR: Set-Distance Rewards for Radiology Report Generation 3 upvotes, #39 of 2026-06-09
  40. Phase Marginalization for Patch-Grid Instability in Vision Transformers 3 upvotes, #39 of 2026-06-09
  41. Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory 3 upvotes, #39 of 2026-06-09
  42. Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path 2 upvotes, #42 of 2026-06-09
  43. OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation 2 upvotes, #42 of 2026-06-09
  44. Pruning and Distilling Mixture-of-Experts into Dense Language Models 1 upvotes, #44 of 2026-06-09
  45. Honest Lying: Understanding Memory Confabulation in Reflexive Agents 1 upvotes, #44 of 2026-06-09
  46. Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense 1 upvotes, #44 of 2026-06-09
  47. Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data 1 upvotes, #44 of 2026-06-09
  48. Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents 1 upvotes, #44 of 2026-06-09
  49. SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices 1 upvotes, #44 of 2026-06-09
  50. Chiaroscuro Attention: Spending Compute in the Dark 1 upvotes, #44 of 2026-06-09
  51. EmpiriGraph-Psy: A Dataset and LLM Pipeline for Extracting Empirical Relation Graphs from Psychology Abstracts 1 upvotes, #44 of 2026-06-09
  52. Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops 1 upvotes, #44 of 2026-06-09
  53. PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment 1 upvotes, #44 of 2026-06-09
  54. EMMA: Extracting Multiple physical parameters from Multimodal Data 0 upvotes, #54 of 2026-06-09
  55. Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation? 0 upvotes, #54 of 2026-06-09
  56. CIPER: A Unified Framework for Cross-view Image-retrieval and Pose-estimation 0 upvotes, #54 of 2026-06-09
  57. Set-Based Transformer for Atmospheric Compensation in Standoff LWIR Hyperspectral Imaging 0 upvotes, #54 of 2026-06-09
  58. PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems 1 upvotes, #54 of 2026-06-09
  59. Precision Is Not Faithfulness: Coverage-Aware Evaluation of Grounded Generation with a Complete Oracle 0 upvotes, #54 of 2026-06-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.