Daily Papers of 2026-08-11

  1. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning 659 upvotes, #1 of 2026-08-11
  2. Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA 334 upvotes, #2 of 2026-08-11
  3. On-Policy Self-Distillation without Any Supervision 210 upvotes, #3 of 2026-08-11
  4. SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 129 upvotes, #4 of 2026-08-11
  5. Stealing Reasoning Traces from Proprietary LLM APIs 107 upvotes, #5 of 2026-08-11
  6. Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution 86 upvotes, #6 of 2026-08-11
  7. Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory 49 upvotes, #7 of 2026-08-11
  8. Motif 3: Technical Report 43 upvotes, #8 of 2026-08-11
  9. Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains 29 upvotes, #9 of 2026-08-11
  10. SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation 28 upvotes, #10 of 2026-08-11
  11. Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation 28 upvotes, #10 of 2026-08-11
  12. What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems 28 upvotes, #10 of 2026-08-11
  13. OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching 25 upvotes, #13 of 2026-08-11
  14. Scaling Inherently Interpretable Language Models 21 upvotes, #14 of 2026-08-11
  15. Business Arena: Benchmarking LLM Agents in a Realistic Marketplace 21 upvotes, #14 of 2026-08-11
  16. Evo-Bench: Can Language Models Improve Agent Harness? 18 upvotes, #16 of 2026-08-11
  17. Evidence-RL: Towards Evidence-intensive Visual Reasoning 16 upvotes, #17 of 2026-08-11
  18. RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance 14 upvotes, #18 of 2026-08-11
  19. RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States 13 upvotes, #19 of 2026-08-11
  20. An End-to-End Agent Auditing Engine 13 upvotes, #19 of 2026-08-11
  21. A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization 9 upvotes, #21 of 2026-08-11
  22. WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks 8 upvotes, #22 of 2026-08-11
  23. The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents 8 upvotes, #22 of 2026-08-11
  24. Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization 8 upvotes, #22 of 2026-08-11
  25. Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation 8 upvotes, #22 of 2026-08-11
  26. Omega-S: A Functional Resilience Index for LLM Fine-Tuning 7 upvotes, #26 of 2026-08-11
  27. The Loss Does Not See the Basis, but Adam Does 7 upvotes, #26 of 2026-08-11
  28. Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval 7 upvotes, #26 of 2026-08-11
  29. Vision-Language Grounding as Bidirectional Concept Correspondence 7 upvotes, #26 of 2026-08-11
  30. Ego-OSCAR: Egocentric Open source Stereo CAptuRe System 7 upvotes, #26 of 2026-08-11
  31. VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use 7 upvotes, #26 of 2026-08-11
  32. Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure 7 upvotes, #26 of 2026-08-11
  33. MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models 6 upvotes, #33 of 2026-08-11
  34. Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers 6 upvotes, #33 of 2026-08-11
  35. MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation 6 upvotes, #33 of 2026-08-11
  36. SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification 6 upvotes, #33 of 2026-08-11
  37. Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness 6 upvotes, #33 of 2026-08-11
  38. CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems 6 upvotes, #33 of 2026-08-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.