Daily Papers of 2026-08-11
- BDH-CQ: In-Context Learning with Recurrent Latent Reasoning 659 upvotes, #1 of 2026-08-11
- Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA 334 upvotes, #2 of 2026-08-11
- On-Policy Self-Distillation without Any Supervision 210 upvotes, #3 of 2026-08-11
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 129 upvotes, #4 of 2026-08-11
- Stealing Reasoning Traces from Proprietary LLM APIs 107 upvotes, #5 of 2026-08-11
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution 86 upvotes, #6 of 2026-08-11
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory 49 upvotes, #7 of 2026-08-11
- Motif 3: Technical Report 43 upvotes, #8 of 2026-08-11
- Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains 29 upvotes, #9 of 2026-08-11
- SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation 28 upvotes, #10 of 2026-08-11
- Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation 28 upvotes, #10 of 2026-08-11
- What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems 28 upvotes, #10 of 2026-08-11
- OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching 25 upvotes, #13 of 2026-08-11
- Scaling Inherently Interpretable Language Models 21 upvotes, #14 of 2026-08-11
- Business Arena: Benchmarking LLM Agents in a Realistic Marketplace 21 upvotes, #14 of 2026-08-11
- Evo-Bench: Can Language Models Improve Agent Harness? 18 upvotes, #16 of 2026-08-11
- Evidence-RL: Towards Evidence-intensive Visual Reasoning 16 upvotes, #17 of 2026-08-11
- RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance 14 upvotes, #18 of 2026-08-11
- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States 13 upvotes, #19 of 2026-08-11
- An End-to-End Agent Auditing Engine 13 upvotes, #19 of 2026-08-11
- A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization 9 upvotes, #21 of 2026-08-11
- WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks 8 upvotes, #22 of 2026-08-11
- The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents 8 upvotes, #22 of 2026-08-11
- Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization 8 upvotes, #22 of 2026-08-11
- Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation 8 upvotes, #22 of 2026-08-11
- Omega-S: A Functional Resilience Index for LLM Fine-Tuning 7 upvotes, #26 of 2026-08-11
- The Loss Does Not See the Basis, but Adam Does 7 upvotes, #26 of 2026-08-11
- Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval 7 upvotes, #26 of 2026-08-11
- Vision-Language Grounding as Bidirectional Concept Correspondence 7 upvotes, #26 of 2026-08-11
- Ego-OSCAR: Egocentric Open source Stereo CAptuRe System 7 upvotes, #26 of 2026-08-11
- VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use 7 upvotes, #26 of 2026-08-11
- Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure 7 upvotes, #26 of 2026-08-11
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models 6 upvotes, #33 of 2026-08-11
- Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers 6 upvotes, #33 of 2026-08-11
- MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation 6 upvotes, #33 of 2026-08-11
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification 6 upvotes, #33 of 2026-08-11
- Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness 6 upvotes, #33 of 2026-08-11
- CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems 6 upvotes, #33 of 2026-08-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.