Daily Papers of 2026-07-31

  1. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents 302 upvotes, #1 of 2026-07-31
  2. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis 298 upvotes, #2 of 2026-07-31
  3. Metis: Memory Foundation Model 270 upvotes, #3 of 2026-07-31
  4. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering 181 upvotes, #4 of 2026-07-31
  5. PhiZero: A World Model Built Around Physical Language 167 upvotes, #5 of 2026-07-31
  6. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System 70 upvotes, #6 of 2026-07-31
  7. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory 58 upvotes, #7 of 2026-07-31
  8. Beacon: Knowing When and How to Perform Agentic Visual Reasoning 54 upvotes, #8 of 2026-07-31
  9. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms 50 upvotes, #9 of 2026-07-31
  10. Flux-OPD: On-Policy Distillation with Evolving Contexts 44 upvotes, #10 of 2026-07-31
  11. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine 44 upvotes, #10 of 2026-07-31
  12. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing 39 upvotes, #12 of 2026-07-31
  13. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation 34 upvotes, #13 of 2026-07-31
  14. RefCaptioner: Multi-Reference Image-Grounded Video Captioning 30 upvotes, #14 of 2026-07-31
  15. See2Think: Do Multimodal Models Really Use Intermediate Visual States? 25 upvotes, #15 of 2026-07-31
  16. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them 24 upvotes, #16 of 2026-07-31
  17. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation 23 upvotes, #17 of 2026-07-31
  18. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers 21 upvotes, #18 of 2026-07-31
  19. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow 20 upvotes, #19 of 2026-07-31
  20. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation 19 upvotes, #20 of 2026-07-31
  21. Can Large Language Models Execute Parent Orders? 18 upvotes, #21 of 2026-07-31
  22. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems 17 upvotes, #22 of 2026-07-31
  23. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models 16 upvotes, #23 of 2026-07-31
  24. MemHarness: Memory Is Reconstructed, Not Replayed 16 upvotes, #23 of 2026-07-31
  25. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger 14 upvotes, #25 of 2026-07-31
  26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability 13 upvotes, #26 of 2026-07-31
  27. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale 13 upvotes, #26 of 2026-07-31
  28. Multi-Head Attention Residuals 11 upvotes, #28 of 2026-07-31
  29. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval 9 upvotes, #29 of 2026-07-31
  30. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models 7 upvotes, #30 of 2026-07-31
  31. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes 7 upvotes, #30 of 2026-07-31
  32. AI Tour Meeting: Group Travel Planning by LLM Agents 6 upvotes, #32 of 2026-07-31
  33. Harness-G: A Graph-Structured Harness for Search Agents 6 upvotes, #32 of 2026-07-31
  34. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing 5 upvotes, #34 of 2026-07-31
  35. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions 4 upvotes, #35 of 2026-07-31
  36. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations 4 upvotes, #35 of 2026-07-31
  37. Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing 2 upvotes, #37 of 2026-07-31
  38. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition 2 upvotes, #37 of 2026-07-31

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.