Daily Papers of 2026-07-01

  1. Dockerless: Environment-Free Program Verifier for Coding Agents 108 upvotes, #1 of 2026-07-01
  2. DOPD: Dual On-policy Distillation 103 upvotes, #2 of 2026-07-01
  3. Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models 76 upvotes, #3 of 2026-07-01
  4. BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding 76 upvotes, #3 of 2026-07-01
  5. Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views 52 upvotes, #5 of 2026-07-01
  6. SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History 43 upvotes, #6 of 2026-07-01
  7. Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks 41 upvotes, #7 of 2026-07-01
  8. Multi-Block Diffusion Language Models 38 upvotes, #8 of 2026-07-01
  9. GEAR: Guided End-to-End AutoRegression for Image Synthesis 34 upvotes, #9 of 2026-07-01
  10. DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation 27 upvotes, #10 of 2026-07-01
  11. MemLearner: Learning to Query Context memory for Video World Models 27 upvotes, #10 of 2026-07-01
  12. PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising 26 upvotes, #12 of 2026-07-01
  13. Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation 24 upvotes, #13 of 2026-07-01
  14. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs 24 upvotes, #13 of 2026-07-01
  15. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement 22 upvotes, #15 of 2026-07-01
  16. Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning 20 upvotes, #16 of 2026-07-01
  17. Xiaomi-GUI-0 Technical Report 19 upvotes, #17 of 2026-07-01
  18. Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? 16 upvotes, #18 of 2026-07-01
  19. RedVox: Safety and Fairness Gaps in Speech Models Across Languages 16 upvotes, #18 of 2026-07-01
  20. Little Brains, Big Feats: Exploring Compact Language Models 15 upvotes, #20 of 2026-07-01
  21. MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training 15 upvotes, #20 of 2026-07-01
  22. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning 13 upvotes, #22 of 2026-07-01
  23. PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation 12 upvotes, #23 of 2026-07-01
  24. QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents 12 upvotes, #23 of 2026-07-01
  25. Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature 11 upvotes, #25 of 2026-07-01
  26. BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language 10 upvotes, #26 of 2026-07-01
  27. Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing 9 upvotes, #27 of 2026-07-01
  28. AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation 9 upvotes, #27 of 2026-07-01
  29. LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents 7 upvotes, #29 of 2026-07-01
  30. SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE 7 upvotes, #29 of 2026-07-01
  31. SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions 6 upvotes, #31 of 2026-07-01
  32. TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive 6 upvotes, #31 of 2026-07-01
  33. MuSViT: A Foundation Vision Model for Sheet Music Representation 6 upvotes, #31 of 2026-07-01
  34. Hierarchical Experimentalist Agents 5 upvotes, #34 of 2026-07-01
  35. FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model 5 upvotes, #34 of 2026-07-01
  36. Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents 4 upvotes, #36 of 2026-07-01
  37. RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue 4 upvotes, #36 of 2026-07-01

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.