Daily Papers of 2026-07-03

  1. Program-as-Weights: A Programming Paradigm for Fuzzy Functions 118 upvotes, #1 of 2026-07-03
  2. AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents 60 upvotes, #2 of 2026-07-03
  3. EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments 50 upvotes, #3 of 2026-07-03
  4. Morphing into Hybrid Attention Models 47 upvotes, #4 of 2026-07-03
  5. Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling 37 upvotes, #5 of 2026-07-03
  6. AgenticDataBench: A Comprehensive Benchmark for Data Agents 35 upvotes, #6 of 2026-07-03
  7. WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory 32 upvotes, #7 of 2026-07-03
  8. Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning 26 upvotes, #8 of 2026-07-03
  9. AutoMem: Automated Learning of Memory as a Cognitive Skill 19 upvotes, #9 of 2026-07-03
  10. Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads 18 upvotes, #10 of 2026-07-03
  11. SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use 18 upvotes, #10 of 2026-07-03
  12. PACE: A Proxy for Agentic Capability Evaluation 18 upvotes, #10 of 2026-07-03
  13. AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition 16 upvotes, #13 of 2026-07-03
  14. Optimizing Visual Generative Models via Distribution-wise Rewards 16 upvotes, #13 of 2026-07-03
  15. When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search 15 upvotes, #15 of 2026-07-03
  16. InstanceControl: Controllable Complex Image Generation without Instance Labeling 14 upvotes, #16 of 2026-07-03
  17. From SRA to Self-Flow: Data Augmentation or Self-Supervision? 14 upvotes, #16 of 2026-07-03
  18. Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs 11 upvotes, #18 of 2026-07-03
  19. DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation 10 upvotes, #19 of 2026-07-03
  20. Discrete Diffusion Language Models for Interactive Radiology Report Drafting 10 upvotes, #19 of 2026-07-03
  21. Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training 10 upvotes, #19 of 2026-07-03
  22. Representation Distribution Matching for One-Step Visual Generation 10 upvotes, #19 of 2026-07-03
  23. WARP: Weight-Space Analysis for Recovering Training Data Portfolios 9 upvotes, #23 of 2026-07-03
  24. AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models 9 upvotes, #23 of 2026-07-03
  25. Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR 7 upvotes, #25 of 2026-07-03
  26. Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting 7 upvotes, #25 of 2026-07-03
  27. Scaling Laws for Grid-Based Approximate Nearest Neighbor Search in High Dimensions 4 upvotes, #27 of 2026-07-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.