Daily Papers of 2026-06-25

  1. Are We Ready For An Agent-Native Memory System? 123 upvotes, #1 of 2026-06-25
  2. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models 111 upvotes, #2 of 2026-06-25
  3. DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation 67 upvotes, #3 of 2026-06-25
  4. ShutterMuse: Capture-Time Photography Guidance with MLLMs 46 upvotes, #4 of 2026-06-25
  5. Improved Large Language Diffusion Models 43 upvotes, #5 of 2026-06-25
  6. Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence 38 upvotes, #6 of 2026-06-25
  7. MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation 35 upvotes, #7 of 2026-06-25
  8. V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning 27 upvotes, #8 of 2026-06-25
  9. UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating 26 upvotes, #9 of 2026-06-25
  10. Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models 24 upvotes, #10 of 2026-06-25
  11. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems 18 upvotes, #11 of 2026-06-25
  12. Autodata: An agentic data scientist to create high quality synthetic data 18 upvotes, #11 of 2026-06-25
  13. IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation 17 upvotes, #13 of 2026-06-25
  14. EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies 16 upvotes, #14 of 2026-06-25
  15. Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do 9 upvotes, #15 of 2026-06-25
  16. RoPE-Aware Bit Allocation for KV-Cache Quantization 8 upvotes, #16 of 2026-06-25
  17. RL-Index: Reinforcement Learning for Retrieval Index Reasoning 6 upvotes, #17 of 2026-06-25
  18. Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods 6 upvotes, #17 of 2026-06-25
  19. TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy 6 upvotes, #17 of 2026-06-25
  20. ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation 5 upvotes, #20 of 2026-06-25
  21. CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression 5 upvotes, #20 of 2026-06-25
  22. What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics 5 upvotes, #20 of 2026-06-25
  23. When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents 4 upvotes, #23 of 2026-06-25
  24. GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods 3 upvotes, #24 of 2026-06-25
  25. PrivacyAlign: Contextual Privacy Alignment for LLM Agents 3 upvotes, #24 of 2026-06-25
  26. Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching 3 upvotes, #24 of 2026-06-25
  27. Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints 3 upvotes, #24 of 2026-06-25
  28. Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation 2 upvotes, #28 of 2026-06-25
  29. Forecasting Future Behavior as a Learning Task 1 upvotes, #29 of 2026-06-25
  30. Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents 1 upvotes, #29 of 2026-06-25
  31. Do Thinking Tokens Help with Safety? 1 upvotes, #29 of 2026-06-25
  32. Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation 1 upvotes, #29 of 2026-06-25
  33. Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach 0 upvotes, #33 of 2026-06-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.