Daily Papers of 2026-06-25
- Are We Ready For An Agent-Native Memory System? 123 upvotes, #1 of 2026-06-25
- Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models 111 upvotes, #2 of 2026-06-25
- DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation 67 upvotes, #3 of 2026-06-25
- ShutterMuse: Capture-Time Photography Guidance with MLLMs 46 upvotes, #4 of 2026-06-25
- Improved Large Language Diffusion Models 43 upvotes, #5 of 2026-06-25
- Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence 38 upvotes, #6 of 2026-06-25
- MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation 35 upvotes, #7 of 2026-06-25
- V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning 27 upvotes, #8 of 2026-06-25
- UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating 26 upvotes, #9 of 2026-06-25
- Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models 24 upvotes, #10 of 2026-06-25
- The Hitchhiker's Guide to Agentic AI: From Foundations to Systems 18 upvotes, #11 of 2026-06-25
- Autodata: An agentic data scientist to create high quality synthetic data 18 upvotes, #11 of 2026-06-25
- IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation 17 upvotes, #13 of 2026-06-25
- EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies 16 upvotes, #14 of 2026-06-25
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do 9 upvotes, #15 of 2026-06-25
- RoPE-Aware Bit Allocation for KV-Cache Quantization 8 upvotes, #16 of 2026-06-25
- RL-Index: Reinforcement Learning for Retrieval Index Reasoning 6 upvotes, #17 of 2026-06-25
- Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods 6 upvotes, #17 of 2026-06-25
- TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy 6 upvotes, #17 of 2026-06-25
- ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation 5 upvotes, #20 of 2026-06-25
- CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression 5 upvotes, #20 of 2026-06-25
- What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics 5 upvotes, #20 of 2026-06-25
- When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents 4 upvotes, #23 of 2026-06-25
- GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods 3 upvotes, #24 of 2026-06-25
- PrivacyAlign: Contextual Privacy Alignment for LLM Agents 3 upvotes, #24 of 2026-06-25
- Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching 3 upvotes, #24 of 2026-06-25
- Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints 3 upvotes, #24 of 2026-06-25
- Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation 2 upvotes, #28 of 2026-06-25
- Forecasting Future Behavior as a Learning Task 1 upvotes, #29 of 2026-06-25
- Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents 1 upvotes, #29 of 2026-06-25
- Do Thinking Tokens Help with Safety? 1 upvotes, #29 of 2026-06-25
- Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation 1 upvotes, #29 of 2026-06-25
- Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach 0 upvotes, #33 of 2026-06-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.