Daily Papers of 2026-06-01
- COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation 109 upvotes, #1 of 2026-06-01
- GrepSeek: Training Search Agents for Direct Corpus Interaction 104 upvotes, #2 of 2026-06-01
- Trust-Region Behavior Blending for On-Policy Distillation 65 upvotes, #3 of 2026-06-01
- Representation Forcing for Bottleneck-Free Unified Multimodal Models 59 upvotes, #4 of 2026-06-01
- SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue 56 upvotes, #5 of 2026-06-01
- Mellum2 Technical Report 53 upvotes, #6 of 2026-06-01
- GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration 43 upvotes, #7 of 2026-06-01
- Function2Scene: 3D Indoor Scene Layout from Functional Specifications 41 upvotes, #8 of 2026-06-01
- LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards 41 upvotes, #8 of 2026-06-01
- Task-Focused Memorization for Multimodal Agents 38 upvotes, #10 of 2026-06-01
- Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer 37 upvotes, #11 of 2026-06-01
- SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer 36 upvotes, #12 of 2026-06-01
- dMoE: dLLMs with Learnable Block Experts 36 upvotes, #12 of 2026-06-01
- Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios 31 upvotes, #14 of 2026-06-01
- SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks 28 upvotes, #15 of 2026-06-01
- Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation 26 upvotes, #16 of 2026-06-01
- VLM3: Vision Language Models Are Native 3D Learners 26 upvotes, #16 of 2026-06-01
- SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search 25 upvotes, #18 of 2026-06-01
- Exploring Autonomous Agentic Data Engineering for Model Specialization 23 upvotes, #19 of 2026-06-01
- LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis 21 upvotes, #20 of 2026-06-01
- Linearizing Vision Transformer with Test-Time Training 20 upvotes, #21 of 2026-06-01
- Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents 20 upvotes, #21 of 2026-06-01
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents 20 upvotes, #21 of 2026-06-01
- PEEK: Picking Essential frames via Efficient Knowledge distillation 20 upvotes, #21 of 2026-06-01
- From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors 18 upvotes, #25 of 2026-06-01
- HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range Videos 13 upvotes, #26 of 2026-06-01
- Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? 12 upvotes, #27 of 2026-06-01
- DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory 12 upvotes, #27 of 2026-06-01
- DEMON: Diffusion Engine for Musical Orchestrated Noise 11 upvotes, #29 of 2026-06-01
- Count Anything 11 upvotes, #29 of 2026-06-01
- Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion 11 upvotes, #29 of 2026-06-01
- Linear Scaling Video VLMs for Long Video Understanding 11 upvotes, #29 of 2026-06-01
- VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies 10 upvotes, #33 of 2026-06-01
- The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement 10 upvotes, #33 of 2026-06-01
- Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring 9 upvotes, #35 of 2026-06-01
- OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents 8 upvotes, #36 of 2026-06-01
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI 8 upvotes, #36 of 2026-06-01
- One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation 8 upvotes, #36 of 2026-06-01
- SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? 8 upvotes, #36 of 2026-06-01
- How can embedding models bind concepts? 8 upvotes, #36 of 2026-06-01
- Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models 8 upvotes, #36 of 2026-06-01
- One-Forcing: Towards Stable One-Step Autoregressive Video Generation 7 upvotes, #42 of 2026-06-01
- AlphaTransit: Learning to Design City-scale Transit Routes 7 upvotes, #42 of 2026-06-01
- FRAPPE: Full Input, Residual Output Autoencoding with Projection Pursuit Encoder 7 upvotes, #42 of 2026-06-01
- GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models 7 upvotes, #42 of 2026-06-01
- MAAT: Multi-phase Adapter-Aware Targeted Unlearning 7 upvotes, #42 of 2026-06-01
- iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning 7 upvotes, #42 of 2026-06-01
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video 7 upvotes, #42 of 2026-06-01
- DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization 6 upvotes, #49 of 2026-06-01
- SurGe: Improved Surface Geometry in Point Maps 6 upvotes, #49 of 2026-06-01
- Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly 3 upvotes, #51 of 2026-06-01
- When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models 3 upvotes, #51 of 2026-06-01
- Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization 2 upvotes, #53 of 2026-06-01
- The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction 2 upvotes, #53 of 2026-06-01
- Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models 2 upvotes, #53 of 2026-06-01
- Benchmarking Composed Image Retrieval for Applied Earth Observation 1 upvotes, #56 of 2026-06-01
- Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization 1 upvotes, #56 of 2026-06-01
- Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode 1 upvotes, #56 of 2026-06-01
- A Topology-Aware Spatiotemporal Handover Framework for Continuous Multi-UAV Tracking 0 upvotes, #59 of 2026-06-01
- Beyond Holistic Models: Systematic Component-level Benchmarking of Deep Multivariate Time-Series Forecasting 1 upvotes, #59 of 2026-06-01
- Frequency-Guided Action Diffusion via Sub-Frequency Manifold Traversal 0 upvotes, #59 of 2026-06-01
- AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling 2 upvotes, #59 of 2026-06-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.