Daily Papers of 2026-06-04
- Cosmos 3: Omnimodal World Models for Physical AI 115 upvotes, #1 of 2026-06-04
- Audio Interaction Model 108 upvotes, #2 of 2026-06-04
- Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories 54 upvotes, #3 of 2026-06-04
- Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning 37 upvotes, #4 of 2026-06-04
- Qwen-Image-Flash: Beyond Objective Design 35 upvotes, #5 of 2026-06-04
- OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs 31 upvotes, #6 of 2026-06-04
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? 29 upvotes, #7 of 2026-06-04
- Streaming Communication in Multi-Agent Reasoning 29 upvotes, #7 of 2026-06-04
- Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation 28 upvotes, #9 of 2026-06-04
- M^3Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks 26 upvotes, #10 of 2026-06-04
- Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems 25 upvotes, #11 of 2026-06-04
- ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning 25 upvotes, #11 of 2026-06-04
- Self-Distilled Policy Gradient 24 upvotes, #13 of 2026-06-04
- ZipSplat: Fewer Gaussians, Better Splats 20 upvotes, #14 of 2026-06-04
- KletterMix: Climbing Toward High-Quality German Pretraining Data 18 upvotes, #15 of 2026-06-04
- MemTrain: Self-Supervised Context Memory Training 17 upvotes, #16 of 2026-06-04
- MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation 17 upvotes, #16 of 2026-06-04
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation 16 upvotes, #18 of 2026-06-04
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching 16 upvotes, #18 of 2026-06-04
- MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills? 14 upvotes, #20 of 2026-06-04
- AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation 14 upvotes, #20 of 2026-06-04
- Large Language Models Hack Rewards, and Society 10 upvotes, #22 of 2026-06-04
- Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions 8 upvotes, #23 of 2026-06-04
- WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts 8 upvotes, #23 of 2026-06-04
- GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors 8 upvotes, #23 of 2026-06-04
- Neural Networks Provably Learn Spectral Representations for Group Composition 6 upvotes, #26 of 2026-06-04
- AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification 6 upvotes, #26 of 2026-06-04
- Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning 5 upvotes, #28 of 2026-06-04
- BraveGuard: From Open-World Threats to Safer Computer-Use Agents 5 upvotes, #28 of 2026-06-04
- BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution 5 upvotes, #28 of 2026-06-04
- MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation 5 upvotes, #28 of 2026-06-04
- Access Sets Matter: Budgeting Expert Reads for Scalable Weight-Space Model Merging 4 upvotes, #32 of 2026-06-04
- DAR: Deontic Reasoning with Agentic Harnesses 4 upvotes, #32 of 2026-06-04
- OpenSTBench: Beyond Semantic Evaluation for Speech Translation 3 upvotes, #34 of 2026-06-04
- SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes 3 upvotes, #34 of 2026-06-04
- PaintBench: Deterministic Evaluation of Precise Visual Editing 3 upvotes, #34 of 2026-06-04
- Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents 3 upvotes, #34 of 2026-06-04
- Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases 3 upvotes, #34 of 2026-06-04
- STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations 3 upvotes, #34 of 2026-06-04
- Score-Control for Hallucination Reduction in Diffusion Models 2 upvotes, #40 of 2026-06-04
- When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models 2 upvotes, #40 of 2026-06-04
- Unlocking Feature Learning in Gated Delta Networks at Scale 2 upvotes, #40 of 2026-06-04
- Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain 1 upvotes, #43 of 2026-06-04
- SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory 1 upvotes, #43 of 2026-06-04
- Measuring the Symmetry--Data Exchange Rate 1 upvotes, #43 of 2026-06-04
- SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing 1 upvotes, #43 of 2026-06-04
- Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning 1 upvotes, #43 of 2026-06-04
- Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game 1 upvotes, #43 of 2026-06-04
- Deep Embedded Multiplicative DMD for Algebra-Preserving Koopman Learning 1 upvotes, #43 of 2026-06-04
- Scalable Inference-Time Annealing with Surrogate Likelihood Estimators 1 upvotes, #50 of 2026-06-04
- Functional Attention: From Pairwise Affinities to Functional Correspondences 2 upvotes, #50 of 2026-06-04
- Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs 1 upvotes, #50 of 2026-06-04
- Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting 0 upvotes, #50 of 2026-06-04
- Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Study 0 upvotes, #50 of 2026-06-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.