Daily Papers of 2026-06-04

  1. Cosmos 3: Omnimodal World Models for Physical AI 115 upvotes, #1 of 2026-06-04
  2. Audio Interaction Model 108 upvotes, #2 of 2026-06-04
  3. Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories 54 upvotes, #3 of 2026-06-04
  4. Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning 37 upvotes, #4 of 2026-06-04
  5. Qwen-Image-Flash: Beyond Objective Design 35 upvotes, #5 of 2026-06-04
  6. OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs 31 upvotes, #6 of 2026-06-04
  7. AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? 29 upvotes, #7 of 2026-06-04
  8. Streaming Communication in Multi-Agent Reasoning 29 upvotes, #7 of 2026-06-04
  9. Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation 28 upvotes, #9 of 2026-06-04
  10. M^3Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks 26 upvotes, #10 of 2026-06-04
  11. Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems 25 upvotes, #11 of 2026-06-04
  12. ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning 25 upvotes, #11 of 2026-06-04
  13. Self-Distilled Policy Gradient 24 upvotes, #13 of 2026-06-04
  14. ZipSplat: Fewer Gaussians, Better Splats 20 upvotes, #14 of 2026-06-04
  15. KletterMix: Climbing Toward High-Quality German Pretraining Data 18 upvotes, #15 of 2026-06-04
  16. MemTrain: Self-Supervised Context Memory Training 17 upvotes, #16 of 2026-06-04
  17. MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation 17 upvotes, #16 of 2026-06-04
  18. Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation 16 upvotes, #18 of 2026-06-04
  19. Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching 16 upvotes, #18 of 2026-06-04
  20. MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills? 14 upvotes, #20 of 2026-06-04
  21. AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation 14 upvotes, #20 of 2026-06-04
  22. Large Language Models Hack Rewards, and Society 10 upvotes, #22 of 2026-06-04
  23. Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions 8 upvotes, #23 of 2026-06-04
  24. WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts 8 upvotes, #23 of 2026-06-04
  25. GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors 8 upvotes, #23 of 2026-06-04
  26. Neural Networks Provably Learn Spectral Representations for Group Composition 6 upvotes, #26 of 2026-06-04
  27. AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification 6 upvotes, #26 of 2026-06-04
  28. Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning 5 upvotes, #28 of 2026-06-04
  29. BraveGuard: From Open-World Threats to Safer Computer-Use Agents 5 upvotes, #28 of 2026-06-04
  30. BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution 5 upvotes, #28 of 2026-06-04
  31. MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation 5 upvotes, #28 of 2026-06-04
  32. Access Sets Matter: Budgeting Expert Reads for Scalable Weight-Space Model Merging 4 upvotes, #32 of 2026-06-04
  33. DAR: Deontic Reasoning with Agentic Harnesses 4 upvotes, #32 of 2026-06-04
  34. OpenSTBench: Beyond Semantic Evaluation for Speech Translation 3 upvotes, #34 of 2026-06-04
  35. SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes 3 upvotes, #34 of 2026-06-04
  36. PaintBench: Deterministic Evaluation of Precise Visual Editing 3 upvotes, #34 of 2026-06-04
  37. Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents 3 upvotes, #34 of 2026-06-04
  38. Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases 3 upvotes, #34 of 2026-06-04
  39. STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations 3 upvotes, #34 of 2026-06-04
  40. Score-Control for Hallucination Reduction in Diffusion Models 2 upvotes, #40 of 2026-06-04
  41. When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models 2 upvotes, #40 of 2026-06-04
  42. Unlocking Feature Learning in Gated Delta Networks at Scale 2 upvotes, #40 of 2026-06-04
  43. Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain 1 upvotes, #43 of 2026-06-04
  44. SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory 1 upvotes, #43 of 2026-06-04
  45. Measuring the Symmetry--Data Exchange Rate 1 upvotes, #43 of 2026-06-04
  46. SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing 1 upvotes, #43 of 2026-06-04
  47. Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning 1 upvotes, #43 of 2026-06-04
  48. Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game 1 upvotes, #43 of 2026-06-04
  49. Deep Embedded Multiplicative DMD for Algebra-Preserving Koopman Learning 1 upvotes, #43 of 2026-06-04
  50. Scalable Inference-Time Annealing with Surrogate Likelihood Estimators 1 upvotes, #50 of 2026-06-04
  51. Functional Attention: From Pairwise Affinities to Functional Correspondences 2 upvotes, #50 of 2026-06-04
  52. Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs 1 upvotes, #50 of 2026-06-04
  53. Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting 0 upvotes, #50 of 2026-06-04
  54. Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Study 0 upvotes, #50 of 2026-06-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.