Daily Papers of 2026-06-08

  1. Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models 119 upvotes, #1 of 2026-06-08
  2. Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings 91 upvotes, #2 of 2026-06-08
  3. ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research 87 upvotes, #3 of 2026-06-08
  4. SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations 52 upvotes, #4 of 2026-06-08
  5. GENEB: Why Genomic Models Are Hard to Compare 46 upvotes, #5 of 2026-06-08
  6. MMAE: A Massive Multitask Audio Editing Benchmark 44 upvotes, #6 of 2026-06-08
  7. AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization 29 upvotes, #7 of 2026-06-08
  8. Robots Need More than VLA and World Models 28 upvotes, #8 of 2026-06-08
  9. OpenSkill: Open-World Self-Evolution for LLM Agents 27 upvotes, #9 of 2026-06-08
  10. Direct 3D-Aware Object Insertion via Decomposed Visual Proxies 26 upvotes, #10 of 2026-06-08
  11. When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents 23 upvotes, #11 of 2026-06-08
  12. UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs 21 upvotes, #12 of 2026-06-08
  13. Watch, Remember, Reason: Human-View Video Understanding with MLLMs 21 upvotes, #12 of 2026-06-08
  14. SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents 19 upvotes, #14 of 2026-06-08
  15. LLM Explainability with Counterfactual Chains and Causal Graphs 17 upvotes, #15 of 2026-06-08
  16. LIMMT: Less is More for Motion Tracking 16 upvotes, #16 of 2026-06-08
  17. Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators 15 upvotes, #17 of 2026-06-08
  18. dots.tts Technical Report 15 upvotes, #17 of 2026-06-08
  19. SIA: Self Improving AI with Harness & Weight Updates 14 upvotes, #19 of 2026-06-08
  20. Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them 14 upvotes, #19 of 2026-06-08
  21. UniSHARP: Universal Sharp Monocular View Synthesis 14 upvotes, #19 of 2026-06-08
  22. PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams 13 upvotes, #22 of 2026-06-08
  23. Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills 12 upvotes, #23 of 2026-06-08
  24. Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models 7 upvotes, #24 of 2026-06-08
  25. SPACENUM: Revisiting Spatial Numerical Understanding in VLMs 6 upvotes, #25 of 2026-06-08
  26. HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems 6 upvotes, #25 of 2026-06-08
  27. CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning 4 upvotes, #27 of 2026-06-08
  28. A Cookbook of 3D Vision: Data, Learning Paradigms, and Application 4 upvotes, #27 of 2026-06-08
  29. Towards Retrieving Interaction Spaces for Agentic Search 4 upvotes, #27 of 2026-06-08
  30. Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors 4 upvotes, #27 of 2026-06-08
  31. Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development 4 upvotes, #27 of 2026-06-08
  32. When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges 3 upvotes, #32 of 2026-06-08
  33. Reinforcement Learning from Rich Feedback with Distributional DAgger 3 upvotes, #32 of 2026-06-08
  34. WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark 3 upvotes, #32 of 2026-06-08
  35. LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models 2 upvotes, #35 of 2026-06-08
  36. Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation 2 upvotes, #35 of 2026-06-08
  37. Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation 2 upvotes, #35 of 2026-06-08
  38. TBD-VLA: Temporal Block Diffusion Vision Language Action Model 2 upvotes, #35 of 2026-06-08
  39. Parametric Social Identity Injection and Diversification in Public Opinion Simulation 1 upvotes, #39 of 2026-06-08
  40. ECI_{sem}: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives 1 upvotes, #39 of 2026-06-08
  41. The Distillation Game: Adaptive Attacks & Efficient Defenses 1 upvotes, #39 of 2026-06-08
  42. Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation 1 upvotes, #39 of 2026-06-08
  43. Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms 1 upvotes, #39 of 2026-06-08
  44. Empirical Study on the Characteristics and Evolution of AI-usage in GitHub Repositories: Evidence from Code Comments 1 upvotes, #39 of 2026-06-08
  45. Streaming Video Generation with Streaming Force Control 1 upvotes, #39 of 2026-06-08
  46. Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity 0 upvotes, #46 of 2026-06-08
  47. Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback 1 upvotes, #46 of 2026-06-08
  48. How Far Can Chord-Symbol Time-Series Adaptation Carry Genre Identity? Capabilities and Boundaries in Multi-Genre Chord-Symbol Modeling 1 upvotes, #46 of 2026-06-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.