Daily Papers of 2026-06-08
- Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models 119 upvotes, #1 of 2026-06-08
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings 91 upvotes, #2 of 2026-06-08
- ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research 87 upvotes, #3 of 2026-06-08
- SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations 52 upvotes, #4 of 2026-06-08
- GENEB: Why Genomic Models Are Hard to Compare 46 upvotes, #5 of 2026-06-08
- MMAE: A Massive Multitask Audio Editing Benchmark 44 upvotes, #6 of 2026-06-08
- AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization 29 upvotes, #7 of 2026-06-08
- Robots Need More than VLA and World Models 28 upvotes, #8 of 2026-06-08
- OpenSkill: Open-World Self-Evolution for LLM Agents 27 upvotes, #9 of 2026-06-08
- Direct 3D-Aware Object Insertion via Decomposed Visual Proxies 26 upvotes, #10 of 2026-06-08
- When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents 23 upvotes, #11 of 2026-06-08
- UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs 21 upvotes, #12 of 2026-06-08
- Watch, Remember, Reason: Human-View Video Understanding with MLLMs 21 upvotes, #12 of 2026-06-08
- SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents 19 upvotes, #14 of 2026-06-08
- LLM Explainability with Counterfactual Chains and Causal Graphs 17 upvotes, #15 of 2026-06-08
- LIMMT: Less is More for Motion Tracking 16 upvotes, #16 of 2026-06-08
- Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators 15 upvotes, #17 of 2026-06-08
- dots.tts Technical Report 15 upvotes, #17 of 2026-06-08
- SIA: Self Improving AI with Harness & Weight Updates 14 upvotes, #19 of 2026-06-08
- Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them 14 upvotes, #19 of 2026-06-08
- UniSHARP: Universal Sharp Monocular View Synthesis 14 upvotes, #19 of 2026-06-08
- PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams 13 upvotes, #22 of 2026-06-08
- Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills 12 upvotes, #23 of 2026-06-08
- Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models 7 upvotes, #24 of 2026-06-08
- SPACENUM: Revisiting Spatial Numerical Understanding in VLMs 6 upvotes, #25 of 2026-06-08
- HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems 6 upvotes, #25 of 2026-06-08
- CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning 4 upvotes, #27 of 2026-06-08
- A Cookbook of 3D Vision: Data, Learning Paradigms, and Application 4 upvotes, #27 of 2026-06-08
- Towards Retrieving Interaction Spaces for Agentic Search 4 upvotes, #27 of 2026-06-08
- Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors 4 upvotes, #27 of 2026-06-08
- Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development 4 upvotes, #27 of 2026-06-08
- When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges 3 upvotes, #32 of 2026-06-08
- Reinforcement Learning from Rich Feedback with Distributional DAgger 3 upvotes, #32 of 2026-06-08
- WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark 3 upvotes, #32 of 2026-06-08
- LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models 2 upvotes, #35 of 2026-06-08
- Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation 2 upvotes, #35 of 2026-06-08
- Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation 2 upvotes, #35 of 2026-06-08
- TBD-VLA: Temporal Block Diffusion Vision Language Action Model 2 upvotes, #35 of 2026-06-08
- Parametric Social Identity Injection and Diversification in Public Opinion Simulation 1 upvotes, #39 of 2026-06-08
- ECI_{sem}: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives 1 upvotes, #39 of 2026-06-08
- The Distillation Game: Adaptive Attacks & Efficient Defenses 1 upvotes, #39 of 2026-06-08
- Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation 1 upvotes, #39 of 2026-06-08
- Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms 1 upvotes, #39 of 2026-06-08
- Empirical Study on the Characteristics and Evolution of AI-usage in GitHub Repositories: Evidence from Code Comments 1 upvotes, #39 of 2026-06-08
- Streaming Video Generation with Streaming Force Control 1 upvotes, #39 of 2026-06-08
- Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity 0 upvotes, #46 of 2026-06-08
- Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback 1 upvotes, #46 of 2026-06-08
- How Far Can Chord-Symbol Time-Series Adaptation Carry Genre Identity? Capabilities and Boundaries in Multi-Genre Chord-Symbol Modeling 1 upvotes, #46 of 2026-06-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.