Daily Papers of 2026-06-12

  1. MiniMax Sparse Attention 140 upvotes, #1 of 2026-06-12
  2. EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments 137 upvotes, #2 of 2026-06-12
  3. WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces 101 upvotes, #3 of 2026-06-12
  4. SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning 101 upvotes, #3 of 2026-06-12
  5. MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling 89 upvotes, #5 of 2026-06-12
  6. InterleaveThinker: Reinforcing Agentic Interleaved Generation 79 upvotes, #6 of 2026-06-12
  7. Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? 78 upvotes, #7 of 2026-06-12
  8. FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents 74 upvotes, #8 of 2026-06-12
  9. LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories 53 upvotes, #9 of 2026-06-12
  10. VIA-SD: Verification via Intra-Model Routing for Speculative Decoding 35 upvotes, #10 of 2026-06-12
  11. From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion 32 upvotes, #11 of 2026-06-12
  12. HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 28 upvotes, #12 of 2026-06-12
  13. EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery 27 upvotes, #13 of 2026-06-12
  14. N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization 24 upvotes, #14 of 2026-06-12
  15. Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning 21 upvotes, #15 of 2026-06-12
  16. VideoMDM: Towards 3D Human Motion Generation From 2D Supervision 20 upvotes, #16 of 2026-06-12
  17. Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback 15 upvotes, #17 of 2026-06-12
  18. MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold 13 upvotes, #18 of 2026-06-12
  19. HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness 12 upvotes, #19 of 2026-06-12
  20. High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation 11 upvotes, #20 of 2026-06-12
  21. TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Search 10 upvotes, #21 of 2026-06-12
  22. Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models 9 upvotes, #22 of 2026-06-12
  23. Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning 7 upvotes, #23 of 2026-06-12
  24. SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling 6 upvotes, #24 of 2026-06-12
  25. Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior 6 upvotes, #24 of 2026-06-12
  26. Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents 5 upvotes, #26 of 2026-06-12
  27. MuJoCo-Drones-Gym: A GPU-Accelerated Multi-Drone Simulator for Control and Reinforcement Learning 4 upvotes, #27 of 2026-06-12
  28. MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training 4 upvotes, #27 of 2026-06-12
  29. Flash-GMM: A Memory-Efficient Kernel for Scalable Soft Clustering 4 upvotes, #27 of 2026-06-12
  30. EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge 4 upvotes, #27 of 2026-06-12
  31. Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents 4 upvotes, #27 of 2026-06-12
  32. Surflo: Consistent 3D Surface Flow Model with Global State 4 upvotes, #27 of 2026-06-12
  33. See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents 3 upvotes, #33 of 2026-06-12
  34. WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation 3 upvotes, #33 of 2026-06-12
  35. The Cold-Start Safety Gap in LLM Agents 2 upvotes, #35 of 2026-06-12
  36. Revisiting Articulated Parts Perception in Robot Manipulation 2 upvotes, #35 of 2026-06-12
  37. WebChallenger: A Reliable and Efficient Generalist Web Agent 2 upvotes, #35 of 2026-06-12
  38. IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder 2 upvotes, #35 of 2026-06-12
  39. ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs 2 upvotes, #35 of 2026-06-12
  40. ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages 2 upvotes, #35 of 2026-06-12
  41. On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance 1 upvotes, #41 of 2026-06-12
  42. Leveraging Morphology for Historical Script Metrological Analysis 1 upvotes, #41 of 2026-06-12
  43. PianoKontext: Expressive Performance Rendering from Deadpan Context 1 upvotes, #41 of 2026-06-12
  44. A Stationary (and Therefore Compatible) Representation is All You Need 1 upvotes, #41 of 2026-06-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.