Daily Papers of 2026-06-23

  1. PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems 95 upvotes, #1 of 2026-06-23
  2. EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions 79 upvotes, #2 of 2026-06-23
  3. OpenRath: Session-Centered Runtime State for Agent Systems 77 upvotes, #3 of 2026-06-23
  4. Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention 76 upvotes, #4 of 2026-06-23
  5. DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams 73 upvotes, #5 of 2026-06-23
  6. World Action Models: A Survey 56 upvotes, #6 of 2026-06-23
  7. KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking 49 upvotes, #7 of 2026-06-23
  8. Unlimited OCR Works 43 upvotes, #8 of 2026-06-23
  9. CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents 37 upvotes, #9 of 2026-06-23
  10. EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory 32 upvotes, #10 of 2026-06-23
  11. Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding 24 upvotes, #11 of 2026-06-23
  12. BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language 24 upvotes, #11 of 2026-06-23
  13. SkillHarness: Harnessing Safe Skills for Computer-Use Agents 20 upvotes, #13 of 2026-06-23
  14. Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation 18 upvotes, #14 of 2026-06-23
  15. HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization 18 upvotes, #14 of 2026-06-23
  16. Self-Compacting Language Model Agents 18 upvotes, #14 of 2026-06-23
  17. Training Open Models for Agentic Phone Use 16 upvotes, #17 of 2026-06-23
  18. Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark 15 upvotes, #18 of 2026-06-23
  19. DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks 14 upvotes, #19 of 2026-06-23
  20. Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents 14 upvotes, #19 of 2026-06-23
  21. Tmax: A simple recipe for terminal agents 14 upvotes, #19 of 2026-06-23
  22. Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills 11 upvotes, #22 of 2026-06-23
  23. PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning 11 upvotes, #22 of 2026-06-23
  24. Exploring the Design Space of Reward Backpropagation for Flow Matching 10 upvotes, #24 of 2026-06-23
  25. Vera: A Layered Diffusion Model for Content-Preserving Video Editing 10 upvotes, #24 of 2026-06-23
  26. Tapered Language Models 9 upvotes, #26 of 2026-06-23
  27. Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning 8 upvotes, #27 of 2026-06-23
  28. Safe Few-Step Generation via Velocity Editing 8 upvotes, #27 of 2026-06-23
  29. Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning 7 upvotes, #29 of 2026-06-23
  30. Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs 7 upvotes, #29 of 2026-06-23
  31. Causal Discovery in the Era of Agents 7 upvotes, #29 of 2026-06-23
  32. Counsel: A Meta-Evaluation Dataset for Agentic Tasks 6 upvotes, #32 of 2026-06-23
  33. PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models 5 upvotes, #33 of 2026-06-23
  34. Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views 5 upvotes, #33 of 2026-06-23
  35. Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild 5 upvotes, #33 of 2026-06-23
  36. Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining 4 upvotes, #36 of 2026-06-23
  37. CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks 4 upvotes, #36 of 2026-06-23
  38. FastMix: Fast Data Mixture Optimization via Gradient Descent 3 upvotes, #38 of 2026-06-23
  39. Comparing Linear Probes with Mahalanobis Cosine Similarity 3 upvotes, #38 of 2026-06-23
  40. Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models 3 upvotes, #38 of 2026-06-23
  41. Go-with-the-Track: Video Compositing and Motion Control with Point Tracking 3 upvotes, #38 of 2026-06-23
  42. Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City 3 upvotes, #38 of 2026-06-23
  43. A Verifiable Search Is Not a Learnable Chain-of-Thought 3 upvotes, #38 of 2026-06-23
  44. UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation 3 upvotes, #38 of 2026-06-23
  45. Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation 3 upvotes, #38 of 2026-06-23
  46. When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents 2 upvotes, #46 of 2026-06-23
  47. Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity? 2 upvotes, #46 of 2026-06-23
  48. MeshFlow: Mesh Generation with Equivariant Flow Matching 2 upvotes, #46 of 2026-06-23
  49. AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining 1 upvotes, #49 of 2026-06-23
  50. Libretto: Giving LLM Agents a Sense of Musical Structure 1 upvotes, #49 of 2026-06-23
  51. HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions 1 upvotes, #49 of 2026-06-23
  52. Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach 0 upvotes, #52 of 2026-06-23
  53. An Exploratory Case Study of LLM-Assisted Refactoring and Gameplay Feature Generation in an Endless Runner Game 1 upvotes, #52 of 2026-06-23
  54. Improving Text-to-Music Generation with Human Preference Rewards 2 upvotes, #52 of 2026-06-23
  55. ShotcreteDepth: A Bi-modal Dataset for Robust Robotic Depth Perception in Shotcrete Construction Environments 1 upvotes, #52 of 2026-06-23
  56. TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization 0 upvotes, #52 of 2026-06-23

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.