Daily Papers of 2026-06-23
- PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems 95 upvotes, #1 of 2026-06-23
- EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions 79 upvotes, #2 of 2026-06-23
- OpenRath: Session-Centered Runtime State for Agent Systems 77 upvotes, #3 of 2026-06-23
- Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention 76 upvotes, #4 of 2026-06-23
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams 73 upvotes, #5 of 2026-06-23
- World Action Models: A Survey 56 upvotes, #6 of 2026-06-23
- KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking 49 upvotes, #7 of 2026-06-23
- Unlimited OCR Works 43 upvotes, #8 of 2026-06-23
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents 37 upvotes, #9 of 2026-06-23
- EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory 32 upvotes, #10 of 2026-06-23
- Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding 24 upvotes, #11 of 2026-06-23
- BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language 24 upvotes, #11 of 2026-06-23
- SkillHarness: Harnessing Safe Skills for Computer-Use Agents 20 upvotes, #13 of 2026-06-23
- Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation 18 upvotes, #14 of 2026-06-23
- HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization 18 upvotes, #14 of 2026-06-23
- Self-Compacting Language Model Agents 18 upvotes, #14 of 2026-06-23
- Training Open Models for Agentic Phone Use 16 upvotes, #17 of 2026-06-23
- Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark 15 upvotes, #18 of 2026-06-23
- DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks 14 upvotes, #19 of 2026-06-23
- Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents 14 upvotes, #19 of 2026-06-23
- Tmax: A simple recipe for terminal agents 14 upvotes, #19 of 2026-06-23
- Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills 11 upvotes, #22 of 2026-06-23
- PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning 11 upvotes, #22 of 2026-06-23
- Exploring the Design Space of Reward Backpropagation for Flow Matching 10 upvotes, #24 of 2026-06-23
- Vera: A Layered Diffusion Model for Content-Preserving Video Editing 10 upvotes, #24 of 2026-06-23
- Tapered Language Models 9 upvotes, #26 of 2026-06-23
- Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning 8 upvotes, #27 of 2026-06-23
- Safe Few-Step Generation via Velocity Editing 8 upvotes, #27 of 2026-06-23
- Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning 7 upvotes, #29 of 2026-06-23
- Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs 7 upvotes, #29 of 2026-06-23
- Causal Discovery in the Era of Agents 7 upvotes, #29 of 2026-06-23
- Counsel: A Meta-Evaluation Dataset for Agentic Tasks 6 upvotes, #32 of 2026-06-23
- PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models 5 upvotes, #33 of 2026-06-23
- Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views 5 upvotes, #33 of 2026-06-23
- Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild 5 upvotes, #33 of 2026-06-23
- Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining 4 upvotes, #36 of 2026-06-23
- CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks 4 upvotes, #36 of 2026-06-23
- FastMix: Fast Data Mixture Optimization via Gradient Descent 3 upvotes, #38 of 2026-06-23
- Comparing Linear Probes with Mahalanobis Cosine Similarity 3 upvotes, #38 of 2026-06-23
- Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models 3 upvotes, #38 of 2026-06-23
- Go-with-the-Track: Video Compositing and Motion Control with Point Tracking 3 upvotes, #38 of 2026-06-23
- Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City 3 upvotes, #38 of 2026-06-23
- A Verifiable Search Is Not a Learnable Chain-of-Thought 3 upvotes, #38 of 2026-06-23
- UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation 3 upvotes, #38 of 2026-06-23
- Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation 3 upvotes, #38 of 2026-06-23
- When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents 2 upvotes, #46 of 2026-06-23
- Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity? 2 upvotes, #46 of 2026-06-23
- MeshFlow: Mesh Generation with Equivariant Flow Matching 2 upvotes, #46 of 2026-06-23
- AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining 1 upvotes, #49 of 2026-06-23
- Libretto: Giving LLM Agents a Sense of Musical Structure 1 upvotes, #49 of 2026-06-23
- HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions 1 upvotes, #49 of 2026-06-23
- Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach 0 upvotes, #52 of 2026-06-23
- An Exploratory Case Study of LLM-Assisted Refactoring and Gameplay Feature Generation in an Endless Runner Game 1 upvotes, #52 of 2026-06-23
- Improving Text-to-Music Generation with Human Preference Rewards 2 upvotes, #52 of 2026-06-23
- ShotcreteDepth: A Bi-modal Dataset for Robust Robotic Depth Perception in Shotcrete Construction Environments 1 upvotes, #52 of 2026-06-23
- TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization 0 upvotes, #52 of 2026-06-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.