Daily Papers of 2026-06-02
- On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters 223 upvotes, #1 of 2026-06-02
- Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs 192 upvotes, #2 of 2026-06-02
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding 143 upvotes, #3 of 2026-06-02
- A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks 65 upvotes, #4 of 2026-06-02
- Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism 62 upvotes, #5 of 2026-06-02
- K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts 55 upvotes, #6 of 2026-06-02
- Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses 51 upvotes, #7 of 2026-06-02
- NITP: Next Implicit Token Prediction for LLM Pre-training 35 upvotes, #8 of 2026-06-02
- SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories 35 upvotes, #8 of 2026-06-02
- X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding 35 upvotes, #8 of 2026-06-02
- Draft-OPD: On-Policy Distillation for Speculative Draft Models 33 upvotes, #11 of 2026-06-02
- Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? 30 upvotes, #12 of 2026-06-02
- Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs 29 upvotes, #13 of 2026-06-02
- VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization 29 upvotes, #13 of 2026-06-02
- VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion 26 upvotes, #15 of 2026-06-02
- Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models 25 upvotes, #16 of 2026-06-02
- ESPO: Early-Stopping Proximal Policy Optimization 19 upvotes, #17 of 2026-06-02
- OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents 19 upvotes, #17 of 2026-06-02
- LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation 19 upvotes, #17 of 2026-06-02
- When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs 17 upvotes, #20 of 2026-06-02
- StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration 16 upvotes, #21 of 2026-06-02
- Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents 16 upvotes, #21 of 2026-06-02
- MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation 16 upvotes, #21 of 2026-06-02
- Joint Agent Memory and Exploration Learning via Novelty Signals 15 upvotes, #24 of 2026-06-02
- Brain-IT-VQA: From Brain Signals to Answers 14 upvotes, #25 of 2026-06-02
- LVSA: Training-Free Sparse Attention for Long Video Diffusion 14 upvotes, #25 of 2026-06-02
- TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation 14 upvotes, #25 of 2026-06-02
- MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft 11 upvotes, #28 of 2026-06-02
- Policy and World Modeling Co-Training for Language Agents 11 upvotes, #28 of 2026-06-02
- PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding 10 upvotes, #30 of 2026-06-02
- Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism 10 upvotes, #30 of 2026-06-02
- RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes 10 upvotes, #30 of 2026-06-02
- Agent Skills Should Go Beyond Text: The Case for Visual Skills 10 upvotes, #30 of 2026-06-02
- ACL-Verbatim: hallucination-free question answering for research 8 upvotes, #34 of 2026-06-02
- Measuring the Depth of LLM Unlearning via Activation Patching 8 upvotes, #34 of 2026-06-02
- SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models 8 upvotes, #34 of 2026-06-02
- FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search 8 upvotes, #34 of 2026-06-02
- AFUN: Towards an Affordance Foundation Model for Functionality Understanding 8 upvotes, #34 of 2026-06-02
- Unified Neural Scaling Laws 7 upvotes, #39 of 2026-06-02
- SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence 7 upvotes, #39 of 2026-06-02
- 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code 7 upvotes, #39 of 2026-06-02
- LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning 7 upvotes, #39 of 2026-06-02
- Multi-Agent Computer Use 7 upvotes, #39 of 2026-06-02
- Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning 7 upvotes, #39 of 2026-06-02
- RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models 7 upvotes, #39 of 2026-06-02
- Not only where, But when: Temporal Scheduling for RLVR 6 upvotes, #46 of 2026-06-02
- Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation 6 upvotes, #46 of 2026-06-02
- Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems 6 upvotes, #46 of 2026-06-02
- HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers 6 upvotes, #46 of 2026-06-02
- MindZero: Learning Online Mental Reasoning With Zero Annotations 4 upvotes, #50 of 2026-06-02
- EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers 3 upvotes, #51 of 2026-06-02
- Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization 2 upvotes, #52 of 2026-06-02
- Can Predicted Dynamics Exist in the Physical World? 2 upvotes, #52 of 2026-06-02
- StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement 2 upvotes, #52 of 2026-06-02
- Confidence-Adaptive SwiGLU for Mixture-of-Experts 2 upvotes, #52 of 2026-06-02
- ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats 2 upvotes, #52 of 2026-06-02
- Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models 2 upvotes, #52 of 2026-06-02
- Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan 1 upvotes, #58 of 2026-06-02
- AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering? 1 upvotes, #58 of 2026-06-02
- Review Arcade: On the Human Alignment and Gameability of LLM Reviews 1 upvotes, #58 of 2026-06-02
- The Hamilton-Jacobi Theory of Deep Learning 1 upvotes, #58 of 2026-06-02
- Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG 1 upvotes, #58 of 2026-06-02
- The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure 1 upvotes, #58 of 2026-06-02
- FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes 1 upvotes, #58 of 2026-06-02
- Model-Based Quality Assessment for Massively Multilingual Parallel Data 1 upvotes, #58 of 2026-06-02
- A Formally Verified Library of Mathematical Finance in Lean 4 1 upvotes, #58 of 2026-06-02
- Geometric Latent Reasoning Induces Shorter Generations in LLMs 1 upvotes, #58 of 2026-06-02
- Show, Don't TELL: Explainable AI-Generated Text Detection 0 upvotes, #68 of 2026-06-02
- Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures 0 upvotes, #68 of 2026-06-02
- τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation 1 upvotes, #68 of 2026-06-02
- DOT-MoE: Differentiable Optimal Transport for MoEfication 0 upvotes, #68 of 2026-06-02
- Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 0 upvotes, #68 of 2026-06-02
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.