Daily Papers of 2026-06-02

  1. On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters 223 upvotes, #1 of 2026-06-02
  2. Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs 192 upvotes, #2 of 2026-06-02
  3. Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding 143 upvotes, #3 of 2026-06-02
  4. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks 65 upvotes, #4 of 2026-06-02
  5. Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism 62 upvotes, #5 of 2026-06-02
  6. K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts 55 upvotes, #6 of 2026-06-02
  7. Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses 51 upvotes, #7 of 2026-06-02
  8. NITP: Next Implicit Token Prediction for LLM Pre-training 35 upvotes, #8 of 2026-06-02
  9. SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories 35 upvotes, #8 of 2026-06-02
  10. X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding 35 upvotes, #8 of 2026-06-02
  11. Draft-OPD: On-Policy Distillation for Speculative Draft Models 33 upvotes, #11 of 2026-06-02
  12. Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? 30 upvotes, #12 of 2026-06-02
  13. Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs 29 upvotes, #13 of 2026-06-02
  14. VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization 29 upvotes, #13 of 2026-06-02
  15. VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion 26 upvotes, #15 of 2026-06-02
  16. Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models 25 upvotes, #16 of 2026-06-02
  17. ESPO: Early-Stopping Proximal Policy Optimization 19 upvotes, #17 of 2026-06-02
  18. OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents 19 upvotes, #17 of 2026-06-02
  19. LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation 19 upvotes, #17 of 2026-06-02
  20. When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs 17 upvotes, #20 of 2026-06-02
  21. StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration 16 upvotes, #21 of 2026-06-02
  22. Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents 16 upvotes, #21 of 2026-06-02
  23. MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation 16 upvotes, #21 of 2026-06-02
  24. Joint Agent Memory and Exploration Learning via Novelty Signals 15 upvotes, #24 of 2026-06-02
  25. Brain-IT-VQA: From Brain Signals to Answers 14 upvotes, #25 of 2026-06-02
  26. LVSA: Training-Free Sparse Attention for Long Video Diffusion 14 upvotes, #25 of 2026-06-02
  27. TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation 14 upvotes, #25 of 2026-06-02
  28. MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft 11 upvotes, #28 of 2026-06-02
  29. Policy and World Modeling Co-Training for Language Agents 11 upvotes, #28 of 2026-06-02
  30. PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding 10 upvotes, #30 of 2026-06-02
  31. Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism 10 upvotes, #30 of 2026-06-02
  32. RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes 10 upvotes, #30 of 2026-06-02
  33. Agent Skills Should Go Beyond Text: The Case for Visual Skills 10 upvotes, #30 of 2026-06-02
  34. ACL-Verbatim: hallucination-free question answering for research 8 upvotes, #34 of 2026-06-02
  35. Measuring the Depth of LLM Unlearning via Activation Patching 8 upvotes, #34 of 2026-06-02
  36. SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models 8 upvotes, #34 of 2026-06-02
  37. FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search 8 upvotes, #34 of 2026-06-02
  38. AFUN: Towards an Affordance Foundation Model for Functionality Understanding 8 upvotes, #34 of 2026-06-02
  39. Unified Neural Scaling Laws 7 upvotes, #39 of 2026-06-02
  40. SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence 7 upvotes, #39 of 2026-06-02
  41. 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code 7 upvotes, #39 of 2026-06-02
  42. LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning 7 upvotes, #39 of 2026-06-02
  43. Multi-Agent Computer Use 7 upvotes, #39 of 2026-06-02
  44. Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning 7 upvotes, #39 of 2026-06-02
  45. RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models 7 upvotes, #39 of 2026-06-02
  46. Not only where, But when: Temporal Scheduling for RLVR 6 upvotes, #46 of 2026-06-02
  47. Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation 6 upvotes, #46 of 2026-06-02
  48. Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems 6 upvotes, #46 of 2026-06-02
  49. HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers 6 upvotes, #46 of 2026-06-02
  50. MindZero: Learning Online Mental Reasoning With Zero Annotations 4 upvotes, #50 of 2026-06-02
  51. EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers 3 upvotes, #51 of 2026-06-02
  52. Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization 2 upvotes, #52 of 2026-06-02
  53. Can Predicted Dynamics Exist in the Physical World? 2 upvotes, #52 of 2026-06-02
  54. StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement 2 upvotes, #52 of 2026-06-02
  55. Confidence-Adaptive SwiGLU for Mixture-of-Experts 2 upvotes, #52 of 2026-06-02
  56. ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats 2 upvotes, #52 of 2026-06-02
  57. Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models 2 upvotes, #52 of 2026-06-02
  58. Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan 1 upvotes, #58 of 2026-06-02
  59. AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering? 1 upvotes, #58 of 2026-06-02
  60. Review Arcade: On the Human Alignment and Gameability of LLM Reviews 1 upvotes, #58 of 2026-06-02
  61. The Hamilton-Jacobi Theory of Deep Learning 1 upvotes, #58 of 2026-06-02
  62. Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG 1 upvotes, #58 of 2026-06-02
  63. The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure 1 upvotes, #58 of 2026-06-02
  64. FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes 1 upvotes, #58 of 2026-06-02
  65. Model-Based Quality Assessment for Massively Multilingual Parallel Data 1 upvotes, #58 of 2026-06-02
  66. A Formally Verified Library of Mathematical Finance in Lean 4 1 upvotes, #58 of 2026-06-02
  67. Geometric Latent Reasoning Induces Shorter Generations in LLMs 1 upvotes, #58 of 2026-06-02
  68. Show, Don't TELL: Explainable AI-Generated Text Detection 0 upvotes, #68 of 2026-06-02
  69. Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures 0 upvotes, #68 of 2026-06-02
  70. τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation 1 upvotes, #68 of 2026-06-02
  71. DOT-MoE: Differentiable Optimal Transport for MoEfication 0 upvotes, #68 of 2026-06-02
  72. Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 0 upvotes, #68 of 2026-06-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.