Daily Papers of 2026-05-19

  1. Code as Agent Harness 204 upvotes, #1 of 2026-05-19
  2. SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution 125 upvotes, #2 of 2026-05-19
  3. LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation 109 upvotes, #3 of 2026-05-19
  4. Lance: Unified Multimodal Modeling by Multi-Task Synergy 74 upvotes, #4 of 2026-05-19
  5. AI for Auto-Research: Roadmap & User Guide 65 upvotes, #5 of 2026-05-19
  6. OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization 62 upvotes, #6 of 2026-05-19
  7. CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? 52 upvotes, #7 of 2026-05-19
  8. Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis 41 upvotes, #8 of 2026-05-19
  9. KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration 37 upvotes, #9 of 2026-05-19
  10. OProver: A Unified Framework for Agentic Formal Theorem Proving 30 upvotes, #10 of 2026-05-19
  11. Post-Trained MoE Can Skip Half Experts via Self-Distillation 30 upvotes, #10 of 2026-05-19
  12. VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation 27 upvotes, #12 of 2026-05-19
  13. LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs 24 upvotes, #13 of 2026-05-19
  14. Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models 22 upvotes, #14 of 2026-05-19
  15. NEWTON: Agentic Planning for Physically Grounded Video Generation 22 upvotes, #14 of 2026-05-19
  16. Measuring Maximum Activations in Open Large Language Models 18 upvotes, #16 of 2026-05-19
  17. AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs 16 upvotes, #17 of 2026-05-19
  18. Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use 15 upvotes, #18 of 2026-05-19
  19. EndPrompt: Efficient Long-Context Extension via Terminal Anchoring 15 upvotes, #18 of 2026-05-19
  20. StableVLA: Towards Robust Vision-Language-Action Models without Extra Data 15 upvotes, #18 of 2026-05-19
  21. Targeted Neuron Modulation via Contrastive Pair Search 14 upvotes, #21 of 2026-05-19
  22. Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement 14 upvotes, #21 of 2026-05-19
  23. CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection 12 upvotes, #23 of 2026-05-19
  24. From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements 12 upvotes, #23 of 2026-05-19
  25. NGM: A Plug-and-Play Training-Free Memory Module for LLMs 10 upvotes, #25 of 2026-05-19
  26. WavFlow: Audio Generation in Waveform Space 10 upvotes, #25 of 2026-05-19
  27. TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents 9 upvotes, #27 of 2026-05-19
  28. FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models 8 upvotes, #28 of 2026-05-19
  29. MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents 8 upvotes, #28 of 2026-05-19
  30. MixSD: Mixed Contextual Self-Distillation for Knowledge Injection 7 upvotes, #30 of 2026-05-19
  31. AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents 7 upvotes, #30 of 2026-05-19
  32. DexHoldem: Playing Texas Hold'em with Dexterous Embodied System 7 upvotes, #30 of 2026-05-19
  33. Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces 6 upvotes, #33 of 2026-05-19
  34. Geometric Phase Transition Enables Extreme Hippocampal Memory Capacity 5 upvotes, #34 of 2026-05-19
  35. Evaluating Cognitive Age Alignment in Interactive AI Agents 5 upvotes, #34 of 2026-05-19
  36. Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models 5 upvotes, #34 of 2026-05-19
  37. SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training 5 upvotes, #34 of 2026-05-19
  38. Actionable World Representation 5 upvotes, #34 of 2026-05-19
  39. A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation 4 upvotes, #39 of 2026-05-19
  40. SNLP: Layer-Parallel Inference via Structured Newton Corrections 4 upvotes, #39 of 2026-05-19
  41. GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions 3 upvotes, #41 of 2026-05-19
  42. Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring 3 upvotes, #41 of 2026-05-19
  43. AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents 3 upvotes, #41 of 2026-05-19
  44. Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers 3 upvotes, #41 of 2026-05-19
  45. AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models 2 upvotes, #45 of 2026-05-19
  46. TopoPrimer: The Missing Topological Context in Forecasting Models 2 upvotes, #45 of 2026-05-19
  47. E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring 2 upvotes, #45 of 2026-05-19
  48. Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics 2 upvotes, #45 of 2026-05-19
  49. SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science 1 upvotes, #49 of 2026-05-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.