Daily Papers of 2026-05-21

  1. Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining 144 upvotes, #1 of 2026-05-21
  2. Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation 131 upvotes, #2 of 2026-05-21
  3. Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos 91 upvotes, #3 of 2026-05-21
  4. HRM-Text: Efficient Pretraining Beyond Scaling 89 upvotes, #4 of 2026-05-21
  5. IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools 83 upvotes, #5 of 2026-05-21
  6. A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook 56 upvotes, #6 of 2026-05-21
  7. You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories 49 upvotes, #7 of 2026-05-21
  8. OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond 39 upvotes, #8 of 2026-05-21
  9. Toto 2.0: Time Series Forecasting Enters the Scaling Era 38 upvotes, #9 of 2026-05-21
  10. It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs 30 upvotes, #10 of 2026-05-21
  11. Generative Recursive Reasoning 29 upvotes, #11 of 2026-05-21
  12. PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models 29 upvotes, #11 of 2026-05-21
  13. Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs 28 upvotes, #13 of 2026-05-21
  14. Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning 22 upvotes, #14 of 2026-05-21
  15. CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing 21 upvotes, #15 of 2026-05-21
  16. LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening 20 upvotes, #16 of 2026-05-21
  17. Stable Audio 3 17 upvotes, #17 of 2026-05-21
  18. Stitched Value Model for Diffusion Alignment 12 upvotes, #18 of 2026-05-21
  19. Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines 12 upvotes, #18 of 2026-05-21
  20. Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation 11 upvotes, #20 of 2026-05-21
  21. On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists 11 upvotes, #20 of 2026-05-21
  22. Learning from Language Feedback via Variational Policy Distillation 10 upvotes, #22 of 2026-05-21
  23. RiT: Vanilla Diffusion Transformers Suffice in Representation Space 10 upvotes, #22 of 2026-05-21
  24. OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization 9 upvotes, #24 of 2026-05-21
  25. MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization 8 upvotes, #25 of 2026-05-21
  26. UniT: Unified Geometry Learning with Group Autoregressive Transformer 8 upvotes, #25 of 2026-05-21
  27. OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation 8 upvotes, #25 of 2026-05-21
  28. SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering 7 upvotes, #28 of 2026-05-21
  29. The Unlearnability Phenomenon in RLVR for Language Models 6 upvotes, #29 of 2026-05-21
  30. Mem-π: Adaptive Memory through Learning When and What to Generate 6 upvotes, #29 of 2026-05-21
  31. Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment 5 upvotes, #31 of 2026-05-21
  32. PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis 4 upvotes, #32 of 2026-05-21
  33. LongMINT: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems 4 upvotes, #32 of 2026-05-21
  34. Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency 4 upvotes, #32 of 2026-05-21
  35. TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload 4 upvotes, #32 of 2026-05-21
  36. SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents 4 upvotes, #32 of 2026-05-21
  37. DrawMotion: Generating 3D Human Motions by Freehand Drawing 3 upvotes, #37 of 2026-05-21
  38. Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection 2 upvotes, #38 of 2026-05-21
  39. DynMuon: A Dynamic Spectral Shaping View of Muon 2 upvotes, #38 of 2026-05-21
  40. Capturing LLM Capabilities via Evidence-Calibrated Query Clustering 2 upvotes, #38 of 2026-05-21
  41. Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation 2 upvotes, #38 of 2026-05-21
  42. Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models 2 upvotes, #38 of 2026-05-21
  43. iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance 2 upvotes, #38 of 2026-05-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.