Daily Papers of 2026-05-21
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining 144 upvotes, #1 of 2026-05-21
- Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation 131 upvotes, #2 of 2026-05-21
- Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos 91 upvotes, #3 of 2026-05-21
- HRM-Text: Efficient Pretraining Beyond Scaling 89 upvotes, #4 of 2026-05-21
- IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools 83 upvotes, #5 of 2026-05-21
- A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook 56 upvotes, #6 of 2026-05-21
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories 49 upvotes, #7 of 2026-05-21
- OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond 39 upvotes, #8 of 2026-05-21
- Toto 2.0: Time Series Forecasting Enters the Scaling Era 38 upvotes, #9 of 2026-05-21
- It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs 30 upvotes, #10 of 2026-05-21
- Generative Recursive Reasoning 29 upvotes, #11 of 2026-05-21
- PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models 29 upvotes, #11 of 2026-05-21
- Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs 28 upvotes, #13 of 2026-05-21
- Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning 22 upvotes, #14 of 2026-05-21
- CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing 21 upvotes, #15 of 2026-05-21
- LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening 20 upvotes, #16 of 2026-05-21
- Stable Audio 3 17 upvotes, #17 of 2026-05-21
- Stitched Value Model for Diffusion Alignment 12 upvotes, #18 of 2026-05-21
- Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines 12 upvotes, #18 of 2026-05-21
- Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation 11 upvotes, #20 of 2026-05-21
- On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists 11 upvotes, #20 of 2026-05-21
- Learning from Language Feedback via Variational Policy Distillation 10 upvotes, #22 of 2026-05-21
- RiT: Vanilla Diffusion Transformers Suffice in Representation Space 10 upvotes, #22 of 2026-05-21
- OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization 9 upvotes, #24 of 2026-05-21
- MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization 8 upvotes, #25 of 2026-05-21
- UniT: Unified Geometry Learning with Group Autoregressive Transformer 8 upvotes, #25 of 2026-05-21
- OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation 8 upvotes, #25 of 2026-05-21
- SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering 7 upvotes, #28 of 2026-05-21
- The Unlearnability Phenomenon in RLVR for Language Models 6 upvotes, #29 of 2026-05-21
- Mem-π: Adaptive Memory through Learning When and What to Generate 6 upvotes, #29 of 2026-05-21
- Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment 5 upvotes, #31 of 2026-05-21
- PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis 4 upvotes, #32 of 2026-05-21
- LongMINT: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems 4 upvotes, #32 of 2026-05-21
- Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency 4 upvotes, #32 of 2026-05-21
- TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload 4 upvotes, #32 of 2026-05-21
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents 4 upvotes, #32 of 2026-05-21
- DrawMotion: Generating 3D Human Motions by Freehand Drawing 3 upvotes, #37 of 2026-05-21
- Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection 2 upvotes, #38 of 2026-05-21
- DynMuon: A Dynamic Spectral Shaping View of Muon 2 upvotes, #38 of 2026-05-21
- Capturing LLM Capabilities via Evidence-Calibrated Query Clustering 2 upvotes, #38 of 2026-05-21
- Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation 2 upvotes, #38 of 2026-05-21
- Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models 2 upvotes, #38 of 2026-05-21
- iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance 2 upvotes, #38 of 2026-05-21
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.