Daily Papers of 2026-05-26
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning 133 upvotes, #1 of 2026-05-26
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation 101 upvotes, #2 of 2026-05-26
- Macaron-A2UI: A Model for Generative UI in Personal Agents 80 upvotes, #3 of 2026-05-26
- Foundation Protocol: A Coordination Layer for Agentic Society 79 upvotes, #4 of 2026-05-26
- TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction 51 upvotes, #5 of 2026-05-26
- Toward Native Multimodal Modeling: A Roadmap 42 upvotes, #6 of 2026-05-26
- ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention 41 upvotes, #7 of 2026-05-26
- QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks 41 upvotes, #7 of 2026-05-26
- Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents 41 upvotes, #7 of 2026-05-26
- ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning 34 upvotes, #10 of 2026-05-26
- CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents 31 upvotes, #11 of 2026-05-26
- AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery 29 upvotes, #12 of 2026-05-26
- Your Embedding Model is SMARTer Than You Think 25 upvotes, #13 of 2026-05-26
- Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World 23 upvotes, #14 of 2026-05-26
- ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement 21 upvotes, #15 of 2026-05-26
- SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills 20 upvotes, #16 of 2026-05-26
- Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion 20 upvotes, #16 of 2026-05-26
- Recursive Flow Matching 19 upvotes, #18 of 2026-05-26
- On-Policy Adversarial Flow Distillation for Autoregressive Video Generation 18 upvotes, #19 of 2026-05-26
- MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing 17 upvotes, #20 of 2026-05-26
- InstructSAM: Segment Any Instance with Any Instructions 17 upvotes, #20 of 2026-05-26
- Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents 16 upvotes, #22 of 2026-05-26
- RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator 15 upvotes, #23 of 2026-05-26
- Channel-wise Vector Quantization 15 upvotes, #23 of 2026-05-26
- Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth 14 upvotes, #25 of 2026-05-26
- Helix4D: Complex 4D Mesh Generation 14 upvotes, #25 of 2026-05-26
- Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild 12 upvotes, #27 of 2026-05-26
- Geometry-Aware Image Flow Matching 11 upvotes, #28 of 2026-05-26
- Language Models Need Sleep 11 upvotes, #28 of 2026-05-26
- Towards Customized Multimodal Role-Play 10 upvotes, #30 of 2026-05-26
- CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models 10 upvotes, #30 of 2026-05-26
- SEAL: Synergistic Co-Evolution of Agents and Learning Environments 10 upvotes, #30 of 2026-05-26
- CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test 9 upvotes, #33 of 2026-05-26
- SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking 8 upvotes, #34 of 2026-05-26
- MetaphorVU: Towards Metaphorical Video Understanding 8 upvotes, #34 of 2026-05-26
- Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution 7 upvotes, #36 of 2026-05-26
- ECHO: Terminal Agents Learn World Models for Free 7 upvotes, #36 of 2026-05-26
- PRISM: Position-encoded Regressive Inverse Spectral Model for Multilayer Thin-Film Design 7 upvotes, #36 of 2026-05-26
- How Far Will They Go? Red-Teaming Online Influence with Large Language Models 6 upvotes, #39 of 2026-05-26
- Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO 5 upvotes, #40 of 2026-05-26
- Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference 5 upvotes, #40 of 2026-05-26
- Reinforcing Few-step Generators via Reward-Tilted Distribution Matching 5 upvotes, #40 of 2026-05-26
- Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints 4 upvotes, #43 of 2026-05-26
- Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries 4 upvotes, #43 of 2026-05-26
- MotiMotion: Motion-Controlled Video Generation with Visual Reasoning 4 upvotes, #43 of 2026-05-26
- HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction 4 upvotes, #43 of 2026-05-26
- Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models 4 upvotes, #43 of 2026-05-26
- Seeing the Needle in the Haystack: Towards Weakly-Supervised Log Instance Anomaly Localization via Counterfactual Perturbation 3 upvotes, #48 of 2026-05-26
- SemBridge: Language Transfer in Sparse Encoders via Multilingual Semantic Bridges 3 upvotes, #48 of 2026-05-26
- Cross-scale Aligned Supervision for Training GANs 3 upvotes, #48 of 2026-05-26
- ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison 1 upvotes, #51 of 2026-05-26
- Decoding the Critique Mechanism in Large Reasoning Models 0 upvotes, #52 of 2026-05-26
- Pixel-Level Pavement Distress Assessment Using Instance Segmentation 1 upvotes, #52 of 2026-05-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.