Daily Papers of 2026-05-26

  1. DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning 133 upvotes, #1 of 2026-05-26
  2. WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation 101 upvotes, #2 of 2026-05-26
  3. Macaron-A2UI: A Model for Generative UI in Personal Agents 80 upvotes, #3 of 2026-05-26
  4. Foundation Protocol: A Coordination Layer for Agentic Society 79 upvotes, #4 of 2026-05-26
  5. TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction 51 upvotes, #5 of 2026-05-26
  6. Toward Native Multimodal Modeling: A Roadmap 42 upvotes, #6 of 2026-05-26
  7. ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention 41 upvotes, #7 of 2026-05-26
  8. QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks 41 upvotes, #7 of 2026-05-26
  9. Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents 41 upvotes, #7 of 2026-05-26
  10. ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning 34 upvotes, #10 of 2026-05-26
  11. CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents 31 upvotes, #11 of 2026-05-26
  12. AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery 29 upvotes, #12 of 2026-05-26
  13. Your Embedding Model is SMARTer Than You Think 25 upvotes, #13 of 2026-05-26
  14. Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World 23 upvotes, #14 of 2026-05-26
  15. ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement 21 upvotes, #15 of 2026-05-26
  16. SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills 20 upvotes, #16 of 2026-05-26
  17. Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion 20 upvotes, #16 of 2026-05-26
  18. Recursive Flow Matching 19 upvotes, #18 of 2026-05-26
  19. On-Policy Adversarial Flow Distillation for Autoregressive Video Generation 18 upvotes, #19 of 2026-05-26
  20. MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing 17 upvotes, #20 of 2026-05-26
  21. InstructSAM: Segment Any Instance with Any Instructions 17 upvotes, #20 of 2026-05-26
  22. Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents 16 upvotes, #22 of 2026-05-26
  23. RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator 15 upvotes, #23 of 2026-05-26
  24. Channel-wise Vector Quantization 15 upvotes, #23 of 2026-05-26
  25. Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth 14 upvotes, #25 of 2026-05-26
  26. Helix4D: Complex 4D Mesh Generation 14 upvotes, #25 of 2026-05-26
  27. Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild 12 upvotes, #27 of 2026-05-26
  28. Geometry-Aware Image Flow Matching 11 upvotes, #28 of 2026-05-26
  29. Language Models Need Sleep 11 upvotes, #28 of 2026-05-26
  30. Towards Customized Multimodal Role-Play 10 upvotes, #30 of 2026-05-26
  31. CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models 10 upvotes, #30 of 2026-05-26
  32. SEAL: Synergistic Co-Evolution of Agents and Learning Environments 10 upvotes, #30 of 2026-05-26
  33. CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test 9 upvotes, #33 of 2026-05-26
  34. SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking 8 upvotes, #34 of 2026-05-26
  35. MetaphorVU: Towards Metaphorical Video Understanding 8 upvotes, #34 of 2026-05-26
  36. Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution 7 upvotes, #36 of 2026-05-26
  37. ECHO: Terminal Agents Learn World Models for Free 7 upvotes, #36 of 2026-05-26
  38. PRISM: Position-encoded Regressive Inverse Spectral Model for Multilayer Thin-Film Design 7 upvotes, #36 of 2026-05-26
  39. How Far Will They Go? Red-Teaming Online Influence with Large Language Models 6 upvotes, #39 of 2026-05-26
  40. Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO 5 upvotes, #40 of 2026-05-26
  41. Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference 5 upvotes, #40 of 2026-05-26
  42. Reinforcing Few-step Generators via Reward-Tilted Distribution Matching 5 upvotes, #40 of 2026-05-26
  43. Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints 4 upvotes, #43 of 2026-05-26
  44. Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries 4 upvotes, #43 of 2026-05-26
  45. MotiMotion: Motion-Controlled Video Generation with Visual Reasoning 4 upvotes, #43 of 2026-05-26
  46. HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction 4 upvotes, #43 of 2026-05-26
  47. Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models 4 upvotes, #43 of 2026-05-26
  48. Seeing the Needle in the Haystack: Towards Weakly-Supervised Log Instance Anomaly Localization via Counterfactual Perturbation 3 upvotes, #48 of 2026-05-26
  49. SemBridge: Language Transfer in Sparse Encoders via Multilingual Semantic Bridges 3 upvotes, #48 of 2026-05-26
  50. Cross-scale Aligned Supervision for Training GANs 3 upvotes, #48 of 2026-05-26
  51. ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison 1 upvotes, #51 of 2026-05-26
  52. Decoding the Critique Mechanism in Large Reasoning Models 0 upvotes, #52 of 2026-05-26
  53. Pixel-Level Pavement Distress Assessment Using Instance Segmentation 1 upvotes, #52 of 2026-05-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.