Daily Papers of 2026-05-29

  1. AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security 142 upvotes, #1 of 2026-05-29
  2. Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments 138 upvotes, #2 of 2026-05-29
  3. OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources 76 upvotes, #3 of 2026-05-29
  4. CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation 61 upvotes, #4 of 2026-05-29
  5. Why Far Looks Up: Probing Spatial Representation in Vision-Language Models 59 upvotes, #5 of 2026-05-29
  6. minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models 56 upvotes, #6 of 2026-05-29
  7. YoCausal: How Far is Video Generation from World Model? A Causality Perspective 51 upvotes, #7 of 2026-05-29
  8. How LoRA Remembers? A Parametric Memory Law for LLM Finetuning 41 upvotes, #8 of 2026-05-29
  9. GenClaw: Code-Driven Agentic Image Generation 38 upvotes, #9 of 2026-05-29
  10. LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training 34 upvotes, #10 of 2026-05-29
  11. Native Audio-Visual Alignment for Generation 33 upvotes, #11 of 2026-05-29
  12. EarlyTom: Early Token Compression Completes Fast Video Understanding 32 upvotes, #12 of 2026-05-29
  13. Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning 31 upvotes, #13 of 2026-05-29
  14. UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering 26 upvotes, #14 of 2026-05-29
  15. Colored Noise Diffusion Sampling 25 upvotes, #15 of 2026-05-29
  16. When Should Models Change Their Minds? Contextual Belief Management in Large Language Models 24 upvotes, #16 of 2026-05-29
  17. LoMo: Local Modality Substitution for Deeper Vision-Language Fusion 23 upvotes, #17 of 2026-05-29
  18. Xetrieval: Mechanistically Explaining Dense Retrieval 21 upvotes, #18 of 2026-05-29
  19. Is Position Bias in Dense Retrievers Built In-or Learned from Data? 20 upvotes, #19 of 2026-05-29
  20. CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists 18 upvotes, #20 of 2026-05-29
  21. WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction 17 upvotes, #21 of 2026-05-29
  22. AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios 16 upvotes, #22 of 2026-05-29
  23. LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents 16 upvotes, #22 of 2026-05-29
  24. Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation 16 upvotes, #22 of 2026-05-29
  25. PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers 15 upvotes, #25 of 2026-05-29
  26. UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents 15 upvotes, #25 of 2026-05-29
  27. When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems 15 upvotes, #25 of 2026-05-29
  28. Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence 14 upvotes, #28 of 2026-05-29
  29. RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains 13 upvotes, #29 of 2026-05-29
  30. NeuROK: Generative 4D Neural Object Kinematics 12 upvotes, #30 of 2026-05-29
  31. AdaState: Self-Evolving Anchors for Streaming Video Generation 12 upvotes, #30 of 2026-05-29
  32. DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation 12 upvotes, #30 of 2026-05-29
  33. PANDO: Efficient Multimodal AI Agents via Online Skill Distillation 11 upvotes, #33 of 2026-05-29
  34. Parallax: Parameterized Local Linear Attention for Language Modeling 11 upvotes, #33 of 2026-05-29
  35. Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention 11 upvotes, #33 of 2026-05-29
  36. PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions 11 upvotes, #33 of 2026-05-29
  37. Thinking Before Constraining: A Unified Decoding Framework for Large Language Models 10 upvotes, #37 of 2026-05-29
  38. ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood 10 upvotes, #37 of 2026-05-29
  39. Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering 10 upvotes, #37 of 2026-05-29
  40. REPOT: Recoverable Program-of-Thought via Checkpoint Repair 10 upvotes, #37 of 2026-05-29
  41. Reflective Prompt Tuning through Language Model Function-Calling 9 upvotes, #41 of 2026-05-29
  42. CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM 9 upvotes, #41 of 2026-05-29
  43. CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval 9 upvotes, #41 of 2026-05-29
  44. Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments 8 upvotes, #44 of 2026-05-29
  45. SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control 8 upvotes, #44 of 2026-05-29
  46. PhoneWorld: Scaling Phone-Use Agent Environments 8 upvotes, #44 of 2026-05-29
  47. Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection 8 upvotes, #44 of 2026-05-29
  48. Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation 7 upvotes, #48 of 2026-05-29
  49. Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases 7 upvotes, #48 of 2026-05-29
  50. Convex Low-resource Accent-Robust Language Detection in Speech Recognition 6 upvotes, #50 of 2026-05-29
  51. Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation 4 upvotes, #51 of 2026-05-29
  52. OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants 3 upvotes, #52 of 2026-05-29
  53. MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation 2 upvotes, #53 of 2026-05-29
  54. Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas 2 upvotes, #53 of 2026-05-29
  55. ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usage 1 upvotes, #55 of 2026-05-29
  56. Reducing Political Manipulation with Consistency Training 1 upvotes, #55 of 2026-05-29
  57. Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning 1 upvotes, #55 of 2026-05-29
  58. Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection 1 upvotes, #55 of 2026-05-29
  59. Towards Consistent Video Geometry Estimation 3 upvotes, #59 of 2026-05-29

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.