Daily Papers of 2026-05-29
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security 142 upvotes, #1 of 2026-05-29
- Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments 138 upvotes, #2 of 2026-05-29
- OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources 76 upvotes, #3 of 2026-05-29
- CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation 61 upvotes, #4 of 2026-05-29
- Why Far Looks Up: Probing Spatial Representation in Vision-Language Models 59 upvotes, #5 of 2026-05-29
- minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models 56 upvotes, #6 of 2026-05-29
- YoCausal: How Far is Video Generation from World Model? A Causality Perspective 51 upvotes, #7 of 2026-05-29
- How LoRA Remembers? A Parametric Memory Law for LLM Finetuning 41 upvotes, #8 of 2026-05-29
- GenClaw: Code-Driven Agentic Image Generation 38 upvotes, #9 of 2026-05-29
- LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training 34 upvotes, #10 of 2026-05-29
- Native Audio-Visual Alignment for Generation 33 upvotes, #11 of 2026-05-29
- EarlyTom: Early Token Compression Completes Fast Video Understanding 32 upvotes, #12 of 2026-05-29
- Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning 31 upvotes, #13 of 2026-05-29
- UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering 26 upvotes, #14 of 2026-05-29
- Colored Noise Diffusion Sampling 25 upvotes, #15 of 2026-05-29
- When Should Models Change Their Minds? Contextual Belief Management in Large Language Models 24 upvotes, #16 of 2026-05-29
- LoMo: Local Modality Substitution for Deeper Vision-Language Fusion 23 upvotes, #17 of 2026-05-29
- Xetrieval: Mechanistically Explaining Dense Retrieval 21 upvotes, #18 of 2026-05-29
- Is Position Bias in Dense Retrievers Built In-or Learned from Data? 20 upvotes, #19 of 2026-05-29
- CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists 18 upvotes, #20 of 2026-05-29
- WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction 17 upvotes, #21 of 2026-05-29
- AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios 16 upvotes, #22 of 2026-05-29
- LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents 16 upvotes, #22 of 2026-05-29
- Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation 16 upvotes, #22 of 2026-05-29
- PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers 15 upvotes, #25 of 2026-05-29
- UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents 15 upvotes, #25 of 2026-05-29
- When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems 15 upvotes, #25 of 2026-05-29
- Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence 14 upvotes, #28 of 2026-05-29
- RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains 13 upvotes, #29 of 2026-05-29
- NeuROK: Generative 4D Neural Object Kinematics 12 upvotes, #30 of 2026-05-29
- AdaState: Self-Evolving Anchors for Streaming Video Generation 12 upvotes, #30 of 2026-05-29
- DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation 12 upvotes, #30 of 2026-05-29
- PANDO: Efficient Multimodal AI Agents via Online Skill Distillation 11 upvotes, #33 of 2026-05-29
- Parallax: Parameterized Local Linear Attention for Language Modeling 11 upvotes, #33 of 2026-05-29
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention 11 upvotes, #33 of 2026-05-29
- PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions 11 upvotes, #33 of 2026-05-29
- Thinking Before Constraining: A Unified Decoding Framework for Large Language Models 10 upvotes, #37 of 2026-05-29
- ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood 10 upvotes, #37 of 2026-05-29
- Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering 10 upvotes, #37 of 2026-05-29
- REPOT: Recoverable Program-of-Thought via Checkpoint Repair 10 upvotes, #37 of 2026-05-29
- Reflective Prompt Tuning through Language Model Function-Calling 9 upvotes, #41 of 2026-05-29
- CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM 9 upvotes, #41 of 2026-05-29
- CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval 9 upvotes, #41 of 2026-05-29
- Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments 8 upvotes, #44 of 2026-05-29
- SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control 8 upvotes, #44 of 2026-05-29
- PhoneWorld: Scaling Phone-Use Agent Environments 8 upvotes, #44 of 2026-05-29
- Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection 8 upvotes, #44 of 2026-05-29
- Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation 7 upvotes, #48 of 2026-05-29
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases 7 upvotes, #48 of 2026-05-29
- Convex Low-resource Accent-Robust Language Detection in Speech Recognition 6 upvotes, #50 of 2026-05-29
- Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation 4 upvotes, #51 of 2026-05-29
- OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants 3 upvotes, #52 of 2026-05-29
- MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation 2 upvotes, #53 of 2026-05-29
- Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas 2 upvotes, #53 of 2026-05-29
- ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usage 1 upvotes, #55 of 2026-05-29
- Reducing Political Manipulation with Consistency Training 1 upvotes, #55 of 2026-05-29
- Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning 1 upvotes, #55 of 2026-05-29
- Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection 1 upvotes, #55 of 2026-05-29
- Towards Consistent Video Geometry Estimation 3 upvotes, #59 of 2026-05-29
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.