Daily Papers of 2026-05-13
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 184 upvotes, #1 of 2026-05-13
- MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents 145 upvotes, #2 of 2026-05-13
- δ-mem: Efficient Online Memory for Large Language Models 119 upvotes, #3 of 2026-05-13
- RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards 74 upvotes, #4 of 2026-05-13
- World Action Models: The Next Frontier in Embodied AI 64 upvotes, #5 of 2026-05-13
- Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics 60 upvotes, #6 of 2026-05-13
- MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments 54 upvotes, #7 of 2026-05-13
- Efficient Pre-Training with Token Superposition 42 upvotes, #8 of 2026-05-13
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward 35 upvotes, #9 of 2026-05-13
- Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization 33 upvotes, #10 of 2026-05-13
- Teaching Language Models to Think in Code 30 upvotes, #11 of 2026-05-13
- L2P: Unlocking Latent Potential for Pixel Generation 29 upvotes, #12 of 2026-05-13
- CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives 28 upvotes, #13 of 2026-05-13
- ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents 27 upvotes, #14 of 2026-05-13
- Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents 21 upvotes, #15 of 2026-05-13
- Useful Memories Become Faulty When Continuously Updated by LLMs 19 upvotes, #16 of 2026-05-13
- Continual Harness: Online Adaptation for Self-Improving Foundation Agents 17 upvotes, #17 of 2026-05-13
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs 17 upvotes, #17 of 2026-05-13
- Learning, Fast and Slow: Towards LLMs That Adapt Continually 17 upvotes, #17 of 2026-05-13
- World Model for Robot Learning: A Comprehensive Survey 16 upvotes, #20 of 2026-05-13
- Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States 16 upvotes, #20 of 2026-05-13
- On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment 16 upvotes, #20 of 2026-05-13
- Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction 16 upvotes, #20 of 2026-05-13
- Relit-LiVE: Relight Video by Jointly Learning Environment Video 15 upvotes, #24 of 2026-05-13
- SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning 14 upvotes, #25 of 2026-05-13
- From Web to Pixels: Bringing Agentic Search into Visual Perception 14 upvotes, #25 of 2026-05-13
- Covering Human Action Space for Computer Use: Data Synthesis and Benchmark 14 upvotes, #25 of 2026-05-13
- One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue 12 upvotes, #28 of 2026-05-13
- EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scales 11 upvotes, #29 of 2026-05-13
- Do not copy and paste! Rewriting strategies for code retrieval 10 upvotes, #30 of 2026-05-13
- PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks 10 upvotes, #30 of 2026-05-13
- Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training 10 upvotes, #30 of 2026-05-13
- Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values 9 upvotes, #33 of 2026-05-13
- LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models 9 upvotes, #33 of 2026-05-13
- Debiased Model-based Representations for Sample-efficient Continuous Control 9 upvotes, #33 of 2026-05-13
- Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs 7 upvotes, #36 of 2026-05-13
- MEME: Multi-entity & Evolving Memory Evaluation 7 upvotes, #36 of 2026-05-13
- A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models 6 upvotes, #38 of 2026-05-13
- A Causal Language Modeling Detour Improves Encoder Continued Pretraining 6 upvotes, #38 of 2026-05-13
- Solve the Loop: Attractor Models for Language and Reasoning 6 upvotes, #38 of 2026-05-13
- AdaPreLoRA: Adafactor Preconditioned Low-Rank Adaptation 5 upvotes, #41 of 2026-05-13
- The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes 5 upvotes, #41 of 2026-05-13
- UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning 5 upvotes, #41 of 2026-05-13
- LychSim: A Controllable and Interactive Simulation Framework for Vision Research 5 upvotes, #41 of 2026-05-13
- Reward Hacking in Rubric-Based Reinforcement Learning 5 upvotes, #41 of 2026-05-13
- Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation 5 upvotes, #41 of 2026-05-13
- LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues 5 upvotes, #41 of 2026-05-13
- Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts 4 upvotes, #48 of 2026-05-13
- VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors 4 upvotes, #48 of 2026-05-13
- AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration -- Learning from Cheap, Optimizing Expensive 4 upvotes, #48 of 2026-05-13
- TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems 3 upvotes, #51 of 2026-05-13
- IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 3 upvotes, #51 of 2026-05-13
- WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting 3 upvotes, #51 of 2026-05-13
- MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics 2 upvotes, #54 of 2026-05-13
- Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation 2 upvotes, #54 of 2026-05-13
- FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation 1 upvotes, #56 of 2026-05-13
- Implicit Preference Alignment for Human Image Animation 1 upvotes, #56 of 2026-05-13
- Reliable Chain-of-Thought via Prefix Consistency 1 upvotes, #56 of 2026-05-13
- Large Language Models over Networks: Collaborative Intelligence under Resource Constraints 1 upvotes, #56 of 2026-05-13
- PAAC: Privacy-Aware Agentic Device-Cloud Collaboration 1 upvotes, #56 of 2026-05-13
- LLM Agents Already Know When to Call Tools -- Even Without Reasoning 1 upvotes, #56 of 2026-05-13
- FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning 1 upvotes, #56 of 2026-05-13
- GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction 1 upvotes, #56 of 2026-05-13
- SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation 1 upvotes, #56 of 2026-05-13
- Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction 1 upvotes, #56 of 2026-05-13
- Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty 1 upvotes, #56 of 2026-05-13
- ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging 1 upvotes, #56 of 2026-05-13
- Geometric Factual Recall in Transformers 1 upvotes, #56 of 2026-05-13
- EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory 1 upvotes, #69 of 2026-05-13
- Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception 2 upvotes, #69 of 2026-05-13
- EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera 0 upvotes, #69 of 2026-05-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.