Daily Papers of 2026-05-18
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence 266 upvotes, #1 of 2026-05-18
- PhysBrain 1.0 Technical Report 141 upvotes, #2 of 2026-05-18
- MMSkills: Towards Multimodal Skills for General Visual Agents 117 upvotes, #3 of 2026-05-18
- FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization 62 upvotes, #4 of 2026-05-18
- Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation 58 upvotes, #5 of 2026-05-18
- Auditing Agent Harness Safety 54 upvotes, #6 of 2026-05-18
- DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo 50 upvotes, #7 of 2026-05-18
- Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding 40 upvotes, #8 of 2026-05-18
- Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization 36 upvotes, #9 of 2026-05-18
- InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation 34 upvotes, #10 of 2026-05-18
- Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR 33 upvotes, #11 of 2026-05-18
- ReactiveGWM: Steering NPC in Reactive Game World Models 28 upvotes, #12 of 2026-05-18
- Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution 22 upvotes, #13 of 2026-05-18
- Hölder Policy Optimisation 19 upvotes, #14 of 2026-05-18
- MetaAgent-X : Breaking the Ceiling of Automatic Multi-Agent Systems via End-to-End Reinforcement Learning 18 upvotes, #15 of 2026-05-18
- Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design 16 upvotes, #16 of 2026-05-18
- PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control 16 upvotes, #16 of 2026-05-18
- From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing 12 upvotes, #18 of 2026-05-18
- Unlocking Dense Metric Depth Estimation in VLMs 12 upvotes, #18 of 2026-05-18
- Steered LLM Activations are Non-Surjective 11 upvotes, #20 of 2026-05-18
- CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage 11 upvotes, #20 of 2026-05-18
- MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware 10 upvotes, #22 of 2026-05-18
- Look Before You Leap: Autonomous Exploration for LLM Agents 9 upvotes, #23 of 2026-05-18
- Efficient Image Synthesis with Sphere Latent Encoder 8 upvotes, #24 of 2026-05-18
- Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models 8 upvotes, #24 of 2026-05-18
- DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules 7 upvotes, #26 of 2026-05-18
- FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction 7 upvotes, #26 of 2026-05-18
- Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution 6 upvotes, #28 of 2026-05-18
- WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes 6 upvotes, #28 of 2026-05-18
- Follow the Mean: Reference-Guided Flow Matching 5 upvotes, #30 of 2026-05-18
- Learning POMDP World Models from Observations with Language-Model Priors 5 upvotes, #30 of 2026-05-18
- HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts 5 upvotes, #30 of 2026-05-18
- Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning 5 upvotes, #30 of 2026-05-18
- Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards 5 upvotes, #30 of 2026-05-18
- ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing 5 upvotes, #30 of 2026-05-18
- Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction 4 upvotes, #36 of 2026-05-18
- OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation 4 upvotes, #36 of 2026-05-18
- Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism 4 upvotes, #36 of 2026-05-18
- GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding 4 upvotes, #36 of 2026-05-18
- AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting 3 upvotes, #40 of 2026-05-18
- No One Knows the State of the Art in Geospatial Foundation Models 3 upvotes, #40 of 2026-05-18
- MLAIRE: Multilingual Language-Aware Information Retrieval Evaluation Protocal 2 upvotes, #42 of 2026-05-18
- Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces 2 upvotes, #42 of 2026-05-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.