Daily Papers of 2026-05-18

  1. CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence 266 upvotes, #1 of 2026-05-18
  2. PhysBrain 1.0 Technical Report 141 upvotes, #2 of 2026-05-18
  3. MMSkills: Towards Multimodal Skills for General Visual Agents 117 upvotes, #3 of 2026-05-18
  4. FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization 62 upvotes, #4 of 2026-05-18
  5. Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation 58 upvotes, #5 of 2026-05-18
  6. Auditing Agent Harness Safety 54 upvotes, #6 of 2026-05-18
  7. DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo 50 upvotes, #7 of 2026-05-18
  8. Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding 40 upvotes, #8 of 2026-05-18
  9. Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization 36 upvotes, #9 of 2026-05-18
  10. InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation 34 upvotes, #10 of 2026-05-18
  11. Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR 33 upvotes, #11 of 2026-05-18
  12. ReactiveGWM: Steering NPC in Reactive Game World Models 28 upvotes, #12 of 2026-05-18
  13. Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution 22 upvotes, #13 of 2026-05-18
  14. Hölder Policy Optimisation 19 upvotes, #14 of 2026-05-18
  15. MetaAgent-X : Breaking the Ceiling of Automatic Multi-Agent Systems via End-to-End Reinforcement Learning 18 upvotes, #15 of 2026-05-18
  16. Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design 16 upvotes, #16 of 2026-05-18
  17. PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control 16 upvotes, #16 of 2026-05-18
  18. From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing 12 upvotes, #18 of 2026-05-18
  19. Unlocking Dense Metric Depth Estimation in VLMs 12 upvotes, #18 of 2026-05-18
  20. Steered LLM Activations are Non-Surjective 11 upvotes, #20 of 2026-05-18
  21. CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage 11 upvotes, #20 of 2026-05-18
  22. MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware 10 upvotes, #22 of 2026-05-18
  23. Look Before You Leap: Autonomous Exploration for LLM Agents 9 upvotes, #23 of 2026-05-18
  24. Efficient Image Synthesis with Sphere Latent Encoder 8 upvotes, #24 of 2026-05-18
  25. Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models 8 upvotes, #24 of 2026-05-18
  26. DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules 7 upvotes, #26 of 2026-05-18
  27. FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction 7 upvotes, #26 of 2026-05-18
  28. Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution 6 upvotes, #28 of 2026-05-18
  29. WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes 6 upvotes, #28 of 2026-05-18
  30. Follow the Mean: Reference-Guided Flow Matching 5 upvotes, #30 of 2026-05-18
  31. Learning POMDP World Models from Observations with Language-Model Priors 5 upvotes, #30 of 2026-05-18
  32. HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts 5 upvotes, #30 of 2026-05-18
  33. Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning 5 upvotes, #30 of 2026-05-18
  34. Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards 5 upvotes, #30 of 2026-05-18
  35. ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing 5 upvotes, #30 of 2026-05-18
  36. Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction 4 upvotes, #36 of 2026-05-18
  37. OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation 4 upvotes, #36 of 2026-05-18
  38. Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism 4 upvotes, #36 of 2026-05-18
  39. GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding 4 upvotes, #36 of 2026-05-18
  40. AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting 3 upvotes, #40 of 2026-05-18
  41. No One Knows the State of the Art in Geospatial Foundation Models 3 upvotes, #40 of 2026-05-18
  42. MLAIRE: Multilingual Language-Aware Information Retrieval Evaluation Protocal 2 upvotes, #42 of 2026-05-18
  43. Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces 2 upvotes, #42 of 2026-05-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.