Daily Papers of 2026-05-27
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding 135 upvotes, #1 of 2026-05-27
- EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation 79 upvotes, #2 of 2026-05-27
- SpatialBench: Is Your Spatial Foundation Model an All-Round Player? 70 upvotes, #3 of 2026-05-27
- MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research 64 upvotes, #4 of 2026-05-27
- Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction 41 upvotes, #5 of 2026-05-27
- D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing 39 upvotes, #6 of 2026-05-27
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence 39 upvotes, #6 of 2026-05-27
- LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
- JLT: Clean-Latent Prediction in Latent Diffusion Transformers 32 upvotes, #9 of 2026-05-27
- Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling 31 upvotes, #10 of 2026-05-27
- LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence 27 upvotes, #11 of 2026-05-27
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation 25 upvotes, #12 of 2026-05-27
- QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents 24 upvotes, #13 of 2026-05-27
- Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini 23 upvotes, #14 of 2026-05-27
- Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models 20 upvotes, #15 of 2026-05-27
- Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective 20 upvotes, #15 of 2026-05-27
- VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions 19 upvotes, #17 of 2026-05-27
- Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning 16 upvotes, #18 of 2026-05-27
- RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models 16 upvotes, #18 of 2026-05-27
- Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement 16 upvotes, #18 of 2026-05-27
- Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments 16 upvotes, #18 of 2026-05-27
- Rethinking VLM Representation for VLA Initialization 15 upvotes, #22 of 2026-05-27
- MobileMoE: Scaling On-Device Mixture of Experts 14 upvotes, #23 of 2026-05-27
- DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs 12 upvotes, #24 of 2026-05-27
- Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals 12 upvotes, #24 of 2026-05-27
- Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation 11 upvotes, #26 of 2026-05-27
- Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows 9 upvotes, #27 of 2026-05-27
- SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent 9 upvotes, #27 of 2026-05-27
- ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models 8 upvotes, #29 of 2026-05-27
- FastKernels: Benchmarking GPU Kernel Generation in Production 8 upvotes, #29 of 2026-05-27
- MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale 8 upvotes, #29 of 2026-05-27
- Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents 7 upvotes, #32 of 2026-05-27
- Learning High-Frequency Continuous Action Chunks in Latent Space 6 upvotes, #33 of 2026-05-27
- CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations 6 upvotes, #33 of 2026-05-27
- Understanding Data Temporality Impact on Large Language Models Pre-training 5 upvotes, #35 of 2026-05-27
- NSF-SciFy: Mining the NSF Awards Database for Scientific Claims 4 upvotes, #36 of 2026-05-27
- EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration 4 upvotes, #36 of 2026-05-27
- STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media 4 upvotes, #36 of 2026-05-27
- Can LLMs Introspect? A Reality Check 4 upvotes, #36 of 2026-05-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.