Daily Papers of 2026-03-16
- LMEB: Long-horizon Memory Embedding Benchmark 71 upvotes, #1 of 2026-03-16
- Can Vision-Language Models Solve the Shell Game? 39 upvotes, #2 of 2026-03-16
- Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation 37 upvotes, #3 of 2026-03-16
- Multimodal OCR: Parse Anything from Documents 34 upvotes, #4 of 2026-03-16
- OmniForcing: Unleashing Real-time Joint Audio-Visual Generation 31 upvotes, #5 of 2026-03-16
- Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously 30 upvotes, #6 of 2026-03-16
- daVinci-Env: Open SWE Environment Synthesis at Scale 29 upvotes, #7 of 2026-03-16
- Visual-ERM: Reward Modeling for Visual Equivalence 21 upvotes, #8 of 2026-03-16
- MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning 19 upvotes, #9 of 2026-03-16
- Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents 16 upvotes, #10 of 2026-03-16
- EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery 15 upvotes, #11 of 2026-03-16
- From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space 14 upvotes, #12 of 2026-03-16
- V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration 13 upvotes, #13 of 2026-03-16
- HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios 11 upvotes, #14 of 2026-03-16
- HybridStitch: Pixel and Timestep Level Model Stitching for Diffusion Acceleration 10 upvotes, #15 of 2026-03-16
- Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models 8 upvotes, #16 of 2026-03-16
- VQQA: An Agentic Approach for Video Evaluation and Quality Improvement 7 upvotes, #17 of 2026-03-16
- Taking Shortcuts for Categorical VQA Using Super Neurons 6 upvotes, #18 of 2026-03-16
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation 6 upvotes, #18 of 2026-03-16
- CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges 6 upvotes, #18 of 2026-03-16
- Steve-Evolving: Open-World Embodied Self-Evolution via Fine-Grained Diagnosis and Dual-Track Knowledge Distillation 6 upvotes, #18 of 2026-03-16
- NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval 5 upvotes, #22 of 2026-03-16
- SimRecon: SimReady Compositional Scene Reconstruction from Real Videos 4 upvotes, #23 of 2026-03-16
- Compression Favors Consistency, Not Truth: When and Why Language Models Prefer Correct Information 3 upvotes, #24 of 2026-03-16
- Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection 3 upvotes, #24 of 2026-03-16
- Can Fairness Be Prompted? Prompt-Based Debiasing Strategies in High-Stakes Recommendations 3 upvotes, #24 of 2026-03-16
- ECoLAD: Deployment-Oriented Evaluation for Automotive Time-Series Anomaly Detection 1 upvotes, #27 of 2026-03-16
- SDF-Net: Structure-Aware Disentangled Feature Learning for Opticall-SAR Ship Re-identification 1 upvotes, #27 of 2026-03-16
- Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction 0 upvotes, #29 of 2026-03-16
- Detecting Intrinsic and Instrumental Self-Preservation in Autonomous Agents: The Unified Continuation-Interest Protocol 1 upvotes, #29 of 2026-03-16
- Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering 0 upvotes, #29 of 2026-03-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.