Daily Papers of 2026-08-17
- Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination 277 upvotes, #1 of 2026-08-17
- Self-Supervised Visual On-Policy Distillation 169 upvotes, #2 of 2026-08-17
- Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence 58 upvotes, #3 of 2026-08-17
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development 55 upvotes, #4 of 2026-08-17
- SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning 34 upvotes, #5 of 2026-08-17
- Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning 33 upvotes, #6 of 2026-08-17
- Marionette: Predicting World States, Rendering Geometry, Painting Appearance 33 upvotes, #6 of 2026-08-17
- DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data 32 upvotes, #8 of 2026-08-17
- MobileMem: Learning from a Year of Mobile Experiences 25 upvotes, #9 of 2026-08-17
- Latent On-Policy Self-Distillation 24 upvotes, #10 of 2026-08-17
- HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark 17 upvotes, #11 of 2026-08-17
- Second Thought: Reasoning in Parallel as LLM Agents Act and Observe 16 upvotes, #12 of 2026-08-17
- CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing 16 upvotes, #12 of 2026-08-17
- Modular Cognitive Architecture Emerges in Large Language Models 15 upvotes, #14 of 2026-08-17
- Scaling Domain Data Repetition in LLM Pretraining 15 upvotes, #14 of 2026-08-17
- PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment 14 upvotes, #16 of 2026-08-17
- Claim-Level Reliability Assessment for Efficient Test-Time Reasoning 12 upvotes, #17 of 2026-08-17
- Multimodal Model Diffing for Feature Discovery and Control 10 upvotes, #18 of 2026-08-17
- Dion3: Full-Stack Orthogonal Updates 9 upvotes, #19 of 2026-08-17
- Forecast Collapse in Time-Series Foundation Models 9 upvotes, #19 of 2026-08-17
- LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure 8 upvotes, #21 of 2026-08-17
- Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems 6 upvotes, #22 of 2026-08-17
- Verifier-Induced Support Reshaping in On-Policy Optimization 5 upvotes, #23 of 2026-08-17
- UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations 4 upvotes, #24 of 2026-08-17
- Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead 4 upvotes, #24 of 2026-08-17
- A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images 4 upvotes, #24 of 2026-08-17
- Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction 3 upvotes, #27 of 2026-08-17
- Is this Citation on Point? 3 upvotes, #27 of 2026-08-17
- Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models 3 upvotes, #27 of 2026-08-17
- UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers 2 upvotes, #30 of 2026-08-17
- Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings 2 upvotes, #30 of 2026-08-17
- SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation 1 upvotes, #32 of 2026-08-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.