Daily Papers of 2026-03-18
- Demystifing Video Reasoning 356 upvotes, #1 of 2026-03-18
- InCoder-32B: Code Foundation Model for Industrial Scenarios 297 upvotes, #2 of 2026-03-18
- SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models 242 upvotes, #3 of 2026-03-18
- MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification 179 upvotes, #4 of 2026-03-18
- Qianfan-OCR: A Unified End-to-End Model for Document Intelligence 145 upvotes, #5 of 2026-03-18
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding 93 upvotes, #6 of 2026-03-18
- Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation 68 upvotes, #7 of 2026-03-18
- WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation 58 upvotes, #8 of 2026-03-18
- TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas 56 upvotes, #9 of 2026-03-18
- Online Experiential Learning for Language Models 55 upvotes, #10 of 2026-03-18
- FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use 43 upvotes, #11 of 2026-03-18
- WiT: Waypoint Diffusion Transformers via Trajectory Conflict Navigation 35 upvotes, #12 of 2026-03-18
- GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent 33 upvotes, #13 of 2026-03-18
- Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training 31 upvotes, #14 of 2026-03-18
- MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games 24 upvotes, #15 of 2026-03-18
- AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents 22 upvotes, #16 of 2026-03-18
- Omnilingual MT: Machine Translation for 1,600 Languages 19 upvotes, #17 of 2026-03-18
- SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? 18 upvotes, #18 of 2026-03-18
- SegviGen: Repurposing 3D Generative Model for Part Segmentation 18 upvotes, #18 of 2026-03-18
- Efficient Reasoning on the Edge 17 upvotes, #20 of 2026-03-18
- SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation 15 upvotes, #21 of 2026-03-18
- Semi-Autonomous Formalization of the Vlasov-Maxwell-Landau Equilibrium 14 upvotes, #22 of 2026-03-18
- One-Eval: An Agentic System for Automated and Traceable LLM Evaluation 11 upvotes, #23 of 2026-03-18
- Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context 11 upvotes, #23 of 2026-03-18
- Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning 10 upvotes, #25 of 2026-03-18
- M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM 10 upvotes, #25 of 2026-03-18
- FlashSampling: Fast and Memory-Efficient Exact Sampling 8 upvotes, #27 of 2026-03-18
- MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation 8 upvotes, #27 of 2026-03-18
- SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation 6 upvotes, #29 of 2026-03-18
- From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation 6 upvotes, #29 of 2026-03-18
- Residual Stream Duality in Modern Transformer Architectures 3 upvotes, #31 of 2026-03-18
- V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising 3 upvotes, #31 of 2026-03-18
- Sparking Scientific Creativity via LLM-Driven Interdisciplinary Inspiration 2 upvotes, #33 of 2026-03-18
- SuperLocalMemory V3: Information-Geometric Foundations for Zero-LLM Enterprise Agent Memory 2 upvotes, #33 of 2026-03-18
- VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents 2 upvotes, #33 of 2026-03-18
- CCTU: A Benchmark for Tool Use under Complex Constraints 2 upvotes, #33 of 2026-03-18
- I Know What I Don't Know: Latent Posterior Factor Models for Multi-Evidence Probabilistic Reasoning 2 upvotes, #33 of 2026-03-18
- Theoretical Foundations of Latent Posterior Factors: Formal Guarantees for Multi-Evidence Reasoning 2 upvotes, #33 of 2026-03-18
- ViT-AdaLA: Adapting Vision Transformers with Linear Attention 2 upvotes, #33 of 2026-03-18
- Mixture of Style Experts for Diverse Image Stylization 2 upvotes, #33 of 2026-03-18
- Anticipatory Planning for Multimodal AI Agents 2 upvotes, #33 of 2026-03-18
- Test-Time Strategies for More Efficient and Accurate Agentic RAG 1 upvotes, #42 of 2026-03-18
- ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation 1 upvotes, #42 of 2026-03-18
- Chain-of-Trajectories: Unlocking the Intrinsic Generative Optimality of Diffusion Models via Graph-Theoretic Planning 1 upvotes, #42 of 2026-03-18
- ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning 1 upvotes, #42 of 2026-03-18
- MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Compute-optimal Scaling of Diffusion Language Models 1 upvotes, #42 of 2026-03-18
- OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder 1 upvotes, #42 of 2026-03-18
- Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR 1 upvotes, #42 of 2026-03-18
- Learning Human-Object Interaction for 3D Human Pose Estimation from LiDAR Point Clouds 1 upvotes, #42 of 2026-03-18
- BERTology of Molecular Property Prediction 0 upvotes, #50 of 2026-03-18
- Measuring Primitive Accumulation: An Information-Theoretic Approach to Capitalist Enclosure in PIK2, Indonesia 0 upvotes, #50 of 2026-03-18
- HistoAtlas: A Pan-Cancer Morphology Atlas Linking Histomics to Molecular Programs and Clinical Outcomes 0 upvotes, #50 of 2026-03-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.