Daily Papers of 2026-03-11
- Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing 141 upvotes, #1 of 2026-03-11
- Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs 70 upvotes, #2 of 2026-03-11
- MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data 51 upvotes, #3 of 2026-03-11
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion 48 upvotes, #4 of 2026-03-11
- InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing 47 upvotes, #5 of 2026-03-11
- Fish Audio S2 Technical Report 33 upvotes, #6 of 2026-03-11
- Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs 28 upvotes, #7 of 2026-03-11
- Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports 26 upvotes, #8 of 2026-03-11
- MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants 14 upvotes, #9 of 2026-03-11
- Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering 12 upvotes, #10 of 2026-03-11
- VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning? 9 upvotes, #11 of 2026-03-11
- Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards 9 upvotes, #11 of 2026-03-11
- Do What I Say: A Spoken Prompt Dataset for Instruction-Following 9 upvotes, #11 of 2026-03-11
- Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications 7 upvotes, #14 of 2026-03-11
- ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning 5 upvotes, #15 of 2026-03-11
- The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness 5 upvotes, #15 of 2026-03-11
- Streaming Autoregressive Video Generation via Diagonal Distillation 5 upvotes, #15 of 2026-03-11
- Towards a Neural Debugger for Python 5 upvotes, #15 of 2026-03-11
- Multi-Head Low-Rank Attention 3 upvotes, #19 of 2026-03-11
- BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation 2 upvotes, #20 of 2026-03-11
- Reward Prediction with Factorized World States 2 upvotes, #20 of 2026-03-11
- SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement 1 upvotes, #22 of 2026-03-11
- Bolbosh: Script-Aware Flow Matching for Kashmiri Text-to-Speech 1 upvotes, #22 of 2026-03-11
- ConFu: Contemplate the Future for Better Speculative Sampling 1 upvotes, #22 of 2026-03-11
- BiCLIP: Domain Canonicalization via Structured Geometric Transformation 1 upvotes, #22 of 2026-03-11
- Compiler-First State Space Duality and Portable O(1) Autoregressive Caching for Inference 1 upvotes, #22 of 2026-03-11
- TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery 0 upvotes, #27 of 2026-03-11
- Micro-Diffusion Compression -- Binary Tree Tweedie Denoising for Online Probability Estimation 1 upvotes, #27 of 2026-03-11
- A Text-Native Interface for Generative Video Authoring 0 upvotes, #27 of 2026-03-11
- Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control 1 upvotes, #27 of 2026-03-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.