Daily Papers of 2026-03-11

  1. Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing 141 upvotes, #1 of 2026-03-11
  2. Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs 70 upvotes, #2 of 2026-03-11
  3. MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data 51 upvotes, #3 of 2026-03-11
  4. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion 48 upvotes, #4 of 2026-03-11
  5. InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing 47 upvotes, #5 of 2026-03-11
  6. Fish Audio S2 Technical Report 33 upvotes, #6 of 2026-03-11
  7. Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs 28 upvotes, #7 of 2026-03-11
  8. Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports 26 upvotes, #8 of 2026-03-11
  9. MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants 14 upvotes, #9 of 2026-03-11
  10. Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering 12 upvotes, #10 of 2026-03-11
  11. VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning? 9 upvotes, #11 of 2026-03-11
  12. Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards 9 upvotes, #11 of 2026-03-11
  13. Do What I Say: A Spoken Prompt Dataset for Instruction-Following 9 upvotes, #11 of 2026-03-11
  14. Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications 7 upvotes, #14 of 2026-03-11
  15. ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning 5 upvotes, #15 of 2026-03-11
  16. The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness 5 upvotes, #15 of 2026-03-11
  17. Streaming Autoregressive Video Generation via Diagonal Distillation 5 upvotes, #15 of 2026-03-11
  18. Towards a Neural Debugger for Python 5 upvotes, #15 of 2026-03-11
  19. Multi-Head Low-Rank Attention 3 upvotes, #19 of 2026-03-11
  20. BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation 2 upvotes, #20 of 2026-03-11
  21. Reward Prediction with Factorized World States 2 upvotes, #20 of 2026-03-11
  22. SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement 1 upvotes, #22 of 2026-03-11
  23. Bolbosh: Script-Aware Flow Matching for Kashmiri Text-to-Speech 1 upvotes, #22 of 2026-03-11
  24. ConFu: Contemplate the Future for Better Speculative Sampling 1 upvotes, #22 of 2026-03-11
  25. BiCLIP: Domain Canonicalization via Structured Geometric Transformation 1 upvotes, #22 of 2026-03-11
  26. Compiler-First State Space Duality and Portable O(1) Autoregressive Caching for Inference 1 upvotes, #22 of 2026-03-11
  27. TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery 0 upvotes, #27 of 2026-03-11
  28. Micro-Diffusion Compression -- Binary Tree Tweedie Denoising for Online Probability Estimation 1 upvotes, #27 of 2026-03-11
  29. A Text-Native Interface for Generative Video Authoring 0 upvotes, #27 of 2026-03-11
  30. Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control 1 upvotes, #27 of 2026-03-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.