Daily Papers of 2026-03-13

  1. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training 90 upvotes, #1 of 2026-03-13
  2. Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections 62 upvotes, #2 of 2026-03-13
  3. IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse 51 upvotes, #3 of 2026-03-13
  4. Video-Based Reward Modeling for Computer-Use Agents 41 upvotes, #4 of 2026-03-13
  5. ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation 34 upvotes, #5 of 2026-03-13
  6. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning 31 upvotes, #6 of 2026-03-13
  7. XSkill: Continual Learning from Experience and Skills in Multimodal Agents 29 upvotes, #7 of 2026-03-13
  8. DVD: Deterministic Video Depth Estimation with Generative Priors 26 upvotes, #8 of 2026-03-13
  9. WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing 24 upvotes, #9 of 2026-03-13
  10. Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation 23 upvotes, #10 of 2026-03-13
  11. One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers 17 upvotes, #11 of 2026-03-13
  12. RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning 15 upvotes, #12 of 2026-03-13
  13. CREATE: Testing LLMs for Associative Creativity 14 upvotes, #13 of 2026-03-13
  14. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing 14 upvotes, #13 of 2026-03-13
  15. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation 13 upvotes, #15 of 2026-03-13
  16. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams 12 upvotes, #16 of 2026-03-13
  17. Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights 11 upvotes, #17 of 2026-03-13
  18. The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training 10 upvotes, #18 of 2026-03-13
  19. EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models 10 upvotes, #18 of 2026-03-13
  20. Mobile-GS: Real-time Gaussian Splatting for Mobile Devices 9 upvotes, #20 of 2026-03-13
  21. Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining 8 upvotes, #21 of 2026-03-13
  22. Meta-Reinforcement Learning with Self-Reflection for Agentic Search 8 upvotes, #21 of 2026-03-13
  23. Training Language Models via Neural Cellular Automata 7 upvotes, #23 of 2026-03-13
  24. Are Video Reasoning Models Ready to Go Outside? 7 upvotes, #23 of 2026-03-13
  25. Tiny Aya: Bridging Scale and Multilingual Depth 7 upvotes, #23 of 2026-03-13
  26. Geometric Autoencoder for Diffusion Models 6 upvotes, #26 of 2026-03-13
  27. FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System 6 upvotes, #26 of 2026-03-13
  28. Automatic Generation of High-Performance RL Environments 6 upvotes, #26 of 2026-03-13
  29. Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data 5 upvotes, #29 of 2026-03-13
  30. DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use 5 upvotes, #29 of 2026-03-13
  31. Coarse-Guided Visual Generation via Weighted h-Transform Sampling 5 upvotes, #29 of 2026-03-13
  32. PACED: Distillation at the Frontier of Student Competence 4 upvotes, #32 of 2026-03-13
  33. Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge 4 upvotes, #32 of 2026-03-13
  34. Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training 4 upvotes, #32 of 2026-03-13
  35. SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving 3 upvotes, #35 of 2026-03-13
  36. Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition 3 upvotes, #35 of 2026-03-13
  37. SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis 2 upvotes, #37 of 2026-03-13
  38. NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks 2 upvotes, #37 of 2026-03-13
  39. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size 2 upvotes, #37 of 2026-03-13
  40. WaDi: Weight Direction-aware Distillation for One-step Image Synthesis 2 upvotes, #37 of 2026-03-13
  41. Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation 2 upvotes, #37 of 2026-03-13
  42. Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks 2 upvotes, #37 of 2026-03-13
  43. Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning 2 upvotes, #37 of 2026-03-13
  44. EmbTracker: Traceable Black-box Watermarking for Federated Language Models 2 upvotes, #37 of 2026-03-13
  45. A Mixed Diet Makes DINO An Omnivorous Vision Encoder 1 upvotes, #45 of 2026-03-13
  46. Causal Attribution of Coastal Water Clarity Degradation to Nickel Processing Expansion at the Indonesia Morowali Industrial Park, Sulawesi 0 upvotes, #46 of 2026-03-13
  47. 4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video 0 upvotes, #46 of 2026-03-13
  48. HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement 0 upvotes, #46 of 2026-03-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.