Daily Papers of 2026-03-13
- Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training 90 upvotes, #1 of 2026-03-13
- Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections 62 upvotes, #2 of 2026-03-13
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse 51 upvotes, #3 of 2026-03-13
- Video-Based Reward Modeling for Computer-Use Agents 41 upvotes, #4 of 2026-03-13
- ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation 34 upvotes, #5 of 2026-03-13
- DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning 31 upvotes, #6 of 2026-03-13
- XSkill: Continual Learning from Experience and Skills in Multimodal Agents 29 upvotes, #7 of 2026-03-13
- DVD: Deterministic Video Depth Estimation with Generative Priors 26 upvotes, #8 of 2026-03-13
- WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing 24 upvotes, #9 of 2026-03-13
- Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation 23 upvotes, #10 of 2026-03-13
- One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers 17 upvotes, #11 of 2026-03-13
- RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning 15 upvotes, #12 of 2026-03-13
- CREATE: Testing LLMs for Associative Creativity 14 upvotes, #13 of 2026-03-13
- GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing 14 upvotes, #13 of 2026-03-13
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation 13 upvotes, #15 of 2026-03-13
- OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams 12 upvotes, #16 of 2026-03-13
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights 11 upvotes, #17 of 2026-03-13
- The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training 10 upvotes, #18 of 2026-03-13
- EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models 10 upvotes, #18 of 2026-03-13
- Mobile-GS: Real-time Gaussian Splatting for Mobile Devices 9 upvotes, #20 of 2026-03-13
- Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining 8 upvotes, #21 of 2026-03-13
- Meta-Reinforcement Learning with Self-Reflection for Agentic Search 8 upvotes, #21 of 2026-03-13
- Training Language Models via Neural Cellular Automata 7 upvotes, #23 of 2026-03-13
- Are Video Reasoning Models Ready to Go Outside? 7 upvotes, #23 of 2026-03-13
- Tiny Aya: Bridging Scale and Multilingual Depth 7 upvotes, #23 of 2026-03-13
- Geometric Autoencoder for Diffusion Models 6 upvotes, #26 of 2026-03-13
- FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System 6 upvotes, #26 of 2026-03-13
- Automatic Generation of High-Performance RL Environments 6 upvotes, #26 of 2026-03-13
- Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data 5 upvotes, #29 of 2026-03-13
- DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use 5 upvotes, #29 of 2026-03-13
- Coarse-Guided Visual Generation via Weighted h-Transform Sampling 5 upvotes, #29 of 2026-03-13
- PACED: Distillation at the Frontier of Student Competence 4 upvotes, #32 of 2026-03-13
- Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge 4 upvotes, #32 of 2026-03-13
- Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training 4 upvotes, #32 of 2026-03-13
- SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving 3 upvotes, #35 of 2026-03-13
- Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition 3 upvotes, #35 of 2026-03-13
- SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis 2 upvotes, #37 of 2026-03-13
- NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks 2 upvotes, #37 of 2026-03-13
- TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size 2 upvotes, #37 of 2026-03-13
- WaDi: Weight Direction-aware Distillation for One-step Image Synthesis 2 upvotes, #37 of 2026-03-13
- Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation 2 upvotes, #37 of 2026-03-13
- Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks 2 upvotes, #37 of 2026-03-13
- Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning 2 upvotes, #37 of 2026-03-13
- EmbTracker: Traceable Black-box Watermarking for Federated Language Models 2 upvotes, #37 of 2026-03-13
- A Mixed Diet Makes DINO An Omnivorous Vision Encoder 1 upvotes, #45 of 2026-03-13
- Causal Attribution of Coastal Water Clarity Degradation to Nickel Processing Expansion at the Indonesia Morowali Industrial Park, Sulawesi 0 upvotes, #46 of 2026-03-13
- 4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video 0 upvotes, #46 of 2026-03-13
- HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement 0 upvotes, #46 of 2026-03-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.