Daily Papers of 2026-03-20
- Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding 92 upvotes, #1 of 2026-03-20
- SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing 65 upvotes, #2 of 2026-03-20
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 60 upvotes, #3 of 2026-03-20
- 3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model 57 upvotes, #4 of 2026-03-20
- FASTER: Rethinking Real-Time Flow VLAs 55 upvotes, #5 of 2026-03-20
- Memento-Skills: Let Agents Design Agents 53 upvotes, #6 of 2026-03-20
- Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer 41 upvotes, #7 of 2026-03-20
- MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction 36 upvotes, #8 of 2026-03-20
- Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens 33 upvotes, #9 of 2026-03-20
- F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World 30 upvotes, #10 of 2026-03-20
- LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs 28 upvotes, #11 of 2026-03-20
- AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents 26 upvotes, #12 of 2026-03-20
- ReactMotion: Generating Reactive Listener Motions from Speaker Utterance 24 upvotes, #13 of 2026-03-20
- VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining 21 upvotes, #14 of 2026-03-20
- Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding 19 upvotes, #15 of 2026-03-20
- EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing 18 upvotes, #16 of 2026-03-20
- Tinted Frames: Question Framing Blinds Vision-Language Models 16 upvotes, #17 of 2026-03-20
- SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation 15 upvotes, #18 of 2026-03-20
- ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents 14 upvotes, #19 of 2026-03-20
- Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models 13 upvotes, #20 of 2026-03-20
- MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning 12 upvotes, #21 of 2026-03-20
- Matryoshka Gaussian Splatting 11 upvotes, #22 of 2026-03-20
- MOSS-TTS Technical Report 10 upvotes, #23 of 2026-03-20
- OSM-based Domain Adaptation for Remote Sensing VLMs 7 upvotes, #24 of 2026-03-20
- Reasoning over mathematical objects: on-policy reward modeling and test time aggregation 6 upvotes, #25 of 2026-03-20
- Prompt-Free Universal Region Proposal Network 3 upvotes, #26 of 2026-03-20
- What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time? 3 upvotes, #26 of 2026-03-20
- Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation 2 upvotes, #28 of 2026-03-20
- COT-FM: Cluster-wise Optimal Transport Flow Matching 2 upvotes, #28 of 2026-03-20
- VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction 1 upvotes, #30 of 2026-03-20
- PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark 1 upvotes, #30 of 2026-03-20
- DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising 1 upvotes, #30 of 2026-03-20
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.