Daily Papers of 2026-03-20

  1. Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding 92 upvotes, #1 of 2026-03-20
  2. SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing 65 upvotes, #2 of 2026-03-20
  3. Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 60 upvotes, #3 of 2026-03-20
  4. 3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model 57 upvotes, #4 of 2026-03-20
  5. FASTER: Rethinking Real-Time Flow VLAs 55 upvotes, #5 of 2026-03-20
  6. Memento-Skills: Let Agents Design Agents 53 upvotes, #6 of 2026-03-20
  7. Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer 41 upvotes, #7 of 2026-03-20
  8. MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction 36 upvotes, #8 of 2026-03-20
  9. Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens 33 upvotes, #9 of 2026-03-20
  10. F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World 30 upvotes, #10 of 2026-03-20
  11. LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs 28 upvotes, #11 of 2026-03-20
  12. AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents 26 upvotes, #12 of 2026-03-20
  13. ReactMotion: Generating Reactive Listener Motions from Speaker Utterance 24 upvotes, #13 of 2026-03-20
  14. VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining 21 upvotes, #14 of 2026-03-20
  15. Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding 19 upvotes, #15 of 2026-03-20
  16. EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing 18 upvotes, #16 of 2026-03-20
  17. Tinted Frames: Question Framing Blinds Vision-Language Models 16 upvotes, #17 of 2026-03-20
  18. SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation 15 upvotes, #18 of 2026-03-20
  19. ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents 14 upvotes, #19 of 2026-03-20
  20. Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models 13 upvotes, #20 of 2026-03-20
  21. MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning 12 upvotes, #21 of 2026-03-20
  22. Matryoshka Gaussian Splatting 11 upvotes, #22 of 2026-03-20
  23. MOSS-TTS Technical Report 10 upvotes, #23 of 2026-03-20
  24. OSM-based Domain Adaptation for Remote Sensing VLMs 7 upvotes, #24 of 2026-03-20
  25. Reasoning over mathematical objects: on-policy reward modeling and test time aggregation 6 upvotes, #25 of 2026-03-20
  26. Prompt-Free Universal Region Proposal Network 3 upvotes, #26 of 2026-03-20
  27. What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time? 3 upvotes, #26 of 2026-03-20
  28. Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation 2 upvotes, #28 of 2026-03-20
  29. COT-FM: Cluster-wise Optimal Transport Flow Matching 2 upvotes, #28 of 2026-03-20
  30. VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction 1 upvotes, #30 of 2026-03-20
  31. PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark 1 upvotes, #30 of 2026-03-20
  32. DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising 1 upvotes, #30 of 2026-03-20

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.