NVIDIA

NVIDIA on Hugging Face Daily Papers: 164 papers, 27 in the top 3 of their day, 8 paper of the day.

  1. Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents 113 upvotes, #7 of 2026-10-01
  2. Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model 15 upvotes, #37 of 2026-10-01
  3. PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents 21 upvotes, #30 of 2026-10-01
  4. PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation 28 upvotes, #28 of 2026-10-01
  5. Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards 9 upvotes, #57 of 2026-09-30
  6. LongLive-Plug: Once-for-All Distillation for Video Generation 41 upvotes, #27 of 2026-09-30
  7. An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning 25 upvotes, #32 of 2026-09-29
  8. TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining 18 upvotes, #42 of 2026-09-29
  9. SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness 130 upvotes, #2 of 2026-09-18
  10. Agora: Git as Shared Memory for Collective AutoResearch 57 upvotes, #5 of 2026-09-17
  11. An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics 40 upvotes, #6 of 2026-09-11
  12. Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training 35 upvotes, #16 of 2026-09-09
  13. ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding 19 upvotes, #34 of 2026-09-09
  14. Post-Training Language Models for Gold-Medal Performance in Coding Competitions 11 upvotes, #19 of 2026-09-03
  15. Hydra-0: Action Flow for Generalist World Modeling and Control 10 upvotes, #12 of 2026-08-24
  16. UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations 4 upvotes, #24 of 2026-08-17
  17. Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation 18 upvotes, #15 of 2026-08-14
  18. Addressable Memory for Video World Models 14 upvotes, #10 of 2026-08-10
  19. Voice Memory for Agentic Speech Recognition 11 upvotes, #16 of 2026-07-30
  20. Parallel Decoding Distillation for Fast Image and Video Generation 16 upvotes, #13 of 2026-07-29
  21. Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification 35 upvotes, #8 of 2026-07-28
  22. Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning 32 upvotes, #3 of 2026-07-27
  23. NVIDIA-labs OO Agents: Native Python Object-Oriented Agents 33 upvotes, #7 of 2026-07-24
  24. SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation 39 upvotes, #5 of 2026-07-24
  25. Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 11 upvotes, #19 of 2026-07-20
  26. RoboTTT: Context Scaling for Robot Policies 22 upvotes, #13 of 2026-07-17
  27. Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE 23 upvotes, #8 of 2026-07-10
  28. ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation 10 upvotes, #14 of 2026-07-10
  29. WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence 14 upvotes, #6 of 2026-07-09
  30. Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model 14 upvotes, #16 of 2026-07-08
  31. Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding 12 upvotes, #18 of 2026-07-08
  32. Unified Audio Intelligence Without Regressing on Text Intelligence 20 upvotes, #18 of 2026-07-07
  33. ASPIRE: Agentic /Skills Discovery for Robotics 22 upvotes, #7 of 2026-07-02
  34. Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis 14 upvotes, #19 of 2026-06-30
  35. One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications 22 upvotes, #13 of 2026-06-30
  36. Vesta: A Generalist Embodied Reasoning Model 10 upvotes, #13 of 2026-06-29
  37. SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation 14 upvotes, #9 of 2026-06-29
  38. PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation 50 upvotes, #2 of 2026-06-29
  39. Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models 24 upvotes, #10 of 2026-06-25
  40. ENPIRE: Agentic Robot Policy Self-Improvement in the Real World 14 upvotes, #14 of 2026-06-19
  41. Adaptive Volumetric Mechanical Property Fields Invariant to Resolution 5 upvotes, #25 of 2026-06-19
  42. Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients 60 upvotes, #4 of 2026-06-17
  43. ProCUA-SFT Technical Report 9 upvotes, #19 of 2026-06-17
  44. Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 15 upvotes, #16 of 2026-06-16
  45. SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning 101 upvotes, #3 of 2026-06-12
  46. SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference 3 upvotes, #29 of 2026-06-11
  47. VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation 5 upvotes, #34 of 2026-06-09
  48. GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors 8 upvotes, #23 of 2026-06-04
  49. Cosmos 3: Omnimodal World Models for Physical AI 115 upvotes, #1 of 2026-06-04
  50. NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation 22 upvotes, #12 of 2026-06-03
  51. Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching 22 upvotes, #12 of 2026-06-03
  52. FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes 1 upvotes, #58 of 2026-06-02
  53. LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation 19 upvotes, #17 of 2026-06-02
  54. SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer 36 upvotes, #12 of 2026-06-01
  55. Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models 2 upvotes, #53 of 2026-06-01
  56. Why Far Looks Up: Probing Spatial Representation in Vision-Language Models 59 upvotes, #5 of 2026-05-29
  57. Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players 419 upvotes, #1 of 2026-05-28
  58. Agent Explorative Policy Optimization for Multimodal Agentic Reasoning 87 upvotes, #2 of 2026-05-28
  59. Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving 17 upvotes, #22 of 2026-05-28
  60. LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding 135 upvotes, #1 of 2026-05-27
  61. PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion 45 upvotes, #6 of 2026-05-25
  62. Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention 30 upvotes, #13 of 2026-05-22
  63. Fast 4D Mesh Generation by Spatio-Temporal Attention Chains 10 upvotes, #21 of 2026-05-20
  64. LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation 109 upvotes, #3 of 2026-05-19
  65. MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models 73 upvotes, #5 of 2026-05-15
  66. SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer 80 upvotes, #4 of 2026-05-15
  67. Retrieval from Within: An Intrinsic Capability of Attention-Based Models 5 upvotes, #30 of 2026-05-14
  68. AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation 96 upvotes, #3 of 2026-05-14
  69. Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence 17 upvotes, #15 of 2026-05-01
  70. Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding 14 upvotes, #6 of 2026-04-30
  71. Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips 57 upvotes, #3 of 2026-04-20
  72. RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies 7 upvotes, #16 of 2026-04-20
  73. Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 36 upvotes, #4 of 2026-04-15
  74. Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation 12 upvotes, #19 of 2026-04-15
  75. Lyra 2.0: Explorable Generative 3D Worlds 37 upvotes, #3 of 2026-04-15
  76. Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music 28 upvotes, #12 of 2026-04-14
  77. SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding 10 upvotes, #23 of 2026-04-14
  78. MoRight: Motion Control Done Right 7 upvotes, #22 of 2026-04-09
  79. FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling 34 upvotes, #5 of 2026-04-09
  80. TriAttention: Efficient Long Reasoning with Trigonometric KV Compression 106 upvotes, #4 of 2026-04-07
  81. PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost 18 upvotes, #16 of 2026-03-24
  82. ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents 14 upvotes, #19 of 2026-03-20
  83. Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 60 upvotes, #3 of 2026-03-20
  84. MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos 13 upvotes, #18 of 2026-03-17
  85. MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data 51 upvotes, #3 of 2026-03-11
  86. Mode Seeking meets Mean Seeking for Fast Long Video Generation 38 upvotes, #5 of 2026-03-02
  87. VGG-T^3: Offline Feed-Forward 3D Reconstruction at Scale 13 upvotes, #10 of 2026-02-27
  88. On Data Engineering for Scaling LLM Terminal Capabilities 90 upvotes, #1 of 2026-02-25
  89. Test-Time Training with KV Binding Is Secretly Linear Attention 29 upvotes, #3 of 2026-02-25
  90. Spanning the Visual Analogy Space with a Weight Basis of LoRAs 14 upvotes, #5 of 2026-02-23
  91. PhyCritic: Multimodal Critic Models for Physical AI 51 upvotes, #4 of 2026-02-12
  92. iGRPO: Self-Feedback-Driven LLM Reasoning 15 upvotes, #20 of 2026-02-11
  93. SAGE: Scalable Agentic 3D Scene Generation for Embodied AI 7 upvotes, #32 of 2026-02-11
  94. DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos 30 upvotes, #8 of 2026-02-09
  95. Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch 9 upvotes, #23 of 2026-02-04
  96. Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text 89 upvotes, #2 of 2026-02-02
  97. FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning 6 upvotes, #13 of 2026-01-29
  98. Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow 21 upvotes, #6 of 2026-01-26
  99. Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning 13 upvotes, #14 of 2026-01-23
  100. Transition Matching Distillation for Fast Video Generation 31 upvotes, #11 of 2026-01-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.