Pavlo Molchanov
Pavlo Molchanov on Hugging Face Daily Papers: 34 papers, 17 in the top 3 of their day, 1,776 upvotes.
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization 191 upvotes, #1 of 2026-01-09
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 28 upvotes, #4 of 2025-12-25
- NVIDIA Nemotron 3: Efficient and Open Intelligence 27 upvotes, #5 of 2025-12-25
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed 12 upvotes, #17 of 2025-12-17
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration 100 upvotes, #2 of 2025-12-03
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models 29 upvotes, #7 of 2025-12-01
- Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs 22 upvotes, #9 of 2025-11-21
- TiDAR: Think in Diffusion, Talk in Autoregression 95 upvotes, #2 of 2025-11-13
- NVIDIA Nemotron Nano V2 VL 25 upvotes, #5 of 2025-11-07
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge 7 upvotes, #19 of 2025-10-23
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM 79 upvotes, #2 of 2025-10-20
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning 15 upvotes, #15 of 2025-10-20
- Fast-dLLM v2: Efficient Block-Diffusion LLM 47 upvotes, #5 of 2025-10-08
- BroRL: Scaling Reinforcement Learning via Broadened Exploration 16 upvotes, #10 of 2025-10-02
- 3D Aware Region Prompted Vision Language Model 12 upvotes, #9 of 2025-09-17
- Universal Deep Research: Bring Your Own Model and Strategy 12 upvotes, #22 of 2025-09-03
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model 31 upvotes, #7 of 2025-08-21
- Scaling RL to Long Videos 132 upvotes, #1 of 2025-07-11
- Small Language Models are the Future of Agentic AI 3 upvotes, #36 of 2025-06-05
- CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training 87 upvotes, #1 of 2025-04-18
- Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning 10 upvotes, #20 of 2025-04-16
- Scaling Vision Pre-Training to 4K Resolution 37 upvotes, #2 of 2025-03-26
- NVILA: Efficient Frontier Visual Language Models 47 upvotes, #3 of 2024-12-06
- Hymba: A Hybrid-head Architecture for Small Language Models 37 upvotes, #4 of 2024-11-22
- EoRA: Training-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation 6 upvotes, #14 of 2024-10-29
- PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation 31 upvotes, #2 of 2024-10-03
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models 43 upvotes, #1 of 2024-09-27
- LLM Pruning and Distillation in Practice: The Minitron Approach 48 upvotes, #2 of 2024-08-22
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos 50 upvotes, #1 of 2024-08-20
- VILA^2: VILA Augmented VILA 36 upvotes, #2 of 2024-07-25
- Compact Language Models via Pruning and Knowledge Distillation 32 upvotes, #1 of 2024-07-23
- LITA: Language Instructed Temporal-Localization Assistant 16 upvotes, #2 of 2024-03-29
- VILA: On Pre-training for Visual Language Models 21 upvotes, #3 of 2023-12-13
- FasterViT: Fast Vision Transformers with Hierarchical Attention 32 upvotes, #1 of 2023-06-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.