Microsoft Research
Microsoft Research on Hugging Face Daily Papers: 60 papers, 9 in the top 3 of their day, 3 paper of the day.
- Reinforcing Agentic Creativity in Scientific Ideation with Night Science 4 upvotes, #67 of 2026-09-29
- Agensh: Scaling Organizational Intelligence to 1,024 Agents 27 upvotes, #12 of 2026-09-23
- ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models 18 upvotes, #16 of 2026-09-22
- BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence 30 upvotes, #11 of 2026-09-21
- ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks 62 upvotes, #4 of 2026-09-17
- Expert-Space Exploration in MoE Reinforcement Learning 5 upvotes, #28 of 2026-09-15
- StudentSim: Training LLM-based Student Simulators 485 upvotes, #1 of 2026-09-02
- Dion3: Full-Stack Orthogonal Updates 9 upvotes, #19 of 2026-08-17
- Full-bandwidth transformer 22 upvotes, #13 of 2026-08-14
- Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 13 upvotes, #13 of 2026-08-10
- Weak-to-Strong On-Policy Distillation 56 upvotes, #4 of 2026-08-03
- Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale 13 upvotes, #26 of 2026-07-31
- LLMs Get Lost in Evolving User Intent 24 upvotes, #9 of 2026-07-24
- Multi-Turn On-Policy Distillation with Prefix Replay 12 upvotes, #14 of 2026-07-24
- LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks 14 upvotes, #19 of 2026-07-21
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources 141 upvotes, #2 of 2026-07-20
- ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving 31 upvotes, #3 of 2026-07-02
- Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement 7 upvotes, #17 of 2026-06-29
- ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction 4 upvotes, #27 of 2026-06-11
- Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts 52 upvotes, #5 of 2026-06-10
- Latent Spatial Memory for Video World Models 66 upvotes, #4 of 2026-06-09
- ECHO: Terminal Agents Learn World Models for Free 7 upvotes, #36 of 2026-05-26
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills 212 upvotes, #1 of 2026-05-25
- From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills 29 upvotes, #9 of 2026-05-25
- Video Models Can Reason with Verifiable Rewards 11 upvotes, #19 of 2026-05-20
- Orchard: An Open-Source Agentic Modeling Framework 19 upvotes, #20 of 2026-05-15
- Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR 16 upvotes, #16 of 2026-05-12
- World-R1: Reinforcing 3D Constraints for Text-to-Video Generation 34 upvotes, #2 of 2026-04-28
- Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs 1 upvotes, #35 of 2026-04-22
- MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation 6 upvotes, #20 of 2026-04-17
- Universal YOCO for Efficient Depth Scaling 17 upvotes, #13 of 2026-04-02
- BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation 11 upvotes, #23 of 2026-04-01
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? 47 upvotes, #2 of 2026-03-26
- Online Experiential Learning for Language Models 55 upvotes, #10 of 2026-03-18
- Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty 11 upvotes, #19 of 2026-03-17
- Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems 5 upvotes, #24 of 2026-03-10
- Scaling Agentic Capabilities, Not Context: Efficient Reinforcement Finetuning for Large Toolspaces 15 upvotes, #12 of 2026-03-10
- Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models 5 upvotes, #24 of 2026-03-10
- Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity 4 upvotes, #27 of 2026-03-10
- Proact-VL: A Proactive VideoLLM for Real-Time AI Companions 31 upvotes, #4 of 2026-03-05
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use 11 upvotes, #13 of 2026-03-04
- Reinforcement World Model Learning for LLM-based Agents 25 upvotes, #11 of 2026-02-06
- Self-Hinting Language Models Enhance Reinforcement Learning 27 upvotes, #13 of 2026-02-05
- Efficient Autoregressive Video Diffusion with Dummy Head 8 upvotes, #31 of 2026-02-05
- LIVE: Long-horizon Interactive Video World Modeling 12 upvotes, #19 of 2026-02-04
- Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 9 upvotes, #37 of 2026-02-03
- Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge 38 upvotes, #2 of 2026-01-20
- X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests 40 upvotes, #5 of 2026-01-13
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding 3 upvotes, #18 of 2025-12-04
- Black-Box On-Policy Distillation of Large Language Models 39 upvotes, #4 of 2025-11-14
- The Collaboration Gap 21 upvotes, #6 of 2025-11-05
- Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets 9 upvotes, #14 of 2025-10-31
- Code Aesthetics with Agentic Reward Feedback 7 upvotes, #21 of 2025-10-28
- LoongRL:Reinforcement Learning for Advanced Reasoning over Long Contexts 58 upvotes, #3 of 2025-10-23
- BitNet Distillation 47 upvotes, #8 of 2025-10-17
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning 1 upvotes, #36 of 2025-10-16
- DocReward: A Document Reward Model for Structuring and Stylizing 26 upvotes, #12 of 2025-10-14
- Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective 7 upvotes, #33 of 2025-10-01
- PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images 4 upvotes, #56 of 2025-09-30
- VibeVoice Technical Report 118 upvotes, #1 of 2025-08-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.