Microsoft Research

Microsoft Research on Hugging Face Daily Papers: 60 papers, 9 in the top 3 of their day, 3 paper of the day.

  1. Reinforcing Agentic Creativity in Scientific Ideation with Night Science 4 upvotes, #67 of 2026-09-29
  2. Agensh: Scaling Organizational Intelligence to 1,024 Agents 27 upvotes, #12 of 2026-09-23
  3. ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models 18 upvotes, #16 of 2026-09-22
  4. BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence 30 upvotes, #11 of 2026-09-21
  5. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks 62 upvotes, #4 of 2026-09-17
  6. Expert-Space Exploration in MoE Reinforcement Learning 5 upvotes, #28 of 2026-09-15
  7. StudentSim: Training LLM-based Student Simulators 485 upvotes, #1 of 2026-09-02
  8. Dion3: Full-Stack Orthogonal Updates 9 upvotes, #19 of 2026-08-17
  9. Full-bandwidth transformer 22 upvotes, #13 of 2026-08-14
  10. Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 13 upvotes, #13 of 2026-08-10
  11. Weak-to-Strong On-Policy Distillation 56 upvotes, #4 of 2026-08-03
  12. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale 13 upvotes, #26 of 2026-07-31
  13. LLMs Get Lost in Evolving User Intent 24 upvotes, #9 of 2026-07-24
  14. Multi-Turn On-Policy Distillation with Prefix Replay 12 upvotes, #14 of 2026-07-24
  15. LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks 14 upvotes, #19 of 2026-07-21
  16. RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources 141 upvotes, #2 of 2026-07-20
  17. ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving 31 upvotes, #3 of 2026-07-02
  18. Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement 7 upvotes, #17 of 2026-06-29
  19. ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction 4 upvotes, #27 of 2026-06-11
  20. Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts 52 upvotes, #5 of 2026-06-10
  21. Latent Spatial Memory for Video World Models 66 upvotes, #4 of 2026-06-09
  22. ECHO: Terminal Agents Learn World Models for Free 7 upvotes, #36 of 2026-05-26
  23. SkillOpt: Executive Strategy for Self-Evolving Agent Skills 212 upvotes, #1 of 2026-05-25
  24. From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills 29 upvotes, #9 of 2026-05-25
  25. Video Models Can Reason with Verifiable Rewards 11 upvotes, #19 of 2026-05-20
  26. Orchard: An Open-Source Agentic Modeling Framework 19 upvotes, #20 of 2026-05-15
  27. Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR 16 upvotes, #16 of 2026-05-12
  28. World-R1: Reinforcing 3D Constraints for Text-to-Video Generation 34 upvotes, #2 of 2026-04-28
  29. Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs 1 upvotes, #35 of 2026-04-22
  30. MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation 6 upvotes, #20 of 2026-04-17
  31. Universal YOCO for Efficient Depth Scaling 17 upvotes, #13 of 2026-04-02
  32. BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation 11 upvotes, #23 of 2026-04-01
  33. Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? 47 upvotes, #2 of 2026-03-26
  34. Online Experiential Learning for Language Models 55 upvotes, #10 of 2026-03-18
  35. Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty 11 upvotes, #19 of 2026-03-17
  36. Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems 5 upvotes, #24 of 2026-03-10
  37. Scaling Agentic Capabilities, Not Context: Efficient Reinforcement Finetuning for Large Toolspaces 15 upvotes, #12 of 2026-03-10
  38. Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models 5 upvotes, #24 of 2026-03-10
  39. Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity 4 upvotes, #27 of 2026-03-10
  40. Proact-VL: A Proactive VideoLLM for Real-Time AI Companions 31 upvotes, #4 of 2026-03-05
  41. Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use 11 upvotes, #13 of 2026-03-04
  42. Reinforcement World Model Learning for LLM-based Agents 25 upvotes, #11 of 2026-02-06
  43. Self-Hinting Language Models Enhance Reinforcement Learning 27 upvotes, #13 of 2026-02-05
  44. Efficient Autoregressive Video Diffusion with Dummy Head 8 upvotes, #31 of 2026-02-05
  45. LIVE: Long-horizon Interactive Video World Modeling 12 upvotes, #19 of 2026-02-04
  46. Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 9 upvotes, #37 of 2026-02-03
  47. Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge 38 upvotes, #2 of 2026-01-20
  48. X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests 40 upvotes, #5 of 2026-01-13
  49. Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding 3 upvotes, #18 of 2025-12-04
  50. Black-Box On-Policy Distillation of Large Language Models 39 upvotes, #4 of 2025-11-14
  51. The Collaboration Gap 21 upvotes, #6 of 2025-11-05
  52. Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets 9 upvotes, #14 of 2025-10-31
  53. Code Aesthetics with Agentic Reward Feedback 7 upvotes, #21 of 2025-10-28
  54. LoongRL:Reinforcement Learning for Advanced Reasoning over Long Contexts 58 upvotes, #3 of 2025-10-23
  55. BitNet Distillation 47 upvotes, #8 of 2025-10-17
  56. Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning 1 upvotes, #36 of 2025-10-16
  57. DocReward: A Document Reward Model for Structuring and Stylizing 26 upvotes, #12 of 2025-10-14
  58. Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective 7 upvotes, #33 of 2025-10-01
  59. PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images 4 upvotes, #56 of 2025-09-30
  60. VibeVoice Technical Report 118 upvotes, #1 of 2025-08-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.