Microsoft

Microsoft on Hugging Face Daily Papers: 68 papers, 6 in the top 3 of their day, 3 paper of the day.

  1. ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization 64 upvotes, #15 of 2026-10-02
  2. Rubric Rewards from Item Response Theory 14 upvotes, #35 of 2026-10-01
  3. Follow the Entities: A Corpus Map for Agentic Search 98 upvotes, #13 of 2026-09-30
  4. The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks 159 upvotes, #1 of 2026-09-23
  5. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation 107 upvotes, #4 of 2026-09-18
  6. When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models 52 upvotes, #10 of 2026-09-18
  7. VibeVoice-ASR-Streaming Technical Report 17 upvotes, #16 of 2026-09-03
  8. Sliding-window beats linear attention 16 upvotes, #18 of 2026-08-31
  9. AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces 64 upvotes, #5 of 2026-08-26
  10. One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows 12 upvotes, #16 of 2026-08-25
  11. Agent Lightning v1.0: Towards Harnessed Agentic RL 31 upvotes, #7 of 2026-08-19
  12. OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching 25 upvotes, #13 of 2026-08-11
  13. AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? 20 upvotes, #19 of 2026-08-04
  14. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model 35 upvotes, #5 of 2026-07-29
  15. OpenForgeRL: Train Harness-native Agents in Any Environment 8 upvotes, #18 of 2026-07-24
  16. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing 72 upvotes, #5 of 2026-07-22
  17. SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference 11 upvotes, #24 of 2026-07-07
  18. ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes 54 upvotes, #5 of 2026-07-07
  19. ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog 61 upvotes, #4 of 2026-07-07
  20. HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents 6 upvotes, #32 of 2026-07-02
  21. Building to the Test: Coding Agents Deliver What You Check, Not What You Requested 8 upvotes, #22 of 2026-07-02
  22. A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets 1 upvotes, #36 of 2026-06-18
  23. FastContext: Training Efficient Repository Explorer for Coding Agents 91 upvotes, #5 of 2026-06-16
  24. WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces 101 upvotes, #3 of 2026-06-12
  25. AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents 4 upvotes, #35 of 2026-06-09
  26. OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents 19 upvotes, #17 of 2026-06-02
  27. Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models 107 upvotes, #3 of 2026-05-25
  28. Covering Human Action Space for Computer Use: Data Synthesis and Benchmark 14 upvotes, #25 of 2026-05-13
  29. Synthetic Computers at Scale for Long-Horizon Productivity Simulation 20 upvotes, #13 of 2026-05-01
  30. Accurate and scalable exchange-correlation with deep learning 4 upvotes, #22 of 2026-04-22
  31. Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks 3 upvotes, #28 of 2026-04-22
  32. AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation 3 upvotes, #26 of 2026-04-13
  33. Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization 8 upvotes, #28 of 2026-04-10
  34. AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding 6 upvotes, #31 of 2026-03-31
  35. Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models 13 upvotes, #20 of 2026-03-20
  36. Phi-4-reasoning-vision-15B Technical Report 19 upvotes, #7 of 2026-03-05
  37. Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization 35 upvotes, #6 of 2026-02-27
  38. Experiential Reinforcement Learning 66 upvotes, #1 of 2026-02-17
  39. CoPE-VideoLM: Codec Primitives For Efficient Video Language Models 29 upvotes, #6 of 2026-02-16
  40. Improving Data and Reward Design for Scientific Reasoning in Large Language Models 39 upvotes, #13 of 2026-02-10
  41. SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks 6 upvotes, #26 of 2026-02-09
  42. MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration 32 upvotes, #7 of 2026-02-09
  43. CUA-Skill: Develop Skills for Computer Using Agent 13 upvotes, #32 of 2026-02-03
  44. RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 17 upvotes, #25 of 2026-02-03
  45. PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards 23 upvotes, #21 of 2026-02-03
  46. Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling 20 upvotes, #10 of 2026-02-02
  47. VIBEVOICE-ASR Technical Report 19 upvotes, #11 of 2026-01-27
  48. WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks 15 upvotes, #13 of 2026-01-07
  49. Animate Any Character in Any World 10 upvotes, #16 of 2025-12-22
  50. DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems 25 upvotes, #9 of 2025-12-09
  51. Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions 1 upvotes, #41 of 2025-12-03
  52. Fara-7B: An Efficient Agentic Model for Computer Use 9 upvotes, #18 of 2025-11-26
  53. UFO^3: Weaving the Digital Agent Galaxy 17 upvotes, #12 of 2025-11-18
  54. GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using Agents 14 upvotes, #7 of 2025-11-07
  55. Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs 20 upvotes, #14 of 2025-10-29
  56. Deep Self-Evolving Reasoning 10 upvotes, #17 of 2025-10-21
  57. QueST: Incentivizing LLMs to Generate Difficult Problems 31 upvotes, #9 of 2025-10-21
  58. Information-Preserving Reformulation of Reasoning Traces for Antidistillation 1 upvotes, #38 of 2025-10-15
  59. SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs 11 upvotes, #17 of 2025-10-07
  60. Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness 6 upvotes, #31 of 2025-10-03
  61. ACON: Optimizing Context Compression for Long-horizon LLM Agents 28 upvotes, #8 of 2025-10-02
  62. InfoAgent: Advancing Autonomous Information-Seeking Agents 10 upvotes, #27 of 2025-10-01
  63. The role of synthetic data in Multilingual, Multi-cultural AI systems: Lessons from Indic Languages 3 upvotes, #33 of 2025-09-29
  64. CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization 1 upvotes, #42 of 2025-09-29
  65. Behind RoPE: How Does Causal Mask Encode Positional Information? 6 upvotes, #22 of 2025-09-26
  66. Contextual Integrity in LLMs via Reasoning and Reinforcement Learning 4 upvotes, #30 of 2025-06-06
  67. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone 222 upvotes, #1 of 2024-04-23
  68. WizardCoder: Empowering Code Large Language Models with Evol-Instruct 34 upvotes, #2 of 2023-06-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.