Microsoft
Microsoft on Hugging Face Daily Papers: 68 papers, 6 in the top 3 of their day, 3 paper of the day.
- ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization 64 upvotes, #15 of 2026-10-02
- Rubric Rewards from Item Response Theory 14 upvotes, #35 of 2026-10-01
- Follow the Entities: A Corpus Map for Agentic Search 98 upvotes, #13 of 2026-09-30
- The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks 159 upvotes, #1 of 2026-09-23
- When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation 107 upvotes, #4 of 2026-09-18
- When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models 52 upvotes, #10 of 2026-09-18
- VibeVoice-ASR-Streaming Technical Report 17 upvotes, #16 of 2026-09-03
- Sliding-window beats linear attention 16 upvotes, #18 of 2026-08-31
- AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces 64 upvotes, #5 of 2026-08-26
- One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows 12 upvotes, #16 of 2026-08-25
- Agent Lightning v1.0: Towards Harnessed Agentic RL 31 upvotes, #7 of 2026-08-19
- OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching 25 upvotes, #13 of 2026-08-11
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? 20 upvotes, #19 of 2026-08-04
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model 35 upvotes, #5 of 2026-07-29
- OpenForgeRL: Train Harness-native Agents in Any Environment 8 upvotes, #18 of 2026-07-24
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing 72 upvotes, #5 of 2026-07-22
- SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference 11 upvotes, #24 of 2026-07-07
- ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes 54 upvotes, #5 of 2026-07-07
- ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog 61 upvotes, #4 of 2026-07-07
- HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents 6 upvotes, #32 of 2026-07-02
- Building to the Test: Coding Agents Deliver What You Check, Not What You Requested 8 upvotes, #22 of 2026-07-02
- A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets 1 upvotes, #36 of 2026-06-18
- FastContext: Training Efficient Repository Explorer for Coding Agents 91 upvotes, #5 of 2026-06-16
- WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces 101 upvotes, #3 of 2026-06-12
- AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents 4 upvotes, #35 of 2026-06-09
- OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents 19 upvotes, #17 of 2026-06-02
- Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models 107 upvotes, #3 of 2026-05-25
- Covering Human Action Space for Computer Use: Data Synthesis and Benchmark 14 upvotes, #25 of 2026-05-13
- Synthetic Computers at Scale for Long-Horizon Productivity Simulation 20 upvotes, #13 of 2026-05-01
- Accurate and scalable exchange-correlation with deep learning 4 upvotes, #22 of 2026-04-22
- Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks 3 upvotes, #28 of 2026-04-22
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation 3 upvotes, #26 of 2026-04-13
- Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization 8 upvotes, #28 of 2026-04-10
- AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding 6 upvotes, #31 of 2026-03-31
- Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models 13 upvotes, #20 of 2026-03-20
- Phi-4-reasoning-vision-15B Technical Report 19 upvotes, #7 of 2026-03-05
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization 35 upvotes, #6 of 2026-02-27
- Experiential Reinforcement Learning 66 upvotes, #1 of 2026-02-17
- CoPE-VideoLM: Codec Primitives For Efficient Video Language Models 29 upvotes, #6 of 2026-02-16
- Improving Data and Reward Design for Scientific Reasoning in Large Language Models 39 upvotes, #13 of 2026-02-10
- SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks 6 upvotes, #26 of 2026-02-09
- MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration 32 upvotes, #7 of 2026-02-09
- CUA-Skill: Develop Skills for Computer Using Agent 13 upvotes, #32 of 2026-02-03
- RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 17 upvotes, #25 of 2026-02-03
- PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards 23 upvotes, #21 of 2026-02-03
- Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling 20 upvotes, #10 of 2026-02-02
- VIBEVOICE-ASR Technical Report 19 upvotes, #11 of 2026-01-27
- WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks 15 upvotes, #13 of 2026-01-07
- Animate Any Character in Any World 10 upvotes, #16 of 2025-12-22
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems 25 upvotes, #9 of 2025-12-09
- Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions 1 upvotes, #41 of 2025-12-03
- Fara-7B: An Efficient Agentic Model for Computer Use 9 upvotes, #18 of 2025-11-26
- UFO^3: Weaving the Digital Agent Galaxy 17 upvotes, #12 of 2025-11-18
- GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using Agents 14 upvotes, #7 of 2025-11-07
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs 20 upvotes, #14 of 2025-10-29
- Deep Self-Evolving Reasoning 10 upvotes, #17 of 2025-10-21
- QueST: Incentivizing LLMs to Generate Difficult Problems 31 upvotes, #9 of 2025-10-21
- Information-Preserving Reformulation of Reasoning Traces for Antidistillation 1 upvotes, #38 of 2025-10-15
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs 11 upvotes, #17 of 2025-10-07
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness 6 upvotes, #31 of 2025-10-03
- ACON: Optimizing Context Compression for Long-horizon LLM Agents 28 upvotes, #8 of 2025-10-02
- InfoAgent: Advancing Autonomous Information-Seeking Agents 10 upvotes, #27 of 2025-10-01
- The role of synthetic data in Multilingual, Multi-cultural AI systems: Lessons from Indic Languages 3 upvotes, #33 of 2025-09-29
- CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization 1 upvotes, #42 of 2025-09-29
- Behind RoPE: How Does Causal Mask Encode Positional Information? 6 upvotes, #22 of 2025-09-26
- Contextual Integrity in LLMs via Reasoning and Reinforcement Learning 4 upvotes, #30 of 2025-06-06
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone 222 upvotes, #1 of 2024-04-23
- WizardCoder: Empowering Code Large Language Models with Evol-Instruct 34 upvotes, #2 of 2023-06-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.