Stanford University

Stanford University on Hugging Face Daily Papers: 27 papers, 1 in the top 3 of their day, 1 paper of the day.

  1. PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX 3 upvotes, #31 of 2026-08-19
  2. AISPA: User-Centric System Prompt Auditing for Large Language Model Applications 37 upvotes, #8 of 2026-08-03
  3. Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? 16 upvotes, #18 of 2026-07-01
  4. InSight: Self-Guided Skill Acquisition via Steerable VLAs 1 upvotes, #27 of 2026-06-24
  5. Decentralized Multi-Agent Systems with Shared Context 3 upvotes, #36 of 2026-06-10
  6. The Distillation Game: Adaptive Attacks & Efficient Defenses 1 upvotes, #39 of 2026-06-08
  7. CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning 4 upvotes, #27 of 2026-06-08
  8. Linear Scaling Video VLMs for Long Video Understanding 11 upvotes, #29 of 2026-06-01
  9. Convex Low-resource Accent-Robust Language Detection in Speech Recognition 6 upvotes, #50 of 2026-05-29
  10. Federation of Experts: Communication Efficient Distributed Inference for Large Language Models 1 upvotes, #51 of 2026-05-14
  11. Asymmetric Flow Models 21 upvotes, #12 of 2026-05-14
  12. Counting as a minimal probe of language model reliability 4 upvotes, #18 of 2026-05-05
  13. PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments 9 upvotes, #11 of 2026-05-05
  14. The Last Human-Written Paper: Agent-Native Research Artifacts 24 upvotes, #10 of 2026-05-01
  15. Recursive Multi-Agent Systems 239 upvotes, #1 of 2026-04-29
  16. AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization 4 upvotes, #19 of 2026-04-20
  17. Zero-shot World Models Are Developmentally Efficient Learners 7 upvotes, #27 of 2026-04-14
  18. TRACE: Capability-Targeted Agentic Training 13 upvotes, #20 of 2026-04-14
  19. Multi-User Large Language Model Agents 26 upvotes, #6 of 2026-04-13
  20. SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation 17 upvotes, #8 of 2026-02-24
  21. Large Language Model Reasoning Failures 11 upvotes, #19 of 2026-02-09
  22. Latent Adversarial Regularization for Offline Preference Optimization 13 upvotes, #19 of 2026-01-30
  23. Endless Terminals: Scaling RL Environments for Terminal Agents 16 upvotes, #8 of 2026-01-26
  24. Learning to Discover at Test Time 40 upvotes, #10 of 2026-01-23
  25. An Information Theoretic Perspective on Agentic System Design 7 upvotes, #22 of 2025-12-30
  26. QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models 10 upvotes, #11 of 2025-12-24
  27. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models 14 upvotes, #17 of 2025-06-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.