ServiceNow-AI

ServiceNow-AI on Hugging Face Daily Papers: 16 papers, 2 in the top 3 of their day, 1 paper of the day.

  1. AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling 20 upvotes, #12 of 2026-09-02
  2. StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments 40 upvotes, #7 of 2026-08-31
  3. SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding 70 upvotes, #4 of 2026-07-15
  4. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents 64 upvotes, #5 of 2026-05-14
  5. Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics 60 upvotes, #6 of 2026-05-13
  6. Therefore I am. I Think 30 upvotes, #9 of 2026-04-03
  7. Terminal Agents Suffice for Enterprise Automation 92 upvotes, #2 of 2026-04-02
  8. EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings 142 upvotes, #5 of 2026-03-17
  9. Privileged Information Distillation for Language Models 25 upvotes, #11 of 2026-02-06
  10. Apriel-1.5-15b-Thinker 106 upvotes, #1 of 2025-10-06
  11. Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval 2 upvotes, #45 of 2025-10-03
  12. DeepCodeSeek: Real-Time API Retrieval for Context-Aware Code Generation 3 upvotes, #40 of 2025-10-01
  13. How to Train Your LLM Web Agent: A Statistical Diagnosis 46 upvotes, #4 of 2025-07-09
  14. Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows 5 upvotes, #28 of 2025-06-02
  15. Multi-task retriever fine-tuning for domain-specific and efficient RAG 10 upvotes, #12 of 2025-01-09
  16. Generating a Low-code Complete Workflow via Task Decomposition and RAG 4 upvotes, #15 of 2024-12-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.