ServiceNow-AI
ServiceNow-AI on Hugging Face Daily Papers: 16 papers, 2 in the top 3 of their day, 1 paper of the day.
- AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling 20 upvotes, #12 of 2026-09-02
- StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments 40 upvotes, #7 of 2026-08-31
- SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding 70 upvotes, #4 of 2026-07-15
- EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents 64 upvotes, #5 of 2026-05-14
- Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics 60 upvotes, #6 of 2026-05-13
- Therefore I am. I Think 30 upvotes, #9 of 2026-04-03
- Terminal Agents Suffice for Enterprise Automation 92 upvotes, #2 of 2026-04-02
- EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings 142 upvotes, #5 of 2026-03-17
- Privileged Information Distillation for Language Models 25 upvotes, #11 of 2026-02-06
- Apriel-1.5-15b-Thinker 106 upvotes, #1 of 2025-10-06
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval 2 upvotes, #45 of 2025-10-03
- DeepCodeSeek: Real-Time API Retrieval for Context-Aware Code Generation 3 upvotes, #40 of 2025-10-01
- How to Train Your LLM Web Agent: A Statistical Diagnosis 46 upvotes, #4 of 2025-07-09
- Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows 5 upvotes, #28 of 2025-06-02
- Multi-task retriever fine-tuning for domain-specific and efficient RAG 10 upvotes, #12 of 2025-01-09
- Generating a Low-code Complete Workflow via Task Decomposition and RAG 4 upvotes, #15 of 2024-12-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.