Scale AI
Scale AI on Hugging Face Daily Papers: 12 papers, 0 in the top 3 of their day, 0 paper of the day.
- SteerDuplex: Steerable Duplex Speech Dialogue Models 6 upvotes, #25 of 2026-09-21
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures 5 upvotes, #28 of 2026-09-15
- Studying Without a Syllabus: Task-Agnostic Environment Preprocessing 4 upvotes, #16 of 2026-09-14
- HarnessOpt-Bench: Evaluating LLMs at Harness Optimization 34 upvotes, #9 of 2026-08-07
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures 10 upvotes, #30 of 2026-08-04
- SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions 6 upvotes, #31 of 2026-07-01
- Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR 6 upvotes, #30 of 2026-05-20
- HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help? 5 upvotes, #15 of 2026-05-05
- SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? 4 upvotes, #32 of 2026-04-14
- Agentic Rubrics as Contextual Verifiers for SWE Agents 10 upvotes, #9 of 2026-01-08
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents 7 upvotes, #11 of 2025-11-14
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training 16 upvotes, #21 of 2025-09-29
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.