Kyle Richardson

Kyle Richardson on Hugging Face Daily Papers: 7 papers, 2 in the top 3 of their day, 204 upvotes.

  1. AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite 3 upvotes, #20 of 2025-10-27
  2. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning 14 upvotes, #13 of 2025-02-04
  3. SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories 6 upvotes, #12 of 2024-09-12
  4. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research 66 upvotes, #2 of 2024-02-02
  5. OLMo: Accelerating the Science of Language Models 86 upvotes, #1 of 2024-02-02
  6. Paloma: A Benchmark for Evaluating Language Model Fit 12 upvotes, #8 of 2023-12-19
  7. Catwalk: A Unified Language Model Evaluation Framework for Many Datasets 7 upvotes, #13 of 2023-12-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.