Kyle Richardson
Kyle Richardson on Hugging Face Daily Papers: 7 papers, 2 in the top 3 of their day, 204 upvotes.
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite 3 upvotes, #20 of 2025-10-27
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning 14 upvotes, #13 of 2025-02-04
- SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories 6 upvotes, #12 of 2024-09-12
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research 66 upvotes, #2 of 2024-02-02
- OLMo: Accelerating the Science of Language Models 86 upvotes, #1 of 2024-02-02
- Paloma: A Benchmark for Evaluating Language Model Fit 12 upvotes, #8 of 2023-12-19
- Catwalk: A Unified Language Model Evaluation Framework for Many Datasets 7 upvotes, #13 of 2023-12-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.