Evaluating LLMs on Real-World Forecasting Against Human Superforecasters
Janna Lu
Evaluating LLMs on Real-World Forecasting Against Human Superforecasters: 3 upvotes on Hugging Face Daily Papers, #26 of 27 papers on 2025-07-08. Day-by-day upvote history.
Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their ability to forecast future events remains understudied. A year ago, large language models struggle to come close to the accuracy of a human crowd. I evaluate state-of-the-art LLMs on 464 forecasting questions from Metaculus, comparing their performance against human superforecasters. Frontier models achieve Brier scores that ostensibly surpass the human crowd but still significantly underperform a group of superforecasters.
Paper page on Hugging Face · arXiv
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.