Leshem Choshen
Leshem Choshen on Hugging Face Daily Papers: 13 papers, 0 in the top 3 of their day, 183 upvotes.
- Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting 4 upvotes, #35 of 2026-06-09
- General Agent Evaluation 11 upvotes, #12 of 2026-02-27
- Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures 9 upvotes, #26 of 2025-10-29
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty 6 upvotes, #17 of 2025-07-29
- TextArena 27 upvotes, #7 of 2025-04-16
- Pretraining Language Models for Diachronic Linguistic Change Discovery 4 upvotes, #16 of 2025-04-10
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation 16 upvotes, #12 of 2024-12-06
- LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content 25 upvotes, #8 of 2024-10-15
- The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community 9 upvotes, #8 of 2024-08-16
- Data Contamination Report from the 2024 CONDA Shared Task 8 upvotes, #6 of 2024-08-01
- Benchmark Agreement Testing Done Right: A Guide for LLM Benchmark Evaluation 3 upvotes, #14 of 2024-07-19
- Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI 23 upvotes, #5 of 2024-01-26
- Genie: Achieving Human Parity in Content-Grounded Datasets Generation 8 upvotes, #14 of 2024-01-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.