Leshem Choshen

Leshem Choshen on Hugging Face Daily Papers: 13 papers, 0 in the top 3 of their day, 183 upvotes.

  1. Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting 4 upvotes, #35 of 2026-06-09
  2. General Agent Evaluation 11 upvotes, #12 of 2026-02-27
  3. Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures 9 upvotes, #26 of 2025-10-29
  4. Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty 6 upvotes, #17 of 2025-07-29
  5. TextArena 27 upvotes, #7 of 2025-04-16
  6. Pretraining Language Models for Diachronic Linguistic Change Discovery 4 upvotes, #16 of 2025-04-10
  7. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation 16 upvotes, #12 of 2024-12-06
  8. LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content 25 upvotes, #8 of 2024-10-15
  9. The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community 9 upvotes, #8 of 2024-08-16
  10. Data Contamination Report from the 2024 CONDA Shared Task 8 upvotes, #6 of 2024-08-01
  11. Benchmark Agreement Testing Done Right: A Guide for LLM Benchmark Evaluation 3 upvotes, #14 of 2024-07-19
  12. Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI 23 upvotes, #5 of 2024-01-26
  13. Genie: Achieving Human Parity in Content-Grounded Datasets Generation 8 upvotes, #14 of 2024-01-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.