Asaf Yehudai

Asaf Yehudai on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 308 upvotes.

  1. Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents 38 upvotes, #30 of 2026-09-30
  2. Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting 4 upvotes, #35 of 2026-06-09
  3. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks 65 upvotes, #4 of 2026-06-02
  4. Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents 7 upvotes, #32 of 2026-05-27
  5. General Agent Evaluation 11 upvotes, #12 of 2026-02-27
  6. CLEAR: Error Analysis via LLM-as-a-Judge Made Easy 17 upvotes, #6 of 2025-07-28
  7. Survey on Evaluation of LLM-based Agents 78 upvotes, #2 of 2025-03-21
  8. WildIFEval: Instruction Following in the Wild 11 upvotes, #10 of 2025-03-13
  9. Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language Models 9 upvotes, #12 of 2025-02-17
  10. JuStRank: Benchmarking LLM Judges for System Ranking 18 upvotes, #9 of 2024-12-13
  11. Benchmark Agreement Testing Done Right: A Guide for LLM Benchmark Evaluation 3 upvotes, #14 of 2024-07-19
  12. Genie: Achieving Human Parity in Content-Grounded Datasets Generation 8 upvotes, #14 of 2024-01-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.