Asaf Yehudai
Asaf Yehudai on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 308 upvotes.
- Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents 38 upvotes, #30 of 2026-09-30
- Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting 4 upvotes, #35 of 2026-06-09
- A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks 65 upvotes, #4 of 2026-06-02
- Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents 7 upvotes, #32 of 2026-05-27
- General Agent Evaluation 11 upvotes, #12 of 2026-02-27
- CLEAR: Error Analysis via LLM-as-a-Judge Made Easy 17 upvotes, #6 of 2025-07-28
- Survey on Evaluation of LLM-based Agents 78 upvotes, #2 of 2025-03-21
- WildIFEval: Instruction Following in the Wild 11 upvotes, #10 of 2025-03-13
- Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language Models 9 upvotes, #12 of 2025-02-17
- JuStRank: Benchmarking LLM Judges for System Ranking 18 upvotes, #9 of 2024-12-13
- Benchmark Agreement Testing Done Right: A Guide for LLM Benchmark Evaluation 3 upvotes, #14 of 2024-07-19
- Genie: Achieving Human Parity in Content-Grounded Datasets Generation 8 upvotes, #14 of 2024-01-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.