Khalid Saifullah

Khalid Saifullah on Hugging Face Daily Papers: 3 papers, 1 in the top 3 of their day, 45 upvotes.

  1. LiveBench: A Challenging, Contamination-Free LLM Benchmark 12 upvotes, #9 of 2024-06-28
  2. Bring Your Own Data! Self-Supervised Evaluation for Large Language Models 16 upvotes, #4 of 2023-06-26
  3. On the Reliability of Watermarks for Large Language Models 6 upvotes, #2 of 2023-06-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.