Jiarui Lu

Jiarui Lu on Hugging Face Daily Papers: 3 papers, 0 in the top 3 of their day, 80 upvotes.

  1. ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities 13 upvotes, #5 of 2024-08-12
  2. MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains 37 upvotes, #6 of 2024-07-30
  3. Can Large Language Models Understand Context? 24 upvotes, #4 of 2024-02-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.