Xiang Gao

Xiang Gao on Hugging Face Daily Papers: 6 papers, 1 in the top 3 of their day, 224 upvotes.

  1. NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents 42 upvotes, #7 of 2025-12-16
  2. DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains 4 upvotes, #14 of 2025-11-17
  3. FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning 29 upvotes, #7 of 2025-09-19
  4. Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? 54 upvotes, #4 of 2025-09-05
  5. FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction 61 upvotes, #3 of 2025-08-21
  6. Use Property-Based Testing to Bridge LLM Code Generation and Validation 10 upvotes, #9 of 2025-06-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.