WeihaoZeng

WeihaoZeng on Hugging Face Daily Papers: 6 papers, 1 in the top 3 of their day, 176 upvotes.

  1. LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth 24 upvotes, #18 of 2026-02-10
  2. The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution 44 upvotes, #6 of 2025-10-30
  3. Pitfalls of Rule- and Model-based Verifiers -- A Case Study on Mathematical Reasoning 6 upvotes, #32 of 2025-05-29
  4. SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild 27 upvotes, #4 of 2025-03-25
  5. B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners 38 upvotes, #2 of 2024-12-24
  6. CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery 14 upvotes, #12 of 2024-06-14

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.