zhongyuan peng

zhongyuan peng on Hugging Face Daily Papers: 9 papers, 1 in the top 3 of their day, 304 upvotes.

  1. ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders 6 upvotes, #12 of 2026-07-23
  2. CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization 38 upvotes, #6 of 2025-07-09
  3. FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models 27 upvotes, #7 of 2025-05-06
  4. IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs 22 upvotes, #8 of 2025-04-23
  5. Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? 26 upvotes, #7 of 2025-02-27
  6. CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models 23 upvotes, #8 of 2025-02-25
  7. SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 92 upvotes, #3 of 2025-02-21
  8. A Comparative Study on Reasoning Patterns of OpenAI's o1 Model 15 upvotes, #15 of 2024-10-18
  9. MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models 17 upvotes, #7 of 2024-10-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.