Ping Nie

Ping Nie on Hugging Face Daily Papers: 22 papers, 4 in the top 3 of their day, 1,241 upvotes.

  1. WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors 30 upvotes, #8 of 2026-05-12
  2. Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction 105 upvotes, #2 of 2026-05-08
  3. ClawBench: Can AI Agents Complete Everyday Online Tasks? 255 upvotes, #3 of 2026-04-10
  4. Watch Before You Answer: Learning from Visually Grounded Post-Training 35 upvotes, #11 of 2026-04-08
  5. ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks 29 upvotes, #11 of 2026-03-31
  6. OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis 91 upvotes, #3 of 2026-03-24
  7. VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction 12 upvotes, #12 of 2026-02-17
  8. Context Forcing: Consistent Autoregressive Video Generation with Long Context 35 upvotes, #6 of 2026-02-06
  9. VisCoder2: Building Multi-Language Visualization Coding Agents 20 upvotes, #14 of 2025-10-29
  10. BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions 27 upvotes, #11 of 2025-10-14
  11. A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports 18 upvotes, #17 of 2025-10-03
  12. EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing 15 upvotes, #11 of 2025-10-02
  13. VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
  14. VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use 64 upvotes, #5 of 2025-09-03
  15. BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent 36 upvotes, #9 of 2025-08-12
  16. Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
  17. VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
  18. StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
  19. VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation 14 upvotes, #15 of 2025-05-21
  20. ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
  21. VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 20 upvotes, #13 of 2025-03-14
  22. ACECODER: Acing Coder RL via Automated Test-Case Synthesis 23 upvotes, #3 of 2025-02-05

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.