Ping Nie
Ping Nie on Hugging Face Daily Papers: 22 papers, 4 in the top 3 of their day, 1,241 upvotes.
- WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors 30 upvotes, #8 of 2026-05-12
- Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction 105 upvotes, #2 of 2026-05-08
- ClawBench: Can AI Agents Complete Everyday Online Tasks? 255 upvotes, #3 of 2026-04-10
- Watch Before You Answer: Learning from Visually Grounded Post-Training 35 upvotes, #11 of 2026-04-08
- ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks 29 upvotes, #11 of 2026-03-31
- OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis 91 upvotes, #3 of 2026-03-24
- VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction 12 upvotes, #12 of 2026-02-17
- Context Forcing: Consistent Autoregressive Video Generation with Long Context 35 upvotes, #6 of 2026-02-06
- VisCoder2: Building Multi-Language Visualization Coding Agents 20 upvotes, #14 of 2025-10-29
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions 27 upvotes, #11 of 2025-10-14
- A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports 18 upvotes, #17 of 2025-10-03
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing 15 upvotes, #11 of 2025-10-02
- VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use 64 upvotes, #5 of 2025-09-03
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent 36 upvotes, #9 of 2025-08-12
- Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
- VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
- StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation 14 upvotes, #15 of 2025-05-21
- ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 20 upvotes, #13 of 2025-03-14
- ACECODER: Acing Coder RL via Automated Test-Case Synthesis 23 upvotes, #3 of 2025-02-05
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.