DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
Xinrun Xu, Pi Bu, Ye Wang, Börje Karlsson, Ziming Wang, Tengtao Song, Qi Zhu, Jun Song, Zhiming Ding, Bo Zheng
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning: 65 upvotes on Hugging Face Daily Papers, #4 of 26 papers on 2025-08-08. Day-by-day upvote history.
Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance. Real-world tasks typically require complex interactions, advanced spatial reasoning, long-term planning, and continuous strategy refinement, usually necessitating understanding the physics rules of the target scenario. However, evaluating these capabilities in real-world scenarios is often prohibitively expensive. To bridge this gap, we introduce DeepPHY, a novel benchmark framework designed to systematically evaluate VLMs' understanding and reasoning about fundamental physical principles through a series of challenging simulated environments. DeepPHY integrates multiple physical reasoning environments of varying difficulty levels and incorporates fine-grained evaluation metrics. Our evaluation finds that even state-of-the-art VLMs struggle to translate descriptive physical knowledge into precise, predictive control.
Paper page on Hugging Face · arXiv
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.