ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems

Yiming Zhang; Yingfan Ma; Yanmei Gu; Zhengkai Yang; Yihong Zhuang; Feng Wang; Zenan Huang; Yuanyuan Wang; Chao Huang; Bowen Song; Cheng Lin; Junbo Zhao

arXiv:2507.04766·cs.LG·July 9, 2025

ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems

Yiming Zhang, Yingfan Ma, Yanmei Gu, Zhengkai Yang, Yihong Zhuang, Feng Wang, Zenan Huang, Yuanyuan Wang, Chao Huang, Bowen Song, Cheng Lin, Junbo Zhao

PDF

Open Access

TL;DR

ABench-Physics introduces a challenging benchmark for evaluating large language models' physical reasoning and generalization abilities through static and dynamic physics problems of graduate and Olympiad difficulty.

Contribution

The paper presents ABench-Physics, a new benchmark with static and dynamic problems to rigorously assess LLMs' physics reasoning and robustness, addressing limitations of previous benchmarks.

Findings

01

State-of-the-art LLMs perform poorly on the benchmark.

02

Models struggle with dynamic problem variants and generalization.

03

Benchmark reveals significant gaps in physical reasoning capabilities.

Abstract

Large Language Models (LLMs) have shown impressive performance in domains such as mathematics and programming, yet their capabilities in physics remain underexplored and poorly understood. Physics poses unique challenges that demand not only precise computation but also deep conceptual understanding and physical modeling skills. Existing benchmarks often fall short due to limited difficulty, multiple-choice formats, and static evaluation settings that fail to capture physical modeling ability. In this paper, we introduce ABench-Physics, a novel benchmark designed to rigorously evaluate LLMs' physical reasoning and generalization capabilities. ABench-Physics consists of two components: Phy_A, a static set of 400 graduate- or Olympiad-level problems; and Phy_B, a dynamic subset of 100 problems equipped with an automatic variation engine to test model robustness across changing conditions.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMachine Learning in Materials Science · Topic Modeling · Text Readability and Simplification