Pith. sign in

REVIEW 1 cited by

Phy-Q as a measure for physical reasoning intelligence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.13696 v3 pith:PVPHHNDB submitted 2021-08-31 cs.AI cs.LG

classification cs.AIcs.LG
keywords physicalagentsreasoninggeneralizationhumanphy-qscenariosagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans are well-versed in reasoning about the behaviors of physical objects and choosing actions accordingly to accomplish tasks, while it remains a major challenge for AI. To facilitate research addressing this problem, we propose a new testbed that requires an agent to reason about physical scenarios and take an action appropriately. Inspired by the physical knowledge acquired in infancy and the capabilities required for robots to operate in real-world environments, we identify 15 essential physical scenarios. We create a wide variety of distinct task templates, and we ensure all the task templates within the same scenario can be solved by using one specific strategic physical rule. By having such a design, we evaluate two distinct levels of generalization, namely the local generalization and the broad generalization. We conduct an extensive evaluation with human players, learning agents with varying input types and architectures, and heuristic agents with different strategies. Inspired by how human IQ is calculated, we define the physical reasoning quotient (Phy-Q score) that reflects the physical reasoning intelligence of an agent using the physical scenarios we considered. Our evaluation shows that 1) all agents are far below human performance, and 2) learning agents, even with good local generalization ability, struggle to learn the underlying physical reasoning rules and fail to generalize broadly. We encourage the development of intelligent agents that can reach the human level Phy-Q score. Website: https://github.com/phy-q/benchmark

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories

    physics.ed-ph 2025-01 conditional novelty 6.0 of 10

    GPT-4o averaged 71% on English physics concept inventories, outperformed average post-instruction undergraduates in most subjects but not laboratory skills, and scored far worse on image-dependent items and in non-Wes...

Pith tools