Pith. sign in

REVIEW 1 cited by

Towards a Benchmark for Scientific Understanding in Humans and Machines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.10327 v2 pith:AXG5DQ2Q submitted 2023-04-20 cs.AI cs.CLcs.HChep-phphysics.hist-ph

classification cs.AIcs.CLcs.HChep-phphysics.hist-ph
keywords understandingscientificdifferentbenchmarkabilityapproachesevaluationhumans
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scientific understanding is a fundamental goal of science, allowing us to explain the world. There is currently no good way to measure the scientific understanding of agents, whether these be humans or Artificial Intelligence systems. Without a clear benchmark, it is challenging to evaluate and compare different levels of and approaches to scientific understanding. In this Roadmap, we propose a framework to create a benchmark for scientific understanding, utilizing tools from philosophy of science. We adopt a behavioral notion according to which genuine understanding should be recognized as an ability to perform certain tasks. We extend this notion by considering a set of questions that can gauge different levels of scientific understanding, covering information retrieval, the capability to arrange information to produce an explanation, and the ability to infer how things would be different under different circumstances. The Scientific Understanding Benchmark (SUB), which is formed by a set of these tests, allows for the evaluation and comparison of different approaches. Benchmarking plays a crucial role in establishing trust, ensuring quality control, and providing a basis for performance evaluation. By aligning machine and human scientific understanding we can improve their utility, ultimately advancing scientific understanding and helping to discover new insights within machines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards a Large Physics Benchmark

    physics.data-an 2025-07 conditional novelty 4.0 of 10

    The paper outlines a multi-format, expert-scored living benchmark for evaluating physics understanding and creativity in LLMs, supported so far only by a small pilot.

Pith tools