Pith. sign in

REVIEW 1 cited by

MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03161 v2 pith:NZ24QCSV submitted 2024-09-05 cs.CL cond-mat.mtrl-sci

MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models

classification cs.CL cond-mat.mtrl-sci
keywords llmsmaterialbenchmultiple-choiceproblemstypeanswerfree-responsematerials
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

A college-level benchmark dataset for large language models (LLMs) in the materials science field, MaterialBENCH, is constructed. This dataset consists of problem-answer pairs, based on university textbooks. There are two types of problems: one is the free-response answer type, and the other is the multiple-choice type. Multiple-choice problems are constructed by adding three incorrect answers as choices to a correct answer, so that LLMs can choose one of the four as a response. Most of the problems for free-response answer and multiple-choice types overlap except for the format of the answers. We also conduct experiments using the MaterialBENCH on LLMs, including ChatGPT-3.5, ChatGPT-4, Bard (at the time of the experiments), and GPT-3.5 and GPT-4 with the OpenAI API. The differences and similarities in the performance of LLMs measured by the MaterialBENCH are analyzed and discussed. Performance differences between the free-response type and multiple-choice type in the same models and the influence of using system massages on multiple-choice problems are also studied. We anticipate that MaterialBENCH will encourage further developments of LLMs in reasoning abilities to solve more complicated problems and eventually contribute to materials research and discovery.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

    cs.AI 2026-07 conditional novelty 7.0

    Materials-science mechanisms are readable and steerable in a Gemma LLM through matched state changes, while absolute hidden-state graphs fail to uniquely encode physical polarity.