Pith. sign in

REVIEW 3 cited by

Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.14572 v2 pith:L5JY554Y submitted 2024-09-22 cs.CL cond-mat.mtrl-scics.AIcs.LG

classification cs.CLcond-mat.mtrl-scics.AIcs.LG
keywords llmsmaterialsrobustnessperformancesciencereliabilityvariousadversarial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have the potential to revolutionize scientific research, yet their robustness and reliability in domain-specific applications remain insufficiently explored. In this study, we evaluate the performance and robustness of LLMs for materials science, focusing on domain-specific question answering and materials property prediction across diverse real-world and adversarial conditions. Three distinct datasets are used in this study: 1) a set of multiple-choice questions from undergraduate-level materials science courses, 2) a dataset including various steel compositions and yield strengths, and 3) a band gap dataset, containing textual descriptions of material crystal structures and band gap values. The performance of LLMs is assessed using various prompting strategies, including zero-shot chain-of-thought, expert prompting, and few-shot in-context learning. The robustness of these models is tested against various forms of 'noise', ranging from realistic disturbances to intentionally adversarial manipulations, to evaluate their resilience and reliability under real-world conditions. Additionally, the study showcases unique phenomena of LLMs during predictive tasks, such as mode collapse behavior when the proximity of prompt examples is altered and performance recovery from train/test mismatch. The findings aim to provide informed skepticism for the broad use of LLMs in materials science and to inspire advancements that enhance their robustness and reliability for practical applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. URSA: The Universal Research and Scientific Agent

    cs.AI 2025-06 unverdicted novelty 4.0 of 10

    URSA is a modular agent ecosystem that uses LLMs and scientific tools to accelerate research tasks of varying complexity.

  2. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0 of 10

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

  3. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 conditional novelty 3.0 of 10

    A cross-disciplinary review of 151 studies concludes LLMs accelerate research workflows while introducing recurring technical and ethical risks, including ten it flags as underexplored.

Pith tools