Pith. sign in

REVIEW 5 cited by

Are large language models superhuman chemists?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01475 v2 pith:7USDXZK3 submitted 2024-04-01 cs.LG cond-mat.mtrl-scics.AIphysics.chem-ph

classification cs.LGcond-mat.mtrl-scics.AIphysics.chem-ph
keywords llmsmodelschemicalchemistslanguagebestcapabilitiesevaluating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have gained widespread interest due to their ability to process human language and perform tasks on which they have not been explicitly trained. However, we possess only a limited systematic understanding of the chemical capabilities of LLMs, which would be required to improve models and mitigate potential harm. Here, we introduce "ChemBench," an automated framework for evaluating the chemical knowledge and reasoning abilities of state-of-the-art LLMs against the expertise of chemists. We curated more than 2,700 question-answer pairs, evaluated leading open- and closed-source LLMs, and found that the best models outperformed the best human chemists in our study on average. However, the models struggle with some basic tasks and provide overconfident predictions. These findings reveal LLMs' impressive chemical capabilities while emphasizing the need for further research to improve their safety and usefulness. They also suggest adapting chemistry education and show the value of benchmarking frameworks for evaluating LLMs in specific domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 19 citations worldwide. Full citation record

  1. Divergence Decoding: Training-Free Capability Fusion

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Divergence Decoding routes each token to either a domain specialist or a general reasoning LLM based on Jensen-Shannon divergence, outperforming either model alone on most tested scientific tasks.

  2. CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    CheMatAgent uses hierarchical Monte Carlo tree search with separate policy and execution models, plus trained reward models, to improve tool selection and parameter filling on a new chemistry benchmark, ChemToolBench.

  3. Scientific-Intention Driven Embodied Intelligent Solar Telescope: Conceptual Design

    astro-ph.IM 2026-07 reject novelty 4.0 of 10

    A three-layer AI-agent design for intention-driven autonomous solar telescopes is proposed; only the precision temperature-control prototype was tested.

  4. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

  5. Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation

    cs.AI 2025-08 reject novelty 2.0 of 10

    This survey claims to be the first systematic review of LLMs for organic synthesis, but its central 'evaluation' is never actually performed.

Pith tools