Pith. sign in

REVIEW 4 cited by

FEABench: Evaluating Language Models on Multiphysics Reasoning Ability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.06260 v1 pith:KWVM7BT4 submitted 2025-04-08 cs.AI cs.CLcs.NAmath.NA

classification cs.AIcs.CLcs.NAmath.NA
keywords problemsabilitylanguagellmsengineeringfeabenchreasoningsoftware
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Building precise simulations of the real world and invoking numerical solvers to answer quantitative problems is an essential requirement in engineering and science. We present FEABench, a benchmark to evaluate the ability of large language models (LLMs) and LLM agents to simulate and solve physics, mathematics and engineering problems using finite element analysis (FEA). We introduce a comprehensive evaluation scheme to investigate the ability of LLMs to solve these problems end-to-end by reasoning over natural language problem descriptions and operating COMSOL Multiphysics$^\circledR$, an FEA software, to compute the answers. We additionally design a language model agent equipped with the ability to interact with the software through its Application Programming Interface (API), examine its outputs and use tools to improve its solutions over multiple iterations. Our best performing strategy generates executable API calls 88% of the time. LLMs that can successfully interact with and operate FEA software to solve problems such as those in our benchmark would push the frontiers of automation in engineering. Acquiring this capability would augment LLMs' reasoning skills with the precision of numerical solvers and advance the development of autonomous systems that can tackle complex problems in the real world. The code is available at https://github.com/google/feabench

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards

    cs.LG 2026-07 conditional novelty 7.0 of 10

    RLVP post-trains one LLM across eight PDE families with hybrid validity-plus-continuous physics rewards, improving solver accuracy and enabling selective compositional transfer to held-out PDEs.

  2. VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    VFEAgent is an end-to-end multi-agent framework that automates FEA modeling and simulation from multimodal inputs, achieving high success rates in generating physically valid simulations across engineering scenarios.

  3. InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Large multimodal models do far worse when collision videos violate familiar physics, and their small gains come from text exemplars, not the videos.

  4. PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models

    cs.LG 2025-05 reject novelty 6.0 of 10

    A new benchmark of 380 principle-based physics problems shows that state-of-the-art LLMs struggle to apply symmetry, conservation, and dimensional-analysis shortcuts, achieving under 50 percent average accuracy with h...

Pith tools