Pith. sign in

REVIEW 3 cited by

EEE-Bench: A Comprehensive Multimodal Electrical And Electronics Engineering Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.01492 v2 pith:YY7RDGAR submitted 2024-11-03 cs.CV

classification cs.CV
keywords engineeringlmmseee-benchmodelsproblemsbenchmarkmultimodalcapability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent studies on large language models (LLMs) and large multimodal models (LMMs) have demonstrated promising skills in various domains including science and mathematics. However, their capability in more challenging and real-world related scenarios like engineering has not been systematically studied. To bridge this gap, we propose EEE-Bench, a multimodal benchmark aimed at assessing LMMs' capabilities in solving practical engineering tasks, using electrical and electronics engineering (EEE) as the testbed. Our benchmark consists of 2860 carefully curated problems spanning 10 essential subdomains such as analog circuits, control systems, etc. Compared to benchmarks in other domains, engineering problems are intrinsically 1) more visually complex and versatile and 2) less deterministic in solutions. Successful solutions to these problems often demand more-than-usual rigorous integration of visual and textual information as models need to understand intricate images like abstract circuits and system diagrams while taking professional instructions, making them excellent candidates for LMM evaluations. Alongside EEE-Bench, we provide extensive quantitative evaluations and fine-grained analysis of 17 widely-used open and closed-sourced LLMs and LMMs. Our results demonstrate notable deficiencies of current foundation models in EEE, with an average performance ranging from 19.48% to 46.78%. Finally, we reveal and explore a critical shortcoming in LMMs which we term laziness: the tendency to take shortcuts by relying on the text while overlooking the visual context when reasoning for technical image problems. In summary, we believe EEE-Bench not only reveals some noteworthy limitations of LMMs but also provides a valuable resource for advancing research on their application in practical engineering tasks, driving future improvements in their capability to handle complex, real-world scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning

    cs.CL 2025-11 conditional novelty 6.0 of 10

    EngTrace, a 1,350-instance symbolic engineering benchmark with gold reasoning traces, shows frontier LLMs outperform math-specialized small models and that trace verification reveals a complexity cliff.

  2. SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new multimodal benchmark for strength of materials shows current foundation models solve at most 56.6% of problems, and expert-written diagram descriptions help more than images.

  3. Patient-Specific Autoregressive Models for Organ Motion Prediction in Radiotherapy

    eess.IV 2025-05 conditional novelty 5.0 of 10

    Auto-RMP, an autoregressive VQGAN plus causal transformer model, predicts future 4D CT phases from prior phases and reports higher lung and heart motion accuracy than DAM and DiffuseRT on public and private datasets.

Pith tools