Pith. sign in

REVIEW 2 cited by

ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.00102 v2 pith:FJADEZIX submitted 2024-11-27 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords digitalmllmselectronicsdatasetmulti-modalproblemsvisualanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal Large Language Models (MLLMs) are gaining significant attention for their ability to process multi-modal data, providing enhanced contextual understanding of complex problems. MLLMs have demonstrated exceptional capabilities in tasks such as Visual Question Answering (VQA); however, they often struggle with fundamental engineering problems, and there is a scarcity of specialized datasets for training on topics like digital electronics. To address this gap, we propose a benchmark dataset called ElectroVizQA specifically designed to evaluate MLLMs' performance on digital electronic circuit problems commonly found in undergraduate curricula. This dataset, the first of its kind tailored for the VQA task in digital electronics, comprises approximately 626 visual questions, offering a comprehensive overview of digital electronics topics. This paper rigorously assesses the extent to which MLLMs can understand and solve digital electronic circuit questions, providing insights into their capabilities and limitations within this specialized domain. By introducing this benchmark dataset, we aim to motivate further research and development in the application of MLLMs to engineering education, ultimately bridging the performance gap and enhancing the efficacy of these models in technical fields.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MatSciBench, a 1,340-question materials science benchmark with detailed solutions and images, shows top LLMs still fall short of college-level mastery.

  2. AITEE -- Agentic Tutor for Electrical Engineering

    cs.CY 2025-05 conditional novelty 6.0 of 10

    AITEE combines YOLO circuit detection, graph-neural-network-based retrieval of lecture material, SPICE simulation, and Socratic prompting to help LLMs answer first-semester electrical engineering circuit questions mor...

Pith tools