{"id":"624f4b23-14fc-4538-ad38-462070e6cf7d","arxiv_id":"2501.17183","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The authors created an aerospace manufacturing QA benchmark with roughly 2,480 questions and found that top LLMs score below 51% accuracy.","lead":"This paper builds a multiple-choice question bank from aerospace manufacturing textbooks and uses it to test eleven large language models, finding that the best model answers only about half the questions correctly. It is a domain-specific evaluation benchmark that could help decide where LLMs can be trusted in aircraft manufacturing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing weak point is the unvalidated Gemini-generated answer key: Table I's accuracy figures may measure agreement with an LLM-produced key rather than aerospace expertise.","rationale":"The reader's weakest assumption is exactly where I would put the weight. The paper's reported numbers are internally coherent: the category sizes sum to 2,480 and the scoring arithmetic in Tables I–IV is consistent with the described +1/−3 rule. The real vulnerability is external validity. Since the benchmark itself, the textbook list, and the answer keys are not released, the only guarantee of ground truth is a roughly 4% expert-review sample whose agreement rate with the automated keys is not reported. The conclusion that LLM aerospace knowledge is lacking is plausible and may well be true, and the consistent deficit in the Aerospace Materials category across models is suggestive. But a benchmark whose keys are produced by the same class of system under evaluation cannot, without independent key validation, support a quantitative claim about absolute capability. This does not require changing the reader's conditional verdict; it reinforces it. I would not reject the paper, because the methodology is sensible and the central claim is probably correct, but acceptance should be contingent on releasing the question bank and demonstrating key accuracy. The minor inconsistencies, such as Qwen2.5-72B-Instruct having 2,464 rather than 2,480 total questions in Table I, are not load-bearing relative to the ground-truth issue.","tokens_in":10421,"tokens_out":5115,"duration_ms":52851,"concrete_test":"Release the 2,480-question final benchmark, the answer keys, and the source textbook list. Have two or more independent aerospace manufacturing experts re-verify a random sample of 300 final questions without seeing Gemini's keys, recording their agreement with each automated key. Then recompute Table I's overall accuracy, average score, and model ordering using only the subset of questions whose keys the experts confirm as correct. If expert agreement on the sample is high (e.g., above 95%) and the model scores on the confirmed-key subset move by only a few percentage points, the concern is settled; if agreement is materially lower or scores shift substantially, the paper's central quantitative claim is not yet supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim rests entirely on Table I's accuracy numbers, and every one of those numbers is scored against an answer key produced by Gemini Pro Vision (Section III.B.2). The paper's own validation covers only a sample of about 300 of ~7,500 generated questions (Section III.B.3); it does not state that all 2,480 questions used in Table I passed expert review, and it does not report per-question agreement between the automated key and the expert reviewers. If the un-reviewed keys contain systematic errors—for example, a distractor marked correct or a correct option omitted—then the reported 38.99–50.75% exact-match accuracies are biased and the model ordering in Table I can change. The -3 penalty for selecting an incorrect option makes this non-benign: a key that omits a true option turns a model's correct knowledge into a negative contribution to average score, so the average-score rankings in Tables I–IV are also at risk. The withheld source list (Section III.A) prevents an independent reader from checking individual keys. The introduction's risk case study with Eqs. (1)–(2) is presented without query logs or reproducible traces, so it cannot substitute for key validation. Thus the headline conclusion that LLM capabilities are in urgent need of improvement is plausible but currently rests on unverified ground truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an evaluation methodology for LLMs in aerospace manufacturing: it extracts text from aerospace textbooks and guidelines, uses the Gemini Pro Vision API to automatically generate multiple-choice questions with multiple correct answers, and evaluates eleven LLMs on 2,480 such questions using a custom scoring system with a strong penalty for incorrect option selections. The headline result is that the best model, Gemini-2.0-flash-exp, achieves 50.75% overall accuracy, leading the authors to conclude that current LLM capabilities in aerospace professional knowledge are in urgent need of improvement. The paper also reports category-specific results for aerospace assembly, aerospace materials, and structural steel, and interprets the results as evidence against relying on LLMs without human oversight in safety-critical manufacturing applications.","tokens_in":10746,"tokens_out":2948,"duration_ms":28069,"significance":"If the ground-truth answer keys are valid, the paper would fill a genuine gap by providing a domain-specific, professionally oriented evaluation benchmark for LLMs in aerospace manufacturing. The category-specific analysis and the discussion of the custom penalty-based scoring are useful contributions, and the paper explicitly identifies an important application risk. However, the central quantitative claims rest on answer keys produced by the same LLM family used for evaluation, with only a small expert-reviewed sample, a withheld source list, no human reference score, and no random-chance baseline. These issues mean the reported accuracy and ranking figures should be treated as provisional; a strengthened validation would make the contribution significantly more valuable.","major_comments":[{"comment":"The evaluation's ground truth is produced by Gemini Pro Vision (Section III.B.2), and expert review covers only approximately 300 of roughly 7,500 generated questions (Section III.B.3); the paper does not state that all 2,480 questions used in Table I passed expert review, nor does it report per-question agreement between the automated answer keys and the expert reviewers. If un-reviewed keys contain systematic errors, the overall accuracy and average-score figures in Tables I–IV measure agreement with an LLM-generated key rather than aerospace expertise.","section":"Section III.B.3 / Table I"},{"comment":"The detailed list of source materials is 'temporarily withheld' (Section III.A). Because the central claim rests entirely on the accuracy numbers, and because individual answer keys cannot be checked without knowing the sources, this withholding blocks independent verification and replication of the benchmark; the source list, or at least the question set with keys, must be released or made available to reviewers.","section":"Section III.A"},{"comment":"No human reference score or random-chance baseline is reported. With six options, multiple correct answers, and a −3 penalty per incorrect selection, even a model that selected all options would obtain a specific non-trivial expected score; without such a baseline, the reported accuracies and average scores cannot be interpreted as evidence of domain competence beyond chance, and the claim that models 'in urgent need of improvement' requires a comparison point.","section":"Section III.C / Tables I–IV"},{"comment":"Gemini-2.0-flash-thinking-exp-1219 is scored on only 717 of 2,480 questions (Table I) because of JSON parsing issues, yet its 12.82% accuracy is listed and ranked alongside models evaluated on nearly the full set; the reported figures for this model are not comparable to the others and should either be recomputed on a common subset of questions, reported with a clear confidence interval, or omitted from the ranking tables.","section":"Section III.D.5 / Tables I–IV"},{"comment":"The risk case study's numerical parameters (P_deviation=0.32, P_failure|deviation=0.18, C_critical=10^6, the adiabatic expansion coefficient alpha=0.78, and the 500 FH−1 threshold) are presented without derivations, source footnotes, or query logs, so they cannot be independently checked; this illustrative calculation does not substitute for validation of the question-answer keys.","section":"Section I.B / Eqs. (1)–(2)"}],"minor_comments":[{"comment":"The sampling criteria for the approximately 300 expert-reviewed questions are not described; state whether sampling was random and stratified by question category and difficulty, and report the outcome of the review (e.g., number of questions whose keys or options were corrected).","section":"Section III.B.3"},{"comment":"The definition of ‘Attempt Rate’ is given in words only; provide a formula and clarify whether the denominator for accuracy is ‘Attempted Questions’ or ‘Total Questions’ when these counts differ.","section":"Section III.C"},{"comment":"The text says DeepSeek-Chat-V3 achieved the highest accuracy in Aerospace Materials at 35.40%, but Table III reports glm-4-plus at 35.89% and DeepSeek-Chat-V3 at 35.40%; reconcile the text and table.","section":"Section IV.B / Table III"},{"comment":"The reference placeholder ‘[?]’ for SHAP explanation frameworks is unresolved; add the intended citation or remove the placeholder.","section":"Section II.B"},{"comment":"Figure 1 is referenced in the text but not included in the provided manuscript; add the figure or remove the reference.","section":"Section I.B"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important evaluation gap, and the category-specific breakdown is a useful step. The main concern is verifiability: the withheld source list, limited expert validation of answer keys, and absence of a human baseline make the central numbers difficult to audit. If the authors can release the question bank and sources (at least for review), provide per-question key-validation statistics, and add a chance-level or human baseline, the contribution would be substantially strengthened. I also recommend that the editors consider whether an artifact-release policy applies to benchmark papers of this type."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is the first aerospace-manufacturing QA benchmark and a clear demonstration that current LLMs perform poorly on it. That's useful for a niche safety-critical domain, and the authors deserve credit for building a domain-specific evaluation from authoritative sources, running eleven models, and reporting a custom scoring system that penalizes wrong selections rather than just rewarding right ones. The expert review of a sample, the difficulty control, and the category breakdowns are sensible design choices.\n\nThe load-bearing problem is exactly what the stress-test note flags: the answer key is generated by Gemini Pro Vision, and only about 300 of the ~7,500 generated questions were expert-validated. The 2,480 questions used in Table I are not explicitly all expert-checked, and the source list is withheld. So the accuracy figures likely measure agreement with an LLM-produced key more than aerospace expertise. The -3 penalty makes this worse: if a key omits a genuinely correct option, a model with correct knowledge gets penalized, biasing both the accuracy and the average-score rankings. Without the released benchmark or a human reference score, an independent reader cannot verify the central numbers. That is a serious flaw.\n\nThe introduction's risk case study also raises eyebrows. Equations (1) and (2) present numbers like P_deviation=0.32 and C_critical=10^6 with no derivation or citation, and the ΔP_cab formula looks like an illustrative estimate rather than a validated calculation. This section reads more like motivation than evidence, and it should be clearly labeled as illustrative or removed.\n\nThat said, the qualitative conclusion that LLMs underperform in aerospace manufacturing is plausible and consistent with findings in other specialized domains. The paper is honest about its limitations, such as the JSON-parsing failures that hurt Gemini-2.0-flash-thinking's results, and the authors do not oversell their work. The benchmark itself, if properly released and validated, could be a genuine asset to the community.\n\nI would send this to peer review, but major revisions are required: release the benchmark and the full answer keys, provide expert agreement rates, report a random-chance baseline and human reference scores, and either remove the risk model or back it with real data. With those changes, the paper would be a solid contribution; as it stands, the numbers should be read as indicative, not definitive.","headline":"Useful first benchmark for LLMs in aerospace manufacturing, but the answer-key validation is too thin to trust the headline numbers.","tokens_in":11202,"tokens_out":1551,"would_cite":false,"duration_ms":15704,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On a 2,480-question aerospace manufacturing exam built from textbooks, the best LLM scores 50.75% and most models fall below half.","keywords":["LLM evaluation","aerospace manufacturing","automated question generation","multiple-choice question answering","hallucination","domain-specific benchmark","safety-critical AI"],"falsifier":"Have aerospace engineers independently answer a random sample of the roughly 2,180 questions that were not expert-reviewed, without seeing Gemini's keys, and compare their answers. If agreement with the keys is substantially below what it would be if the keys were authoritative, the reported model accuracies do not measure aerospace knowledge.","tokens_in":10249,"feed_emoji":"✈️","tokens_out":4864,"duration_ms":40890,"temperature":0.7,"pith_summary":"This paper tries to measure whether large language models actually know aerospace manufacturing, and it concludes they mostly do not. The authors build a multiple-choice exam from aerospace manufacturing textbooks and guidelines, generate questions automatically with an LLM, and have eleven models sit the test. The top model, Gemini-2.0-flash-exp, reaches 50.75% overall accuracy, while most models score below half and every model drops in the aerospace materials category. If the numbers are right, LLM output in this safety-critical domain cannot be trusted without external verification.","feed_headline":"Top LLM gets only 50.75% on aerospace manufacturing quiz","feed_subtitle":"Eleven models, 2,480 questions: most score below half and all stumble on aerospace materials.","key_machinery":"The machinery is a pipeline that turns authoritative aerospace textbooks and guidelines into a benchmark: OCR to Markdown, text-snippet selection, automated multiple-choice question generation with six options and one or more correct answers, difficulty control through prompting, and a penalty-aware scoring scheme (+1 per correct option, -3 per incorrect option, 0 for unselected options). The average score under this asymmetric penalty is what differentiates models, rewarding caution as well as raw accuracy.","core_discovery":"The central discovery is that current LLMs' aerospace manufacturing expertise is far short of what deployment would require. On the constructed 2,480-question set, the best-performing model, Gemini-2.0-flash-exp, achieves 50.75% overall accuracy and Claude-3-5-sonnet-20241022 achieves 49.37%, while GPT-4o, DeepSeek-Chat-V3, and others fall below 46%. In the Aerospace Materials category the best model scores only 42.92%, and several models obtain negative average scores under the asymmetric penalty scheme. The paper also presents a concrete hallucination case where a leading model recommended nickel plating for titanium fasteners on Boeing 787 fuselage assemblies, contradicting HB 8752-2023, AS9100D, and NASA-STD-6012B. The authors read this as evidence that models lack version-aware standards knowledge, process integration, and the ability to quantify airworthiness risk.","pith_inferences":["Because only about 300 of the roughly 7,500 generated questions were expert-reviewed, the reported accuracies should be read as agreement with an LLM-generated answer key until the remaining keys are independently checked.","The withheld textbook list and non-public question bank make external replication currently impossible; releasing an expert-validated subset would let other labs verify the numbers.","A natural testable extension is to measure whether retrieval-augmented generation or fine-tuning on the source documents lifts accuracy above the roughly 50% ceiling reported here, which the paper's own future-work list points toward."],"forward_implications":["Aerospace manufacturers cannot yet rely on an off-the-shelf LLM for process design, material selection, or tool information retrieval without human review.","The Aerospace Materials category is a consistent failure point across all models, indicating a gap in training data depth for material science.","The asymmetric scoring system supplies a practical way to rank models by reliability, not just accuracy, in safety-critical domains.","The generated question bank, with 7,500 questions produced and 2,480 used in the final evaluation, can serve as a reusable benchmark for future aerospace LLMs."],"supporting_citations":[{"why":"China Aviation Industry Standards HB 8751/8752-2023 fastener surface treatment specifications, used to expose the nickel-plating hallucination in the case study.","marker":"[1]"},{"why":"AS9100D quality-management requirements, cited to check coating thickness tolerance control in the same case.","marker":"[2]"},{"why":"NASA-STD-6012B corrosion-protection standard, cited for hydrogen-concentration limits that the suggested process would violate.","marker":"[3]"},{"why":"MMLU provides the general multi-task evaluation baseline against which this domain-specific aerospace question set is positioned.","marker":"[4]"},{"why":"MedExpQA provides a medical-domain QA benchmark precedent that motivates adapting evaluation frameworks to specialized fields.","marker":"[11]"}],"fun_headline_variants":["Top LLM scores 50.75% on aerospace manufacturing quiz","Aerospace quiz: best LLM only halfway to expert level","GPT-4, Gemini, Claude fail aerospace expertise test","2,480 aerospace questions: LLMs average below half","Aerospace materials quiz exposes LLM hallucinations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation's ground truth is a set of answer keys that were mostly written by an LLM, with only about 300 of roughly 7,500 questions expert-reviewed, so the reported accuracy numbers depend on those keys being correct.","fun_headline_variants_meta":{"raw":{"variants":["Top LLM scores 50.75% on aerospace manufacturing quiz","Aerospace quiz: best LLM only halfway to expert level","GPT-4, Gemini, Claude fail aerospace expertise test","2,480 aerospace questions: LLMs average below half","Aerospace materials quiz exposes LLM hallucinations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000372,"raw_usage":{"total_tokens":1991,"prompt_tokens":952,"completion_tokens":1039,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":956}},"tokens_in":568,"tokens_out":1039,"duration_ms":9501,"temperature":1.0,"reasoning_tokens":956,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:30:20.445498+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have aerospace engineers independently answer a random sample of the roughly 2,180 questions that were not expert-reviewed, without seeing Gemini's keys, and compare their answers. If agreement with the keys is substantially below what it would be if the keys were authoritative, the reported model accuracies do not measure aerospace knowledge.","supporting_citations":[{"cited_title":"Alonso, M","cited_arxiv_id":null,"evidence_quote":"MedExpQA provides a medical-domain QA benchmark precedent that motivates adapting evaluation frameworks to specialized fields."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"China Aviation Industry Standards HB 8751/8752-2023 fastener surface treatment specifications, used to expose the nickel-plating hallucination in the case study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"AS9100D quality-management requirements, cited to check coating thickness tolerance control in the same case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"NASA-STD-6012B corrosion-protection standard, cited for hydrogen-concentration limits that the suggested process would violate."},{"cited_title":"Hendrycks, T","cited_arxiv_id":null,"evidence_quote":"MMLU provides the general multi-task evaluation baseline against which this domain-specific aerospace question set is positioned."}],"review_version":1}