{"id":"cc411308-a8d3-4721-8edd-3eb390967fec","arxiv_id":"2607.15418","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"On real construction drawings, the best AI model scores 71.7% versus 94.9% for experienced engineers, with the largest gaps in expert-level reasoning and quantity take-off.","lead":"DrawingVQA is a new benchmark that tests AI vision-language models on real construction drawings, using 92 expert-written questions at three reasoning depths. Early results show the best model beats average civil-engineering students but falls far below experienced professionals, especially on quantity take-off and multi-step cross-referencing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human/model comparison mixes denominators: §4.1 humans answer 20 items while models are scored on all 92, so the reported 71.7 vs 94.9 expert gap is not yet quantitatively established.","rationale":"The reader's weakest assumption correctly identifies the human/model denominator mismatch as the most load-bearing issue. My independent reading of §4.1, Table 2, and the abstract confirms that the central quantitative claim depends on comparing a 20-question human sample with 92-question model scores, with no representativeness check and no uncertainty quantification. This makes the specific gap numbers unreliable, though the benchmark construction and the qualitative direction of results remain plausible. Since the issue is fixable by a side-by-side comparison on the same subset and does not invalidate the dataset resource, CONDITIONAL remains the appropriate verdict. I also noticed the Supp. Table 7 API-endpoint swap, but that affects only which Gemini variant is best, not the core model-vs-expert comparison; it is secondary to the denominator problem.","tokens_in":27535,"tokens_out":4154,"duration_ms":43805,"concrete_test":"Run Gemini-2.5-pro (and at least the next two best models) on exactly the same 20 QA pairs used for the human study, using the same prompt and answer-parsing protocol, and report cohort-wise accuracy with bootstrap confidence intervals. If model performance on the 20-item subset is close to its 92-item performance and professionals still score far higher, the central gap claim stands; if the model matches or exceeds professionals on the 20-item subset, the headline quantitative conclusion fails as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim is that Gemini-2.5-pro (71.7%) surpasses average human performance (68.4%) but remains far below experienced professionals (94.9%), especially on R3 expert reasoning and QTO. However, §4.1 states that the human baseline used a randomly selected subset of 20 QA pairs (4 perceptual, 8 contextual, 8 expert) answered by 52 participants, while Table 2 and §4.2 report model scores over the full 92-question set. No analysis shows the 20-item subset is representative of the full benchmark in difficulty, reasoning-depth mix, or construction-domain coverage. With only 8 expert-level items, a single item shifts the professional R3 score by ~6 points, and the per-domain human rows in Table 2 (e.g., QTO, Compliance) rest on very few or even zero items depending on how the stratified subset landed. The reported model-vs-expert gap and the R3/QTO conclusions therefore lose their quantitative meaning as stated. This is not a cosmetic issue: the abstract's 'substantial gap between model and expert performance' is the central contribution. The benchmark itself may still be valuable, and the qualitative gap may survive a corrected comparison, but the current paper does not demonstrate it with the numbers it reports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"DrawingVQA is a benchmark of 92 expert-authored VQA pairs on 33 real-world 'Issued for Construction' structural drawings, spanning three reasoning depths and a dual categorization (seven construction-engineering dimensions x four MLLM capabilities). The paper evaluates proprietary and open-weight MLLMs and reports that Gemini-2.5-Pro (71.7%) surpasses the average human score (68.4%) but falls far short of experienced professionals (94.9%), with the largest gaps at expert-reasoning depth and quantity take-off. Ablations include text-only collapsing performance, option-order permutation, PDF-hybrid input, and an open-ended format check. The paper argues that current MLLMs have entry-level document literacy but not professional-grade drawing reasoning.","tokens_in":1537,"tokens_out":2603,"duration_ms":69212,"significance":"The dataset's authentic source, expert curation, contamination controls, and several falsifiable ablations are strengths. The qualitative finding that models degrade sharply at expert-reasoning depth and QTO is plausible and internally consistent. However, the headline human/model comparison mixes denominators (92-item model scores vs. 20-item human scores), so the quantitative claim is not yet established. With a corrected statistical footing, the benchmark would be a valuable community resource.","major_comments":[{"comment":"Model accuracy is computed on all 92 questions, while human scores come from a stratified 20-question subset (4/8/8) with a 20-minute time limit. The reported human rows in Table 2 thus rest on very small per-cell counts, and no analysis shows the subset is representative. The headline 71.7 vs. 94.9 comparison is therefore not quantitatively established. Please report model scores on the same 20 items, collect a full human baseline, or provide confidence intervals and a representativeness check.","section":"Sec. 4.1, Table 2, Sec. 4.2"},{"comment":"Gemini-2.5-Pro is listed with endpoint 'gemini-2.5-flash' and Gemini-2.5-Flash with 'gemini-2.5-pro'. If this is not a typo, the model attributions in Table 2 are unreliable. Please verify endpoints and correct the table.","section":"Supplementary F.2, Table 7"},{"comment":"No error bars or confidence intervals are reported for model or human accuracies. With n=92, Gemini-2.5-Pro (71.7) vs. average human (68.4) is within sampling noise; with n=20 for human subcohorts, per-domain rows are unstable. Report binomial/bootstrap CIs and exact item counts.","section":"Sec. 4.2, Table 2"}],"minor_comments":[{"comment":"Realism ratings differ: main text says 32.7% rated '4' and 38.5% rated '5'; Table 8 reports 32.1% and 39.6%. Please make consistent.","section":"Sec. 4.2 vs Supplementary Table 8"},{"comment":"'52 Responses' is ambiguous: 52 participants or 52 forms? Please report participant counts per cohort and per-question response counts.","section":"Sec. 4.1"},{"comment":"The MCQ-vs-open and unseen-drawings comparisons report only p-values (0.09, 0.07). Please include sample sizes, effect sizes, and test statistics.","section":"Supplementary D.6, D.7"},{"comment":"The reference list has duplicate numeric keys and some citations do not match their intended entries. Please clean up to allow verification of related-work claims.","section":"References"},{"comment":"The main-paper text-only ablation shows only two models. Since this is a strong sanity check, consider reporting the full table or explicitly directing readers to the supplementary table.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The dataset, ablations, and error analysis are solid and worth publishing; the main obstacle is the incommensurable human/model comparison, which is fixable by scoring models on the same 20 human-answered items or by collecting a full human baseline. The API endpoint transposition in Table 7 must be resolved before the numerical claims can be trusted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The dataset is the real thing: first expert-grounded VQA benchmark on actual IFC construction drawings, with a sensible three-depth reasoning schema and a dual categorization that maps engineering workflows to MLLM capabilities. The QA pairs are expert-authored with hand-crafted distractors, the ablations (text-only collapse, option permutation, PDF hybrid) are the right checks, and the error analysis with attention maps is a step beyond the usual benchmark paper. Any group working on document-grounded MLLM evaluation in AEC gets value from this resource.\n\nThe soft spot is the headline comparison. Human scores come from a 20-question stratified subset (4/8/8 split) answered by 52 participants, while model scores in Table 2 are over all 92 questions. There's no analysis showing the 20-item sample is representative, no error bars, and several per-domain human cells rest on a handful of items or zero. So the specific claim that Gemini-2.5-pro (71.7) beats average human (68.4) but trails professionals (94.9) is not quantitatively established as reported. The qualitative finding — models crater on R3 and QTO while experts stay strong — is internally consistent across many models and survives the ablations, so I'd bet it holds. But the abstract's 'substantial gap' needs to be re-anchored to data that actually supports it.\n\nTwo more issues, one minor and one worth fixing. Supp Table 7 swaps the Gemini-2.5-pro and Gemini-2.5-flash API endpoints, so the identity of the best scorer is uncertain; that's a typo but it lands on the central result. And the 'strictly contamination-free / true zero-shot' claim is asserted as confirmation for proprietary models whose training data you can't inspect; it should be worded as a mitigation, not a guarantee.\n\nNone of this sinks the benchmark, and I'd send it to peer review rather than desk reject. The authors need to redo the human comparison either by collecting human responses on the full set or by reporting model scores on the 20-item subset, add error bars, and qualify the contamination language. If they fix those, this becomes a standard reference for AEC-AI evaluation.","headline":"Solid benchmark resource, but the headline human-vs-model numbers are built on mismatched denominators and need a corrected rewrite.","tokens_in":28325,"tokens_out":2048,"would_cite":true,"duration_ms":20762,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DrawingVQA, a benchmark on real construction drawings, shows AI matches interns but not expert engineers.","keywords":["construction drawings","multimodal large language models","visual question answering","benchmark","expert reasoning","quantity take-off","reasoning depth","dual categorization"],"falsifier":"Recruit experienced professionals to answer all 92 DrawingVQA questions under the same timed protocol; if their score drops well below 94.9% or the best model's score rises above 71.7%, the reported model-versus-expert gap is not stable.","tokens_in":27344,"feed_emoji":"🏗️","tokens_out":4360,"duration_ms":42919,"temperature":0.7,"pith_summary":"DrawingVQA is a proposed benchmark for measuring how well multimodal large language models read real-world construction drawings — the dense, legally binding 'Issued for Construction' sheet sets used on building sites. Its central claim is that these documents test a kind of visual–textual reasoning that natural-image and schematic-plan benchmarks do not: combined reading of symbols, tables, cross-references, and annotations at professional depth. On 92 expert-written questions across three reasoning levels, the best model scores 71.7%, above the average human participant (68.4%) but far below seasoned professionals (94.9%). The authors take this as evidence that current AI has entry-level drawing literacy but not expert reasoning, and they propose a dual categorization framework — engineering tasks crossed with AI capabilities — to pinpoint exactly where models fail. The concrete payoff would be a diagnostic tool for deciding when models can be trusted in construction workflows.","feed_headline":"Top AI scores 71.7% on real drawings; pros hit 94.9%","feed_subtitle":"Benchmark on issued-for-construction sheets reveals the gap: expert reasoning and quantity take-off.","key_machinery":"The benchmark itself is the central object: 33 'Issued for Construction' structural drawings paired with 92 expert-written QA items, each assigned a reasoning depth (perceptual, contextual, expert). The second load-bearing mechanism is the dual categorization framework, which tags every question along an engineering dimension (administration, element identification, language understanding, dimensional reasoning, design semantics, quantity take-off, compliance) and an AI capability dimension (visual perception, knowledge, reasoning, OCR), so accuracy can be diagnosed by failure type rather than by a single score.","core_discovery":"On its own terms, the paper's discovery is that today's multimodal models read construction drawings competently at the perceptual and contextual levels, then collapse at the expert level: while almost every model's score falls from reasoning level 1 to level 3, experienced professionals score highest exactly at level 3. The largest model shortfall is quantity take-off — counting and extracting quantities from drawings — where the best model reaches 41.7% against 96.6% for professionals. The paper also shows that removing the image drops performance near to random guessing, which it takes as confirmation that the benchmark genuinely tests grounded visual reasoning rather than memorized domai","pith_inferences":["Because the model-vs-human comparison rests on only 20 of the 92 questions, the exact 71.7-vs-94.9 spread is a sample estimate; a full human pass could move the numbers, though the qualitative conclusion that experts outperform models on expert-level items would likely stand.","The standardization of US construction drawings (NCS) means the benchmark's difficulty may not transfer to non-US or non-standard drafting conventions; the same benchmark structure applied to those conventions would be a natural test.","The authors' PDF-hybrid result suggests resolution is not the bottleneck — even machine-readable vector text fails to ground spatially — which points toward training on technical drawings, not bigger images, as the extension path they leave implicit.","An extended inference: a model that cannot handle QTO cannot be trusted for cost estimation; so DrawingVQA-style probes could serve as a practical gate for AI adoption in construction, not just a research score."],"forward_implications":["Benchmark scores on DrawingVQA would let builders and engineers separate model competence in document retrieval from competence in spatial reasoning about a specific project's sheets.","If the paper is right, high pass rates on engineering certification exams tell us little about whether a model can trace a grid line or verify a weld callout — the benchmark re-anchors evaluation onto practical task performance.","The consistent quantity take-off failure isolates a single capability — precise object counting in dense line drawings — that future multimodal systems must improve before they can support estimating workflows.","The dual categorization framework would let benchmark users attribute an error to perception, OCR, knowledge, or reasoning per engineering task, enabling targeted rather than global model fixes."],"fun_headline_variants":["AI flunks expert-level construction drawing tasks","Quantity take-off: the task where AI scores 41.7% vs pros' 96.6%","Why AI can't read construction drawings like an expert","New benchmark: AI models collapse on construction expert reasoning","Real-world drawings benchmark exposes AI's expert gap"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparison that produces the paper's headline numbers assumes that the 20-question human sample fairly represents all 92 questions, and that the three experts' authored answers are correct ground truth for every item.","fun_headline_variants_meta":{"raw":{"variants":["AI flunks expert-level construction drawing tasks","Quantity take-off: the task where AI scores 41.7% vs pros' 96.6%","Why AI can't read construction drawings like an expert","New benchmark: AI models collapse on construction expert reasoning","Real-world drawings benchmark exposes AI's expert gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001221,"raw_usage":{"total_tokens":4844,"prompt_tokens":719,"completion_tokens":4125,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":4040}},"tokens_in":463,"tokens_out":4125,"duration_ms":25602,"temperature":1.0,"reasoning_tokens":4040,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T23:25:08.065834+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit experienced professionals to answer all 92 DrawingVQA questions under the same timed protocol; if their score drops well below 94.9% or the best model's score rises above 71.7%, the reported model-versus-expert gap is not stable.","supporting_citations":[],"review_version":1}