{"id":"b95d1e9d-0327-4530-a7b7-2a76bcfae5b6","arxiv_id":"2507.03013","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A 201-question multimodal STEM dataset with student records shows that the best model (58.5%) trails students (62.7%), with the largest gaps on essential-image and multi-concept questions.","lead":"The authors built a dataset of 201 exam questions with images, student scores, and model answers to compare how well AI handles visual STEM questions. It shows AI still trails students overall, and that adding essential images and multiple concepts can challenge AI without hurting human performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline interaction claim—models drop on crucial images and complex questions while students stay flat—is never statistically tested; the reported bootstrap CIs overlap, and Table 5 contains an impossible CI.","rationale":"The central contribution is not the overall 58.5% vs 62.7% mean—those are close—but the interaction pattern that yields actionable assessment design. For that pattern to license the conclusion, the difference between human and model accuracy must vary by feature. The paper does not report any test of that variation. The bootstrap CIs shown in Tables 5-6 are the only quantitative support cited ('We report statistical significance ... in Tables 5 and 6'), yet the crucial/supplemental and simple/complex intervals overlap within each group. An overlapping CI cannot support 'students are stable' nor can it support 'models drop'; non-overlap is a sufficient but not necessary condition for a difference. A logistic mixed-effects model with a group-by-feature interaction is the direct test and should be feasible because item-level model predictions and per-question student aggregates are released. The impossible MCQMA CI suggests the reported statistics need a correctness pass before they can be used to draw conclusions. I agree with the reader that the historical baseline creates comparability concerns, but the more fundamental issue is that the headline interaction is untested; fixing the baseline would not salvage a missing interaction test. These are addressable, so CONDITIONAL remains the right verdict; my read does not move it.","tokens_in":12754,"tokens_out":5299,"duration_ms":62754,"concrete_test":"Using the released dataset, fit a logistic mixed-effects model to item-level correctness with fixed effects for group (model/student), image purpose (crucial/supplemental), and their interaction, plus random intercepts for question and subject; report the interaction coefficient and confidence interval. Run the same model for problem complexity. If the interaction is not significant, the conclusion should be reframed as descriptive rather than causal. Also recompute Table 5's MCQMA bootstrap CI; if the interval still excludes the mean, re-check the bootstrap code.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's actionable conclusion is an interaction claim: crucial images and multiple concepts 'pose a greater challenge for models without increasing difficulty for students.' But the analysis never tests that interaction. The only inferential statistics presented are bootstrap 95% CIs in Tables 5 and 6, and for the key comparisons those intervals overlap: model crucial 0.561 [0.49, 0.63] vs supplemental 0.740 [0.60, 0.86]; student crucial 0.637 [0.60, 0.67] vs supplemental 0.588 [0.51, 0.67]. Overlap does not establish equality, and non-overlap is not required; what is needed is a formal group-by-feature interaction test, e.g., a logistic mixed-effects regression on item-level predictions with random intercepts for questions. Without it, the visual gap could be driven by correlated features (diagrams, complexity, subject) or by the acknowledged mismatches between human baselines (historical course records, original French, partial-credit grading) and model evaluation (English translations, exact match, no partial credit for MCQMA). The reliability of the reported statistics is also in question: Table 5 lists MCQMA model accuracy 0.457 with 95% CI [0.68, 0.84], which is impossible for a mean inside its own interval. This suggests an error in the exact table used to support 'statistical significance.' The dataset release is a real asset, and the descriptive pattern may survive re-analysis, but as reported the central claim is underdetermined.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a manually curated dataset of 201 university-level STEM exam questions with images, annotated for image type, image purpose (crucial vs. supplemental), question format, and problem complexity. The authors evaluate four multimodal LLM families under five prompting strategies, aggregate model responses by majority vote and by max, and compare the resulting accuracies with historical student performance from course records (546 respondents per question on average). Their main descriptive findings are that the best model (GPT-4o under majority vote) reaches 58.5% accuracy versus 62.7% for students; that student performance is relatively stable across image and question features while model performance drops on crucial images, diagrams, line plots, MCQMA questions, and complex problems; and that these patterns motivate design recommendations for assessments that are harder for AI without adding difficulty for students.","tokens_in":13002,"tokens_out":2108,"duration_ms":26213,"significance":"If the central claim holds, the paper provides a useful, field-relevant result: multimodal LLMs are measurably weaker than students on visually grounded STEM reasoning, and this weakness is predictable from question features. The dataset release is a genuine asset; manual annotation of image purpose and problem complexity, use of multiple model families and prompting strategies, and grounding in real student performance data are all strengths. The paper also makes a concrete, falsifiable design recommendation. However, the headline interaction claim—that crucial images and multiple concepts challenge models more than students—is not supported by the statistical evidence presented, and one reported confidence interval is internally impossible. The significance of the paper therefore depends on a re-analysis that the current manuscript does not provide.","major_comments":[{"comment":"The central claim that crucial images and multiple concepts 'pose a greater challenge for models without increasing difficulty for students' is an interaction claim, but the paper never tests the group-by-feature interaction. The only inferential evidence is the bootstrap 95% confidence intervals in Tables 5 and 6, and for the key comparisons those intervals overlap: model crucial accuracy is 0.561 [0.49, 0.63] versus supplemental 0.740 [0.60, 0.86], while student crucial accuracy is 0.637 [0.60, 0.67] versus supplemental 0.588 [0.51, 0.67]. Overlapping intervals do not establish an interaction, and the descriptive differences could be driven by correlated features such as subject, question length, or format. I request a formal interaction test, for example a logistic mixed-effects regression on item-level predictions with random intercepts for questions, with group (student vs. model) interacted with image purpose and problem complexity. Without such a test, the actionable conclusion in the abstract and conclusion is underdetermined.","section":"Section 4.3 and Conclusion, Tables 5 and 6"},{"comment":"Table 5 reports MCQMA model accuracy as 0.457 with a 95% confidence interval of [0.68, 0.84]. This interval does not contain the reported mean and is therefore impossible for a bootstrap CI computed from the same data. This suggests an indexing or transcription error in the very table used to support statements about statistical significance. The authors should correct the table and re-verify all other intervals, since a reader cannot currently assess which of the reported intervals are reliable. This issue directly affects the discussion of MCQMA performance in Section 4.3.","section":"Table 5"},{"comment":"The human baseline is historical student accuracy from original course records, while the model is evaluated on English translations of the question text with exact-match grading and no partial credit for MCQMA. The manuscript acknowledges these mismatches only in the Limitations section, but they are load-bearing for the central comparison. If the original exams were in French (Appendix A.1, fields 13–14) or if instructor-specific grading included partial credit, then the student-model accuracy gap on visual and MCQMA questions could be partly an artifact of translation and grading differences rather than a genuine human visual advantage. Please either re-analyze with a more comparable human baseline (e.g., scoring the model with partial credit for MCQMA, or evaluating students on the same translated items) or substantially temper the causal-sounding recommendations in the Conclusion.","section":"Section 3, Section 4, and Limitations"}],"minor_comments":[{"comment":"The selection of the five prompting strategies is based on performance on only 10 questions. This is a small selection set and risks overfitting the prompt choice to those items; the manuscript should acknowledge this explicitly.","section":"Appendix B.3"},{"comment":"The reference for Claude 3.7 Sonnet is incomplete ('Claude 3.7 Sonnet, 2025' without a publisher or technical report identifier), and the SciBench reference appears twice in essentially identical form (Wang et al., 2023 and 2024b). Please clean up the bibliography.","section":"References"},{"comment":"The Limitations section states 'we confirmed statistical significance,' but the only inferential statistics presented are bootstrap confidence intervals, and no formal significance tests appear in the main text or appendices. Please either report the tests or remove the phrase.","section":"Limitations"},{"comment":"The ablation study in Appendix D.2 is described as showing that removing supplemental images 'slightly improves model performance, though the effect is minimal,' but no numerical comparison or confidence interval is provided for that ablation. Adding the numbers would make the claim checkable.","section":"Section 4.2 / Figure 3"},{"comment":"The statement that 'all prompting strategies perform similarly' in Figure 9 is based on visual inspection; reporting the per-strategy accuracies and their variability would make the claim more precise.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The dataset and descriptive analyses are a solid contribution, and the paper is likely publishable after a re-analysis. The main concern is statistical: the headline interaction is asserted without an interaction test, and Table 5 contains an impossible confidence interval that undermines confidence in the reported numbers. I would encourage the editor to request a revision that adds a formal mixed-effects analysis and corrects the table, rather than rejecting, because the descriptive pattern and the released data are valuable to the community. I would also gently note that the authors should be careful with the phrase 'confirmed statistical significance,' which currently overstates what the paper reports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful thing: this paper releases a hand-annotated dataset of 201 multimodal STEM questions with real student accuracy per question, and it reports a clear descriptive pattern—students stay flat across image role and complexity, while GPT-4o drops on crucial images and multi-concept questions. That dataset is a real asset, and the ablation that removes supplemental images is a nice touch. The limitations section is honest about the grading mismatch on MCQMA.\n\nThe soft spot is that the headline claim is an interaction claim and it is never tested as one. The only inferential statistics are bootstrap CIs in Tables 5 and 6, and for the two key comparisons the intervals overlap (model crucial 0.561 [0.49, 0.63] vs supplemental 0.740 [0.60, 0.86]; student crucial 0.637 [0.60, 0.67] vs supplemental 0.588 [0.51, 0.67]). Overlap does not disprove the effect, but it does not support the confident conclusion either. What is needed is a formal group-by-feature interaction test, e.g., a logistic mixed-effects model on item-level predictions with random intercepts for questions. Without that, the pattern could be driven by correlated features like subject or question length.\n\nThere is also a concrete error: Table 5 reports MCQMA model accuracy of 0.457 with a 95% CI of [0.68, 0.84]. A mean cannot sit outside its own interval; that looks like a swapped row or a mislabeled bound. Since the paper says it \"confirmed statistical significance,\" this table is load-bearing, and it needs to be fixed and rechecked.\n\nThe human baseline comes from historical course records, often in French with instructor-specific grading, while models saw English translations and were graded by exact match. That is a real but addressable mismatch. The authors acknowledge part of it but do not quantify how translation or partial credit would shift the comparison.\n\nBottom line: the dataset and the descriptive pattern are worth engaging with. The inferential layer, as written, does not support the actionable recommendation. This deserves a serious referee—an editor should send it out—but the revision should add proper interaction tests, fix the impossible CI, and address the baseline mismatch head-on.","headline":"A useful dataset and a plausible descriptive pattern, but the main interaction claim is untested and Table 5 contains an impossible CI.","tokens_in":13561,"tokens_out":1925,"would_cite":true,"duration_ms":19600,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Best AI model scores 58.5% on visual STEM exam questions; students reach 62.7% and keep the edge when images matter.","keywords":["multimodal LLM evaluation","STEM assessment","human-AI comparison","question design","academic integrity","visual reasoning","prompting strategies","university exams"],"falsifier":"Re-run the benchmark on the same questions in their original language, score multiple-answer questions with partial credit, and have the same instructors' students answer an identical machine-gradable version; if the model's deficit on crucial-image and multi-concept questions disappears or reverses, the design guidance does not survive.","tokens_in":12520,"feed_emoji":"🎓","tokens_out":5667,"duration_ms":60934,"temperature":0.7,"pith_summary":"The paper introduces a dataset of 201 university-level STEM exam questions with images, each annotated for image type, whether the image is essential or merely supplemental, question format, and conceptual complexity, and each paired with historical student performance averaging 546 respondents per question. Against that baseline, the best model with majority voting answers 58.5% of questions correctly, below the student average of 62.7%. The paper's central finding is that students consistently outperform the model precisely where images carry essential information or where questions combine multiple concepts, while model performance also fluctuates more across subjects. The authors conclude that questions can be designed to challenge current AI systems without making them harder for students, by making images crucial and keeping questions concise.","feed_headline":"Best AI still trails students on visual STEM questions","feed_subtitle":"Scores drop when images are essential and concepts multiply, while student scores stay flat — a lever for AI-resistant exams.","key_machinery":"The load-bearing instrument is the hand-curated, feature-annotated question set: 201 real university STEM questions, each labelled for image type (diagram, line plot, algorithm, picture), image purpose (crucial versus supplemental), question type (single-answer, multiple-answer, and compound questions), and problem complexity (simple versus complex), with historical student accuracy as a human baseline. The comparison protocol uses a fixed set of five prompting strategies—direct zero-shot, chain-of-thought, image-first, text-first, and a two-stage description pipeline—aggregated by majority vote and by maximum score, with exact-match grading for all multiple-choice variants. The annotation allows the authors to attribute model failures to specific features, such as showing that removing supplemental images barely changes model accuracy, which is what turns a benchmark into guidance for question design.","core_discovery":"The paper claims that current multimodal large language models are systematically weaker than human students on visually grounded STEM reasoning, and that this weakness is predictable from two question features: whether the image is essential and how many concepts must be integrated. On its 201-question benchmark the strongest system answers 58.5% of questions with a majority-vote aggregation of five prompting strategies, against a 62.7% human average; across at least one prompting strategy the model retrieves the correct answer on 75.5% of questions, meaning the limitation is as much about reliably selecting the right answer as about not knowing it. Human performance is nearly flat across image type, image role, question format, and complexity, whereas model accuracy drops markedly on crucial images (56.1% versus 74.0% for supplemental images), on multiple-answer formats, and on complex problems. From this the authors draw a design principle: an assessment that requires a crucial image and multiple interacting concepts is harder for models while imposing no added difficulty on students, so such features can protect take-home exams without penalising learners.","pith_inferences":["The paper's historical student baseline may not be strictly comparable to model scoring: students answered original-language exams with instructor-specific grading while models saw English translations scored by exact match. We infer that re-evaluating on original-language, partial-credit multiple-answer questions could narrow or widen the reported gaps by subject.","The flat student performance across question features suggests the human advantage on visual questions is durable, but a stronger test would give students and models identical, machine-gradable questions with image purpose randomly varied; we infer such matched-pair experiments would sharpen the causal claim about crucial images.","Because models retrieve the right answer in at least one prompt on three-quarters of questions, we infer that the practical vulnerability of remote assessments is not lack of knowledge but answer selection; calibration or voting schemes that force consistency may be a better defensive test than harder questions.","The design recommendations are correlational, not causal; we infer that prospective testing, writing fresh exam questions that vary image necessity and concept count while controlling content, is the natural next step before institutions rely on these features for integrity."],"forward_implications":["Educators can make take-home exams more AI-resistant by writing questions whose image is essential to solving the problem and whose conditions draw on multiple concepts, without adding measured difficulty for students.","A model that answers text-heavy STEM questions well should not be taken as evidence of multimodal competence; on questions where images matter, the best model scores 56.1 percent versus 63.7 percent for students.","Because the model succeeds on 75.5 percent of questions with at least one prompting strategy but only 58.5 percent under majority voting, students who sample several answer strategies could still pass, so single-attempt assessment is a weaker safeguard.","Subject matter matters more for models than for students: the model is strongest in astronomy, computer science, and microfabrication and weakest in quantum physics, chemistry, and neuroscience, so AI-proofing advice should be subject-aware.","Supplemental images confer no meaningful advantage to the model, since removing them leaves performance essentially unchanged, so accessibility-friendly text descriptions do not by themselves make questions AI-answerable."],"supporting_citations":[{"why":"Shows LLMs can perform well on university assessments, the prior result this paper refines by adding a human baseline and visual questions.","marker":"Borges et al. (2024)"},{"why":"Provides a multimodal physics QA dataset with multi-image chain-of-thought prompting, a methodological source for the benchmark design.","marker":"Anand et al. (2024)"},{"why":"SceMQA supplies a scientific multimodal QA benchmark that motivates comparing model performance against student ability.","marker":"Liang et al. (2024)"},{"why":"Exams-V covers multilingual multimodal exam questions and informs the translation and image-handling pipeline.","marker":"Das et al. (2024)"},{"why":"M3Exam offers a multilingual multimodal multilevel benchmark that frames the need for feature-level analysis.","marker":"Zhang et al. (2023)"},{"why":"Chain-of-thought prompting is one of the five strategies the paper uses to elicit model answers.","marker":"Wei et al. (2023)"},{"why":"Describes the GPT-4 family used as the primary evaluated model and the baseline image description generator.","marker":"OpenAI (2023)"}],"fun_headline_variants":["AI lags students on visual, multi-concept STEM questions","Best AI still stumbles on image-dependent STEM problems","Visual STEM tasks expose AI's gap vs. human students","Images and complexity thwart AI in STEM assessments"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that historical student scores from real exams, with instructor-specific grading and the original language, measure the same difficulty as exact-match scoring of translated questions given to models; if translation or grading choices shift difficulty, the human advantage on visual questions could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["AI lags students on visual, multi-concept STEM questions","Best AI still stumbles on image-dependent STEM problems","Visual STEM tasks expose AI's gap vs. human students","Images and complexity thwart AI in STEM assessments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1215,"prompt_tokens":938,"completion_tokens":277,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":214}},"tokens_in":554,"tokens_out":277,"duration_ms":3593,"temperature":1.0,"reasoning_tokens":214,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:46:43.954440+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the benchmark on the same questions in their original language, score multiple-answer questions with partial credit, and have the same instructors' students answer an identical machine-gradable version; if the model's deficit on crucial-image and multi-concept questions disappears or reverses, the design guidance does not survive.","supporting_citations":[],"review_version":1}