{"id":"3810379c-936c-4c5b-8bae-738628f76801","arxiv_id":"2506.19662","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Across four image-based physics concept inventories, tested AI models scored between 21% and 81.5% accuracy, and cheaper models often matched or beat far more expensive ones.","lead":"This paper benchmarked multimodal AI models on 102 image-based physics questions and found accuracy ranging from 21% to 81.5%, with price not reliably predicting quality. It gives educators a cost-aware, evidence-based way to choose AI tools for tutoring, feedback, and grading in physics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark may be inflated by memorization of public concept inventories; the central claim's visual-task validity is unestablished.","rationale":"The reader's weakest assumption was protocol representativeness (single minimal prompt, temperature 0.7, 10 runs). While prompt sensitivity is a legitimate measurement-stability concern, the paper explicitly frames the minimal prompt as a deliberate reflection of default user scenarios and acknowledges it as a limitation. The more load-bearing issue is construct validity: the benchmark uses publicly available, established concept inventories that may have appeared in training data. If memorization inflates scores, the central claim's absolute performance levels and the advice about cheaper models being 'sufficiently capable' lose their basis. This is not an internal inconsistency but an unexamined threat to what the benchmark measures. The paper deserves credit for transparent methodology, a Zenodo dataset, repeated runs, and careful cost accounting; these do not mitigate the contamination risk. The observed low FTGOT scores are suggestive but do not rule out selective memorization of the older, more widely used instruments. A held-out perturbation test could settle the question. Since the concern warrants additional analysis but does not by itself falsify the comparative ranking, the conditional verdict stands; the paper should be revised to address memorization or to temper the claim that the benchmark measures visual task ability.","tokens_in":18272,"tokens_out":11540,"duration_ms":116156,"concrete_test":"Run the same 102 screenshots with a systematic visual perturbation—e.g., mirror-image flip, recoloring, or relabeling of axes/component values—on the top five models, keeping all text identical. If accuracy on high-scoring items drops substantially (e.g., >10 percentage points) under visual perturbation while text-only performance remains stable, the original scores depended on visual content; if accuracy is unchanged, the models are keyed to textual memorization, and the claim to measure visual tasks is unsupported. A complementary check would compare performance on 20–30 newly written isomorphic items with novel diagrams.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that MLLMs range from 81.5% to 21% and that cheaper models may be sufficiently capable for visual physics tasks—depends on the measured scores reflecting genuine visual reasoning. Section 2.2 selects four established, published concept inventories (TUG-K, BEMA, QMVI, FTGOT) that are freely accessible via PhysPort (footnote, p. 6). These tests are old and widely used, so they are plausibly present in the training corpora of the tested models; the paper never addresses this contamination risk. The observed score pattern is consistent with selective memorization: GPT-5 scores 93.2% and 92.3% on the oldest, most widely circulated tests (BEMA, TUG-K) but drops to 48.5% on the edited, less standardized FTGOT; the small-parameter Gemma models, with presumably less exposure, sit near chance. If models answer from memorized text-answer associations rather than the visual content of the screenshots, the benchmark does not measure 'physics visual tasks,' and the practical recommendation that institutions use cheaper models based on these absolute scores could be misleading. This is a construct-validity threat to the central claim, distinct from the acknowledged prompt-sensitivity limitation.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks 17 multimodal large language models (MLLMs) from Anthropic, Google, and OpenAI on four established, image-based physics concept inventories (TUG-K, BEMA, QMVI, FTGOT), comprising 102 items. Each item was submitted as a screenshot with a minimal prompt, 10 times per model, and responses were scored by the final selected letter. The authors report per-inventory and total accuracy, standard deviations, and standard errors, alongside token-based cost estimates for a single pass through the benchmark. The main empirical findings are: total accuracy ranges from 81.5% (GPT-5) to 21.0% (Gemma 3-4b); performance is substantially higher on BEMA and TUG-K than on QMVI and FTGOT; and cost does not strictly track performance, with GPT-5 mini being a notable low-cost, high-performance outlier. The paper concludes that institutions should select models on the basis of measured benchmark performance and cost rather than provider reputation or list price, and that some cheaper models may be adequate for certain educational tasks.","tokens_in":18436,"tokens_out":3867,"duration_ms":37863,"significance":"If the results are valid, the paper provides a useful, independent, and reusable benchmark for physics educators and administrators choosing among commercial MLLMs. The authors have made their response dataset publicly available on Zenodo (ref. [71]), which is a concrete strength and supports reproducibility. The paper also connects performance to per-use cost in a transparent way, addressing a gap in the physics education literature. However, the central claim that these scores reflect competence on 'physics visual tasks' rests on the assumption that the models actually use the image content rather than memorized text-answer associations from the published concept inventories. The statistical reporting also contains a technical error that undermines the precision of the reported confidence intervals. These issues are fixable but need to be addressed before the practical recommendations can be fully trusted.","major_comments":[{"comment":"The reported SEM is not the standard error of the averaged performance. The authors write that SD is the square root of the sum of item-score variances, and that SEM is SD divided by sqrt(10). For item i with 10 Bernoulli trials and success probability p_i, the item-score variance is p_i(1-p_i)/10, so the computed quantity is sqrt(Σ p_i(1-p_i)/10), while the SEM of the mean across k items is sqrt(Σ p_i(1-p_i)/10)/k. The reported values therefore overstate the uncertainty of the total score by a factor of roughly k (the number of items, e.g., 31 for BEMA). Because Table 3 uses these SEMs to indicate confidence in the performance estimates, the current numbers do not support statements like 'most SEM values below 2.5%' as a statement about the precision of the reported averages. The calculation should be corrected or the quantity relabeled as an aggregate variability measure.","section":"Section 2.4 and Table 3"},{"comment":"The paper does not address the risk that the four concept inventories—especially the widely circulated BEMA and TUG-K—are part of the models' training corpora. Since the text of the questions and answer options is publicly available via PhysPort, the high scores on BEMA (93.2% for GPT-5) and TUG-K (92.3%) could partly reflect memorization of answer-key patterns rather than visual interpretation of the screenshots. The drop to 48.5% on the edited FTGOT items, whose format was modified to remove the four-tier structure, is consistent with this concern. Because the central claim is that the benchmark measures performance on 'physics visual tasks,' the authors should either provide evidence that the visual content is necessary (e.g., a text-only ablation or a set of novel, non-public items) or substantially temper the claim that the scores reflect image-based reasoning. As written, the construct validity of the benchmark for its stated purpose is not established.","section":"Section 2.2 and Discussion"}],"minor_comments":[{"comment":"The abstract and research framing state that 15 models were benchmarked, but Table 1 lists 17 models and Table 3 includes an additional row for 'Gemini 2.5 Flash (no reasoning),' for 18 rows total. The number should be corrected and the status of the no-reasoning condition clarified as a configuration rather than a separate model.","section":"Abstract, Table 1, Table 3"},{"comment":"The prompt instructs models to answer with the letter N when no option is correct, but the original concept inventories do not contain an N option. This changes the response space and may interact with model behavior in ways not analyzed; the authors should at least report how often N was selected and whether any model used it disproportionately.","section":"Section 2.3"},{"comment":"There is a typo in the fourth paragraph of the Discussion: 'praticular' should be 'particular.'","section":"Section 4"},{"comment":"The figure includes a 'Student' marker with reference to multiple sources [67, 68, 74, 75], but it is not explained how student performance was aggregated across four different instruments with different student populations. The caption should state the source and construction of this reference point.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The core empirical contribution—an open, reproducible, cost-aware benchmark of current MLLMs on physics concept inventories—is valuable and within the journal's scope. The two major issues I raised are both addressable: the SEM calculation is a straightforward statistical correction, and the memorization risk can be mitigated by adding a text-only ablation or by reframing the claims to focus on the screenshots as given rather than on visual reasoning per se. I would not reject on these grounds, but the manuscript needs another round. I also note the abstract's '15 models' inconsistency as a sign that the final version should be carefully checked before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The authors have put together a genuinely useful dataset: 17 accessible MLLMs from three providers, four standard concept inventories, 102 items, 10 runs each, with token costs and a clean Zenodo archive of the raw responses. The main empirical finding—accuracy ranges from roughly 81% to 21%, and price is a weak predictor of quality—is well supported by the tables. GPT-5 mini at about $0.26 per pass looking like a much better buy than Claude Opus 4 at $4.50 is exactly the kind of information an instructor or an IT office needs.\n\nWhat is new here is the breadth. Previous work from this group and others mostly covered one or two models on one or two inventories. Combining multiple providers and four inventories with cost accounting is a real step forward, and the transparent methodology (screenshots, minimal prompt, 10 repeated runs, manual verification of a 30% subset) is a credit to the authors.\n\nThe biggest soft spot is one the paper does not address. BEMA, TUG-K, QMVI, and FTGOT are public, decades-old instruments that have been in circulation long enough to be in the training corpora of these models. The score pattern—best on the oldest, most widely distributed tests, worse on the newer four-tier FTGOT—is consistent with partial memorization. If models answer from text associations rather than from interpreting the screenshots, the benchmark does not measure physics visual tasks, and the practical recommendation to adopt cheaper models based on these absolute scores could mislead. The authors need to discuss this contamination risk and ideally test on a held-out set of novel items. This is not fatal to the paper's usefulness as a rough guide to model tiers, but it is fatal to the claim that these scores measure visual reasoning.\n\nThree smaller points. The abstract says 15 models, but Tables 1 and 3 list 17 (plus a \"no reasoning\" variant). The SEM calculation in Section 2.4 and the table notes describe something that is not the standard error of the averaged score; the wording or the formula needs fixing. And while raw responses are on Zenodo, the benchmark script and post-processing code are not; that is easy to add.\n\nThe paper is a practical contribution for physics educators and for anyone doing MLLM benchmarking in education. It deserves a serious referee, conditional on the authors addressing the contamination issue and cleaning up the minor inconsistencies. I would accept and ask for major revision rather than desk reject.","headline":"A useful, thoroughly documented benchmark of 17 MLLMs on standard physics concept inventories, but the unaddressed contamination risk of public test items keeps it from fully earning the claim that it measures visual reasoning.","tokens_in":18972,"tokens_out":4249,"would_cite":true,"duration_ms":43050,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multimodal AI models scored between 81.5 percent and 21 percent on visual physics questions, and price did not predict performance.","keywords":["multimodal large language models","physics education","concept inventories","visual reasoning","cost-performance benchmarking","geometrical optics","AI model selection"],"falsifier":"Re-run the same 102 items with several alternative minimal prompts, such as no answer-letter instruction, a 'think step by step' cue, or different temperatures, and check whether the top models and the cheap-model advantage persist. If GPT-5 mini's 75 percent collapses or Claude Opus 4 jumps above the mid-tier models under a plausible prompt, the cost-performance ranking is protocol-dependent rather than a stable property of the models.","tokens_in":18002,"feed_emoji":"⚛️","tokens_out":7645,"duration_ms":75007,"temperature":0.7,"pith_summary":"This paper asks whether commercially available multimodal AI models can handle the kinds of image-based conceptual physics questions that students encounter, and whether the cost of using them is justified. The authors benchmark a range of publicly available models from three major providers on 102 items drawn from four validated concept inventories covering kinematics, electromagnetism, quantum mechanics, and geometrical optics. Their central empirical claim is that performance varies enormously, from 81.5 percent down to 21 percent, and that price is not a reliable guide to quality: some inexpensive models perform close to the best, while one of the most expensive models underperforms mid-tier rivals. The paper argues that educational institutions should select models on measured performance and cost rather than provider reputation or listed prices. It also finds that all models struggle most with geometrical optics, suggesting that the type of visual reasoning required, rather than the physics topic by itself, is what drives difficulty.","feed_headline":"Cheap AI matched pricey rivals on physics visual test","feed_subtitle":"A 102-item benchmark shows accuracy from 81.5% to 21%, with cost failing to predict performance.","key_machinery":"The central mechanism is a standardized, image-based benchmark: screenshots of 102 items from four research-validated concept inventories, including BEMA for electricity and magnetism, TUG-K for kinematics graphs, QMVI for quantum mechanics visualization, and FTGOT for geometrical optics. Each item is submitted to every model through its API under one minimal prompt, with temperature set to 0.7 where possible, ten fresh runs per item, and scoring based only on the final answer letter. This protocol converts the vague question of whether a model understands physics pictures into a comparable performance number. The same runs supply the cost analysis, which tracks input, output, and hidden reasoning tokens and multiplies them by listed per-token prices to give a normalized single-pass cost for the entire benchmark.","core_discovery":"On the paper's own terms, the best multimodal models of mid-2025 already surpass post-instruction university student averages on established concept inventories: GPT-5 reaches 81.5 percent, with o3, Gemini 2.5 Pro, and GPT-5 mini in the 75 to 76 percent range. Every model's score drops sharply on items requiring spatial and geometric interpretation, most notably geometrical optics, where the best score is 51.5 percent. The cost analysis shows a loose, non-proportional relationship between price and accuracy: GPT-5 mini scores 75 percent at roughly $0.27 for one pass over the full benchmark, while Claude Opus 4 scores 57 percent at roughly $4.54 and Gemini 2.5 Pro scores 75.8 percent at roughly $4.68. The paper interprets this as evidence that cheaper models can be sufficiently capable for some educational uses, and that high per-token prices do not guarantee high performance on visual physics tasks.","pith_inferences":["I infer that the reported scores are lower bounds on what targeted prompting could achieve, particularly for non-reasoning models; if prompt engineering raises cheap models more than reasoning models, the cost-performance advantage could grow.","I infer that scoring only the final letter, without analyzing the reasoning text, could hide systematic differences in how models arrive at correct answers, so a qualitative pass might alter the practical ranking for tutoring.","I infer that the benchmark protocol could be tested for robustness by repeating the same 102 items with several minimal prompt variants and temperatures; if the rank order changes substantially, the conclusions are protocol-dependent.","I infer that the finding that visual format matters more than topic points toward building a physics-specific visual-reasoning benchmark that could predict model performance on unseen diagram types."],"forward_implications":["Institutions deploying AI for multiple-choice physics diagnostics can get near-top accuracy from a model costing less than a dollar for a full 102-item pass, rather than assuming the flagship model is required.","Price tier and provider reputation are not reliable proxies for visual physics ability: one of the most expensive models scored roughly 24 percentage points below a model costing about one-seventeenth as much.","Free open-weight models scored between 21 and 35 percent, below what the paper considers acceptable for student-facing physics work involving images.","Even the best models score below 52 percent on geometrical optics items, so AI support in that domain should be treated as unreliable for now.","Benchmarks like this need to be re-run regularly, because model capabilities, pricing, and API availability change quickly.","The finding that visual task type matters more than physics topic suggests that future work should analyze which specific visual formats cause failures.","A testable extension is to run the same 102 items under several prompt variants and temperatures to map sensitivity; large rank shifts would show that the current single-protocol numbers are protocol-dependent.","Because only the final answer letter is scored, a model could reach a high accuracy while generating flawed explanations; a qualitative analysis might rank models differently for tutoring uses where reasoning quality matters."],"supporting_citations":[{"why":"supplies the 31-item electricity and magnetism inventory that anchors the BEMA scores.","marker":"[66]"},{"why":"supplies the 20 geometrical optics items where all models scored lowest.","marker":"[67]"},{"why":"supplies the 25 quantum-mechanics visualization items.","marker":"[68]"},{"why":"supplies the 26 kinematics graph items.","marker":"[69]"},{"why":"earlier work showing that ChatGPT's failures came largely from misreading graphs, motivating image-based benchmarking.","marker":"[33]"},{"why":"earlier BEMA results showing models answer consistently or repeat the same wrong option, justifying ten runs per item.","marker":"[37]"},{"why":"a free-versus-subscription comparison that informs the choice of a minimal prompt and the task-type interpretation.","marker":"[40]"},{"why":"shows that feedback and grading quality tracks the model's own performance, making benchmark scores educationally relevant.","marker":"[49]"},{"why":"the released response dataset that the cost and performance calculations draw on.","marker":"[71]"}],"fun_headline_variants":["Cheap AI matches pricey rivals on physics visual test","Pricey AI models don't guarantee better physics answers","Physics visual benchmark: cost fails to predict accuracy","AI physics scores range 21–81.5%, price no guarantee"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire ranking rests on the assumption that a single minimal prompt, repeated ten times at one temperature, gives a stable and representative estimate of each model's ability on these image-based items; if answers depend heavily on prompt wording or formatting, the reported performance gaps could be artifacts of the protocol.","fun_headline_variants_meta":{"raw":{"variants":["Cheap AI matches pricey rivals on physics visual test","Pricey AI models don't guarantee better physics answers","Physics visual benchmark: cost fails to predict accuracy","AI physics scores range 21–81.5%, price no guarantee"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1299,"prompt_tokens":975,"completion_tokens":324,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":257}},"tokens_in":591,"tokens_out":324,"duration_ms":4037,"temperature":1.0,"reasoning_tokens":257,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:29:07.404850+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same 102 items with several alternative minimal prompts, such as no answer-letter instruction, a 'think step by step' cue, or different temperatures, and check whether the top models and the cheap-model advantage persist. If GPT-5 mini's 75 percent collapses or Claude Opus 4 jumps above the mid-tier models under a plausible prompt, the cost-performance ranking is protocol-dependent rather than a stable property of the models.","supporting_citations":[{"cited_title":"American Journal of Physics70(3), 238–251 (2002) https://doi.org/10","cited_arxiv_id":null,"evidence_quote":"supplies the 25 quantum-mechanics visualization items."},{"cited_title":"Physical Review Special Topics - Physics Education Research2(1), 010105 (2006) https: //doi.org/10.1103/PhysRevSTPER.2.010105","cited_arxiv_id":null,"evidence_quote":"supplies the 31-item electricity and magnetism inventory that anchors the BEMA scores."},{"cited_title":"Research in Science & Technological Education35(2), 238–260 (2017) https://doi.org/10.1080/02635143.2017.1310094","cited_arxiv_id":null,"evidence_quote":"supplies the 20 geometrical optics items where all models scored lowest."},{"cited_title":"Dataset on Zenodo (2025)","cited_arxiv_id":null,"evidence_quote":"the released response dataset that the cost and performance calculations draw on."}],"review_version":2}