{"id":"2960883a-5429-454e-8f91-d2ae5d489695","arxiv_id":"2501.06143","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4o averaged 71% on English physics concept inventories, outperformed average post-instruction undergraduates in most subjects but not laboratory skills, and scored far worse on image-dependent items and in non-Western languages.","lead":"Researchers tested GPT-4o on thousands of physics quiz questions in 35 languages, submitting them as images just as students would see them. The AI beat the average undergraduate on most topics, but failed on lab-skills questions and on questions that required reading diagrams and graphs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Outperformance claim depends on non-representative, mostly English student benchmarks; language mismatch and sparse %Post values could flip category conclusions.","rationale":"The reader's weakest_assumption and our load-bearing concern coincide: the %Post benchmarks are best-effort, mostly English, and not systematically collected. The outperformance claim is the central headline and it is exactly the place where this weakness bites. I considered whether the image-requirement coding (Section IV.D) was more load-bearing; while inter-rater reliability is absent, the 81%/79%/49% pattern is an internal comparison that does not depend on student data and is corroborated by prior work (Polverini & Gregorcic 2024), so it is less fragile. The language-comparison findings (Section IV.B) are similarly descriptive and not undermined by benchmark issues. Thus the concrete test targets the one claim for which a flawed baseline is fatal. Because the authors explicitly acknowledge the limitation and frame the study as exploratory, the appropriate verdict remains CONDITIONAL with the reader's conditions; no change.","tokens_in":32656,"tokens_out":4959,"duration_ms":51295,"concrete_test":"Recompute the Section IV.C category comparison under two perturbations: (a) restrict to English-language inventories, comparing GPT-4o's English scores only to the English-speaking %Post values; (b) for each inventory with multiple %Post values (e.g., FCI 38/56/66, BEMA 42/43/61), use the maximum reported value instead of the mean and re-run the category averages. If GPT-4o no longer outperforms students in one or more subject categories under either perturbation, the central outperformance claim depends on benchmark selection. Where any non-English post-instruction data exist (e.g., FCI in Spanish or Swedish), perform a language-matched comparison; if the AI falls below those student scores, the claim should be restricted to English administrations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim that GPT-4o 'outperforms average post-instruction undergraduate students in all subject categories except laboratory skills' rests entirely on the %Post column in Tables II-V, which Section II describes as 'collected best-effort and not necessarily representative.' Section VII concedes that 'most of the human data came from English-speaking students taking the English versions.' The comparison is therefore not language-matched: for example, FCI %Post values (38, 56, 66; averaged ~53) are compared against GPT-4o's 20% in Punjabi and 22% in Tamil (Table VI), so the 'average undergraduate' is effectively an English-speaking student. Category-level claims are also fragile: Relativity has a single inventory (RCI) and Laboratory skills has only CDPA with a student baseline (MUQ has none); THERM and other categories rest on one or two %Post studies. No sensitivity analysis is given, so varying the benchmark within the reported ranges could plausibly flip several inventory-level comparisons and potentially category averages. The paper's own hedging about 'rough proxy' does not protect the strong categorical claim in the abstract and conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an empirical measurement study in which GPT-4o (Azure version 2024-08-06) was asked to solve items from 54 physics concept inventories sourced from PhysPort, using screenshots of the items rather than text-only inputs, across 35 languages. The authors report inventory-level and category-level performance, language differences, an item-level cross-lingual difficulty analysis, a comparison with published post-instruction undergraduate student scores, and an analysis of text-only versus image-required items. The headline findings are that GPT-4o performs worst on laboratory-skills items, performs better in English and European languages than in several non-European languages, shows largely language-independent item difficulty, outperforms average post-instruction students in most subject categories except laboratory skills, and performs markedly worse on items requiring visual interpretation (49% correct) than on text-only items (81% correct).","tokens_in":32820,"tokens_out":8285,"duration_ms":81350,"significance":"The study's empirical base is substantial and well documented: 3,662 image files, 4,674 items, 14,022 scored responses, and clear descriptions of the prompting protocol, JSON output schema, and data-processing decisions. The central descriptive result that required-image items are much harder for GPT-4o than text-only items (49% vs 81%) replicates and extends earlier work on vision-capable chatbots, and the cross-language difficulty correlation is an interesting and falsifiable observation. The promised data release on PhysPort is a further strength. However, the comparison with student benchmarks, which supports the strongest claim in the abstract and conclusion, is built on best-effort, mostly English, non-representative literature values and is not language-matched; this issue must be addressed before the categorical outperformance claim can be accepted as stated.","major_comments":[{"comment":"The claim that GPT-4o 'outperforms average post-instruction undergraduate students in all subject categories except laboratory skills' is not supported by a language-matched comparison. The %Post values in Tables II–V are, as the authors state in Section II, 'collected best-effort and not necessarily representative,' and Section VII concedes that most human data come from English-speaking students taking English versions. Figure 5, however, plots GPT-4o scores pooled over all languages against these benchmarks. For example, FCI %Post values of 38, 56, and 66 are compared with AI scores in 32 languages, several of which (20% in Punjabi, 22% in Tamil) are far below any of those student values. Because several categories rest on one or two inventories (RELA has only RCI; LAB has only CDPA with a student baseline, while MUQ has none), plausible variation in the benchmark values or restricting the comparison to English-language AI scores could change category-level conclusions. The authors should either redo the comparison with English-only or otherwise language-matched AI scores, or substantially soften the categorical wording in the abstract and conclusion.","section":"Section IV.C, Figure 5, and abstract"},{"comment":"The coding of items into 'text-only,' 'unneeded image,' and 'required image' is done manually, but no coding rubric, second coder, or inter-rater reliability check is reported. This coding is load-bearing for RQ4 and for the abstract's claim that the AI 'performs worse on items requiring visual interpretation of images,' since the 49% versus 81% gap is computed entirely from these labels. The qualitative direction of the result is likely robust given the size of the gap, but the per-category comparisons in Figure 6, several of which involve small numbers of inventories, need at least a transparent coding protocol or a reliability check to support the reported magnitudes.","section":"Section IV.D and Figure 6"},{"comment":"No uncertainty quantification is provided for the reported percentages, even though each item contributes three non-independent stochastic responses from a probabilistic system. Inventory-level and category-level differences, such as the LAB average of 35% versus the THERM average of 85%, or the FCI Portuguese score of 74% versus the Punjabi score of 20%, are discussed as exact values without confidence intervals or tests. A simple bootstrap by items or by inventories would establish which category and language differences are robust. The paper is explicitly exploratory, so this is not a fatal flaw, but the categorical claims in the abstract would be better supported by such an analysis.","section":"Sections IV.B–IV.C and Table VI"}],"minor_comments":[{"comment":"The text contains a typo: 'particularlyimportantnt' should read 'particularly important.'","section":"Section V, paragraph 2"},{"comment":"The version identifiers such as '2.0,' 'F06,' '5.5.7,' and 'vf' are not explained anywhere; readers cannot tell what distinguishes these versions or which exact version was used for each inventory.","section":"Tables II–V, column 'Vers.'"},{"comment":"The statement that 'all inventories — except TUG-K2.6 — were presented to the AI in English' is misleading: the screenshots contained text in the nominal language, and only the system prompt and JSON schema were in English. The wording should be corrected to 'prompted in English.'","section":"Section IV.B"},{"comment":"The random-incorrect-answer baseline is computed assuming five-option items with one correct answer, but the dataset includes inventories with four or other numbers of options. The theoretical probabilities are therefore only approximately applicable; the analysis should either restrict itself to five-option items or recompute the baselines by item type.","section":"Appendix C, Table X"},{"comment":"The figure reports percentages without the number of items or submissions in each image-category and subject-category cell; these sample sizes matter, especially for categories with few inventories such as LAB and RELA.","section":"Figure 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a physics education research journal and the underlying measurement effort is valuable. The main issue is that the abstract's strongest claim—that GPT-4o outperforms average post-instruction undergraduates in all subject categories except laboratory skills—depends on a non-representative, mostly English student benchmark and is not language-matched. This is fixable by reworking the comparison or the wording, so I recommend major revision rather than rejection. I see no concerns about citation practice or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The descriptive core is solid and worth knowing: 4,674 items across 54 inventories and 35 languages, administered as screenshots, with 14,022 scored responses. The two robust findings are the image-interpretation bottleneck (49% on required-image items vs. 81% on text-only) and the lab-skills weakness (35%). I also buy the cross-lingual item-difficulty invariance in Table IX — it is a real, non-obvious empirical pattern, and the conditional analysis (correctness in other languages given English correctness) is appropriate. The inconsistent-answer analysis is a nice extra: the model gravitates toward the same wrong choices, which is worth following up.\n\nThe soft spot is exactly what the stress test flagged. The headline claim that GPT-4o outperforms average post-instruction students in all categories except laboratory skills rests on the %Post columns, which the paper itself describes as best-effort and not necessarily representative. Most of the human data come from English-speaking students taking English versions, while the AI scores include languages like Punjabi and Tamil where the model is near random. Comparing those is apples to oranges. The paper does hedge in Sections II and VII, but the abstract and conclusion state the outperformance claim without that hedge. Category-level conclusions are also fragile where a category has one or two inventories, especially Relativity and Lab Skills. I do not think the claim is fabricated — the direction is probably right for English-language comparisons — but it is overstated as written. A sensitivity analysis varying the %Post benchmarks within their reported ranges, or restricting the comparison to a language-matched subset, would fix it.\n\nMinor soft spots are mostly self-acknowledged: three runs per item leaves sampling noise, the image-requirement coding has no reported inter-rater reliability, and the data availability is still a promise rather than a verified link. None of these undercut the descriptive findings.\n\nFor peer review: this deserves referee time. The descriptive map of multimodal LLM performance across physics inventories is useful to PER and to anyone benchmarking AI on physics tasks. I would send it to review but request a revision that softens or properly bounds the student-comparison claim. The paper is honest and the central image-related finding will hold up.","headline":"A useful descriptive benchmark with a real image-interpretation finding; the student-outperformance claim is overstated because the human baselines are best-effort and mostly English.","tokens_in":33363,"tokens_out":1407,"would_cite":true,"duration_ms":16759,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On screenshots of physics concept inventories, GPT-4o outscored average post-instruction students in every subject category except laboratory skills.","keywords":["physics education research","concept inventories","GPT-4o","multimodal language models","multilingual performance","visual interpretation","undergraduate physics assessment","educational equity"],"falsifier":"Collect a matched, representative sample of post-instruction undergraduates for the same inventories, in the same languages and screenshot formats, with the same scoring rules; if the AI's per-inventory average falls below the student average in more than one subject category, the paper's central outperformance claim fails.","tokens_in":32466,"feed_emoji":"🤖","tokens_out":7280,"duration_ms":72234,"temperature":0.7,"pith_summary":"This paper sets out to map where a current multimodal AI system stands on the kind of conceptual physics questions used to evaluate undergraduate instruction. It screenshots 54 validated concept inventories in up to 35 languages and finds that the AI's average score beats the average post-instruction undergraduate in every subject category except laboratory skills. The two sharpest limits are visual: performance falls from 81% on text-only items to 49% when interpreting a diagram or graph is required, and scores approach random guessing in several non-European languages. The authors argue instructors should therefore treat AI's inventory scores as a partial capability profile, not as evidence of conceptual mastery.","feed_headline":"GPT-4o beats average students on physics tests, except labs","feed_subtitle":"Scored 81 percent on text-only items but 49 percent when a diagram had to be read; language gaps ran even deeper.","key_machinery":"The carrying mechanism is the multimodal screenshot protocol: 3,662 item images, each submitted three times, for 14,022 solutions, to the model with a structured JSON prompt, then scored against transcribed answer keys. Each item is coded by whether it is text-only, contains an unneeded image, or requires image interpretation, and each inventory is compared to best-effort published post-instruction student scores. This design is what turns raw model outputs into the per-language, per-subject accuracy tables and the visual-interpretation contrast.","core_discovery":"The central claim is that GPT-4o, when given screenshots of items from 54 validated physics concept inventories in up to 35 languages, outperforms the average post-instruction undergraduate on most instruments and in every subject category except laboratory skills. In English, its average is 71.1%; thermodynamics (85.2%) and astronomy (80.4%) are strongest, while laboratory skills are weakest at 35.0%. On the multimodal axis, the model answers 81% of text-only items correctly, 79% of items with decorative images, and only 49% of items where reading the image is necessary. The authors also find language dependence, with English and most European languages performing best and near-random scores in Punjabi (20%) and Tamil (22%) on the most widely translated inventory, while items that are hard in English tend to be hard in other languages. The paper frames these as exploratory findings about capability boundaries, not evidence of conceptual understanding.","pith_inferences":["A testable extension the authors do not run: feed the same items as text with image content transcribed into captions; if accuracy on the 49% 'required image' items rises toward text-only levels, the bottleneck is visual encoding rather than the physics content itself.","If the English prompt biases the comparison, re-running with native-language prompts on a subset of inventories would isolate language-of-prompt effects; the paper notes structured outputs were unreliable for non-English prompts, so this would need a different output format.","The 66% repeat-incorrect pattern suggests the model has systematic attractors in its wrong answers, and linking those attractors to specific distractors would give instructors a map of where AI and student misconceptions coincide or diverge.","Because the student post-instruction benchmarks are mostly English-speaking, the outperformance claim is strongest for English-language assessments; how the gap looks in non-English classrooms remains open."],"forward_implications":["If a multimodal model scores 81% on text-only concept items, many current conceptual homework and quiz questions can be completed by AI, weakening the validity of unproctored versions of such assessments.","If required-image items drop to 49% accuracy, graph- and diagram-based questions are a more resilient format for assessing students without AI assistance.","If English and European languages score far above South Asian languages, AI tutoring tools will be least reliable for the students who may most need them, creating an equity risk.","If the AI's incorrect answers are consistent across languages rather than random, its error profile can be studied separately from student misconceptions instead of being treated as a noisy student proxy."],"supporting_citations":[{"why":"source of the validated concept inventories and their translations; the dataset is drawn from it.","marker":"[82]"},{"why":"documents the validation classifications that determined which inventories qualified for inclusion.","marker":"[83]"},{"why":"identifies the model under test, GPT-4o, and its release.","marker":"[32]"},{"why":"describes the API service used to submit screenshots with a data-use contract that keeps inventories out of training.","marker":"[154]"},{"why":"provides the Force Concept Inventory, whose many translations drive the language-difficulty analysis.","marker":"[33]"},{"why":"provides the Brief Electricity and Magnetism Assessment, an image-heavy inventory that contributes to the visual-interpretation result and the student-score comparison.","marker":"[35]"},{"why":"prior study of the model family on a kinematics graph inventory that frames the visual-interpretation weakness revisited here.","marker":"[36]"},{"why":"prior study of image-layout effects on a circuit inventory that the paper's item-level image coding extends.","marker":"[37]"}],"fun_headline_variants":["GPT-4o tops physics students, but flops on lab skills","AI beats physics undergrads, lags on diagrams and labs","GPT-4o excels at physics text, stumbles on images","Language gaps and diagram struggles mar GPT-4o's physics win","GPT-4o outdoes students, except in labs and image questions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the published post-instruction student scores gathered from the literature are representative, comparable benchmarks for undergraduates across languages and institutions; the paper itself notes they are best-effort, heterogeneous, and mostly from English-speaking samples, so the outperformance comparison stands only as firmly as those numbers do.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o tops physics students, but flops on lab skills","AI beats physics undergrads, lags on diagrams and labs","GPT-4o excels at physics text, stumbles on images","Language gaps and diagram struggles mar GPT-4o's physics win","GPT-4o outdoes students, except in labs and image questions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000145,"raw_usage":{"total_tokens":1202,"prompt_tokens":994,"completion_tokens":208,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":115}},"tokens_in":610,"tokens_out":208,"duration_ms":2895,"temperature":1.0,"reasoning_tokens":115,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:05:56.670934+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a matched, representative sample of post-instruction undergraduates for the same inventories, in the same languages and screenshot formats, with the same scoring rules; if the AI's per-inventory average falls below the student average in more than one subject category, the paper's central outperformance claim fails.","supporting_citations":[{"cited_title":"Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment","cited_arxiv_id":"2412.10019","evidence_quote":"prior study of image-layout effects on a circuit inventory that the paper's item-level image coding extends."}],"review_version":1}