{"id":"707a7479-b60a-46c6-aaa2-6f867983a14d","arxiv_id":"2508.21452","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"On a new 50-item thermodynamics benchmark, the best LLM scored 82%, below the authors' 95% tutoring-safety threshold, with diagram-based questions near chance.","lead":"Researchers tested 19 large language models on a new 50-question undergraduate thermodynamics test and found that none passed their 95% accuracy bar; the best model scored 82%. The failures concentrated on irreversible processes and on reading thermodynamic diagrams, suggesting current AI assistants cannot yet tutor this subject unsupervised.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No human-expert baseline on UTQA: without evidence that competent thermodynamics instructors reach ≥95% on the same 50 items, the 'unsuitable for tutoring' conclusion remains unanchored from item quality.","rationale":"The reader's conditional verdict is well aligned with my read: the benchmark is plausible and the text/diagram gap is large, but the central suitability conclusion depends on an unvalidated threshold and an unvalidated instrument. I am partial rather than full agreement because the reader frames the 95% threshold as the weakest assumption, whereas I see the lack of human-expert performance on UTQA as the more fundamental concern: even a well-justified threshold would be meaningless if the items are not demonstrably answerable by competent humans at that level. The paper's own internal note ('OLD STUFF before problem with question set spottet') strengthens this concern by suggesting that some reported evaluations may not correspond to the final benchmark. I considered alternative concerns—the cross-model σ≈0.05 being borrowed from gpt-4o prompt runs, and the below-chance 6% score for gpt-4.1 on diagrams—but these are secondary: the headline gap (best 82% vs 95%; diagram 32% vs text 67%) is large enough that moderate statistical corrections would not overturn it. A human-baseline test is the single check that would settle whether the 'not suitable for unsupervised tutoring' claim is an artifact of the test or a real property of current LLMs. The verdict remains CONDITIONAL (equivalently UNCHANGED from the reader): the paper should be revised to add expert human data and to resolve the internal figure/version inconsistency before its deployment guidance is used.","tokens_in":12535,"tokens_out":6699,"duration_ms":79069,"concrete_test":"Recruit at least 10 thermodynamics experts (faculty, postdocs, or advanced PhD students) not involved in item creation. Administer the published UTQA in identical single-choice format, including all 17 diagrams, under untimed conditions. Report mean accuracy, per-item agreement (e.g., Fleiss' kappa), and flag every item where >20% of experts select a non-keyed option. In parallel, compare the question set used for Fig. 3 (prompt experiments) against the released dataset: if items changed after the 'problem with question set' note, rerun the gpt-4o baseline and prompt comparisons on the final set. Decision rule: if expert mean accuracy ≥95% with high agreement, the LLM shortfall is supported; if experts score below 95% or several items are disputed, the threshold must be recalibrated or the items revised before any deployment conclusion is drawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—no 2025 LLM reaches 95% accuracy, hence none is suitable for unsupervised thermodynamics tutoring—presupposes that UTQA items are a fair, unambiguous measure of what a competent tutor knows. The paper provides no human-expert baseline: no data on how thermodynamics faculty or PhD-level experts perform on the same 50 items, no per-item agreement statistics, and no demonstration that the 17 diagram items are decidable at the image resolution used. Without such a baseline, the absolute 95% bar is unanchored: if expert accuracy on UTQA is itself below 95% because of item ambiguity, diagram artifacts, or deliberate emphasis on irreversible edge cases, then the 'not suitable' conclusion would reflect instrument difficulty rather than an LLM-specific deficit. This concern is amplified by internal evidence: immediately preceding the Fig. 3 caption is the stray note 'OLD STUFF before problem with question set spottet', indicating that some displayed prompt results may have been generated before a known problem with the question set was identified. The final released benchmark may therefore differ from what was actually evaluated in those figures. The reader's weakest_assumption correctly notes the imported 95% threshold; but the deeper load-bearing issue is the unmeasured construct validity of the instrument, which is required for any threshold interpretation. A low human score on the same items would make the headline claim misleading even if all raw accuracy numbers are accurate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces UTQA, a 50-item single-choice undergraduate thermodynamics benchmark (33 text-only, 17 diagram-based), and uses it to evaluate 19 commercial LLMs plus 17 prompting variants on gpt-4o. The headline finding is that no 2025-era model reaches the authors' 95% accuracy threshold for unsupervised tutoring; the best model (gpt-o3) scores 82% overall. Text-only accuracy is substantially higher than diagram accuracy (67% vs. 32% averaged across models), and the authors attribute the gap to failures in finite-rate/irreversible reasoning and in mapping visual diagram features to thermodynamic meaning. The paper concludes that current LLMs are not yet reliable enough for unsupervised undergraduate thermodynamics tutoring, while noting that accuracy is necessary but not sufficient for tutoring quality.","tokens_in":12914,"tokens_out":4800,"duration_ms":55049,"significance":"If the results hold, UTQA is a useful, publicly released benchmark for a domain that is underrepresented in existing LLM evaluation suites. The paper has clear strengths: the benchmark is downloadable, full prompts and solutions are promised in the SI, items were expert-validated, controlled linguistic degradations are systematically designed, and the text-only vs. diagram gap is a striking and actionable finding. The strongest raw result—no leading model exceeds 82% on a 50-item expert-constructed test—appears robust in direction, and the evidence that diagram comprehension is the main bottleneck is plausible. However, the paper's suitability conclusion depends on an absolute accuracy threshold whose calibration is not established, and several supporting claims (human equivalence, finite-rate difficulty concentration) lack direct measurement. The benchmark itself is a contribution; the interpretive framing needs substantial reinforcement.","major_comments":[{"comment":"The central claim—that no LLM is suitable for unsupervised tutoring because none reaches 95% on UTQA—requires an anchor showing that UTQA items are a fair measure of competent human tutor knowledge. No human-expert baseline is reported: no data on how thermodynamics faculty or PhD-level experts perform on the same 50 items, no per-item expert agreement, and no demonstration that the 17 diagram items are decidable from the image resolution used. If expert accuracy on UTQA is itself below 95% (due to item ambiguity or deliberate emphasis on edge cases), the conclusion would reflect instrument difficulty rather than an LLM-specific deficit. The paper should add a human-expert baseline or explicitly reframe the claim as 'below an aspirational threshold' rather than 'unsuitable for tutoring.'","section":"Conclusions / Fig. 8"},{"comment":"Immediately before the Fig. 3 caption, the manuscript contains the leftover text 'OLD STUFF before problem with question set spottet' (sic). This is internal evidence that some displayed prompt results may have been generated before a known problem with the question set was identified. The paper must clarify whether the prompt-comparison results in Figs. 2–4, and the gpt-4o points in Figs. 5 and 8, were obtained with the final, released UTQA item set. If any displayed results predate a question-set fix, those numbers cannot be interpreted as evaluating the released benchmark.","section":"Figure 3 caption / Results"},{"comment":"The run-to-run scatter σ≈0.05 is estimated from 51 gpt-4o prompt-variation batches (Methods, Fig. 2a), but it is then applied as a universal per-condition uncertainty to all 19 cross-model comparisons, including the 17-item diagram subset. For a 17-item binary test, the binomial standard error at p≈0.5 is roughly 0.12, so several adjacent model differences in Fig. 6 are not significant under a test-specific error. The paper should report run counts and standard errors for each model and subset rather than a single pooled σ, or explicitly justify why the pooled estimate is transferable.","section":"Methods / Figs. 5–6"},{"comment":"The claim that the performance gap 'concentrates in finite-rate/irreversible scenarios' is not supported by a defined item subset or statistical comparison. The text gives illustrative examples but no per-category accuracies, item IDs, or a significance test for the category. Similarly, the statement that the strongest text-only models 'approach the level of a well-prepared graduate tutor' is a human-equivalence claim with no human data behind it. Define the finite-rate/irreversible item subset and report its accuracy separately, or soften these claims to what the data directly show.","section":"Results: Common strengths and weaknesses / Conclusions"},{"comment":"The 95% competence threshold is imported from VanLehn (2011), which is a review of tutoring effectiveness expressed in effect sizes, not a calibration of accuracy thresholds for an MCQ test. The paper does not justify translating that literature into a 95% accuracy cut on this specific 50-item instrument. Because the threshold is decisive for the suitability conclusion, its calibration needs direct support—for example, human expert scores, an analysis of the consequences of an 18% vs. 5% error rate in tutoring, or a learning-outcome criterion. Without this, the suitability conclusion is an interpretive assertion, even though the raw accuracy numbers remain informative.","section":"Introduction / Conclusions"}],"minor_comments":[{"comment":"The caption states the figure shows '33 text-only questions,' but the panel labels mention 'no image interpretation' and 'image interpretation problems,' and the caption contains the typo 'spottet.' Please align labels and text and remove the leftover note.","section":"Figure 3"},{"comment":"The diagram-based results appear in both Fig. 5(b) and Fig. 6 with inconsistent panel labels ('image interpretation problems' vs. 'diagram based items') and model ordering. Consolidate or clearly distinguish the two figures to avoid duplication confusion.","section":"Figures 5–6"},{"comment":"There are minor typographical issues in the reference list, e.g., 'Langauge' in ref. 31 and 'arXiv preprint' duplicated in ref. 26. Please proofread.","section":"Bibliography"},{"comment":"The abstract says 'the best LLMs achieved 82% accuracy,' which matches gpt-o3 in Fig. 8, but the Results section also emphasizes DeepSeek R1 on text-only items. Make the exact aggregation basis (overall vs. text-only) consistent wherever numbers are quoted.","section":"Abstract/Conclusions"}],"recommendation":"major_revision","confidential_remarks":"The paper contains a leftover 'OLD STUFF' note in a figure caption, which is embarrassing but fixable; the more serious issue is that the suitability conclusion depends on a human baseline and a calibrated threshold that are not present. The benchmark itself appears useful and the text-vs-diagram gap is likely real, so I would not reject. The editor may want to ask the authors to deposit the full item-by-item results and human-baseline data as a condition of revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper contributes a real, reusable benchmark (UTQA) and a clean measurement of a large text/image accuracy gap across 19 LLMs. The main conclusion—that current models aren't reliable for unsupervised thermodynamics tutoring—is plausible but not as well supported as the raw numbers suggest. It should go to peer review, and the authors should be asked to anchor or soften the suitability claim.\n\nWhat's genuinely new: a 50-item single-choice benchmark targeting conceptual thermodynamics (entropy, reversibility, finite-rate processes) that existing benchmarks under-serve, with 17 carefully drawn p–V/T–S diagram items. The paper ships the benchmark, full prompts, and command invocations; that reproducibility is the right standard. The gpt-4o prompt-sensitivity study is careful about run-to-run noise (σ=0.05, σ₃≈0.03), and the degradation experiments give educators a usable answer: clarity loss hurts more than orthographic noise.\n\nThe soft spots are in the headline inference. The 95% threshold is imported from a tutoring-effectiveness review and treated as a competence bar without evidence that competent instructors score ≥95% on these exact items. There is no human-expert baseline, no per-item agreement data, and only 17 diagram items—so the categorical claim that diagram-binding is the bottleneck, while suggestive, rests on a small sample. The cross-model error bars are borrowed from gpt-4o runs. On their own these are limitations, not fatal flaws.\n\nTwo things in the manuscript itself need attention. First, the stray note 'OLD STUFF before problem with question set spottet' in the Fig. 3 caption is easy to miss but important: it suggests some displayed prompt results may predate a known problem with the question set. The authors need to state exactly which results are affected and whether the released benchmark matches the evaluated version. Second, the 'well-prepared graduate tutor' comparison in the discussion is asserted, not measured—same missing baseline issue in a different form.\n\nWho is this for? People building AI tutoring systems or domain-specific LLM benchmarks; thermodynamics instructors curious about model limits. The paper is a solid benchmark contribution with a shaky policy conclusion. A serious referee should require the human baseline (even a small one) and cleanup of the stale-note ambiguity before the suitability claim is published. Would I cite it? Yes, as a benchmark source.","headline":"Useful new benchmark with a credible text-vs-diagram gap, but the tutoring-suitability claim rests on an unvalidated 95% bar and a missing human baseline; needs revision, not desk rejection.","tokens_in":13317,"tokens_out":1904,"would_cite":true,"duration_ms":21098,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current LLMs fall short of the 95% accuracy needed for unsupervised thermodynamics tutoring, topping out at 82%.","keywords":["large language models","thermodynamics education","benchmarking","prompt engineering","diagram-based reasoning","reversibility","entropy","educational measurement"],"falsifier":"A concrete falsifier: run UTQA on a future or untested model with the same protocol; if it scores above 95% on both text and diagram subsets, the central suitability claim is directly refuted. A more targeted check: take items the models miss and present the same diagrams with axis labels added or rotated; if accuracy jumps sharply, the bottleneck is label-pairing rather than conceptual diagram binding, weakening the paper's interpretation.","tokens_in":12479,"feed_emoji":"📊","tokens_out":5051,"duration_ms":48887,"temperature":0.7,"pith_summary":"The paper introduces UTQA, a 50-question single-choice benchmark in undergraduate thermodynamics, and reports that the leading 2025-era language models score at most 82% overall, below a 95% accuracy bar the authors take as the reliability threshold for unsupervised tutoring. The shortfall is not uniform: text-only items are answered far better than questions requiring interpretation of p–V, T–S, and similar diagrams, where many models fall to near-chance levels. The benchmark isolates what the models lack: reasoning about finite-rate, irreversible processes and binding visual diagram features to thermodynamic meaning. The authors conclude that current LLMs are not yet suitable as unsupervised tutors in this domain, while noting that accuracy alone is not sufficient for tutoring.","feed_headline":"No LLM clears 95% on thermodynamics tutoring benchmark","feed_subtitle":"Best model scores 82%; diagrams and irreversible-process questions trip up all 19 tested models.","key_machinery":"The central object is UTQA, a 50-item benchmark with 33 text-only and 17 diagram-based single-choice questions. Its design deliberately isolates single constructs (state vs path functions, q/w sign conventions, quasistatic vs non-quasistatic, entropy bookkeeping) and uses distractors that correspond to known student misconceptions, so accuracy measures principle-grounded reasoning rather than surface cueing. The benchmark's power comes from the contrast between items with canonical structure, solvable by one state-function identity, and items requiring multi-constraint integration or diagram-to-thermodynamics binding.","core_discovery":"The central claim is that no leading 2025-era LLM reached the authors' 95% competence threshold on UTQA, a 50-item instrument covering ideal-gas processes, reversibility, and diagram interpretation. The best model achieved 82%; text-only items averaged 67% across 19 models, while diagram-based items averaged 32%. Error analysis locates the bottleneck in two specific capabilities: recognizing when an ideal-gas process is finite-rate or irreversible and therefore resisting quasistatic templates, and mapping perceptual features of diagrams (signed areas, path orientation, leg types, cycle constraints) to thermodynamic quantities. The authors argue this pattern shows fluent recall of canonical r","pith_inferences":["This suggests a testable extension: adding axis labels or verbal descriptions of diagrams to exam prompts could disproportionately improve scores if the binding deficit is genuine.","The 95% threshold, imported from tutoring-effectiveness research, may be conservative for hybrid human-AI tutoring; a supervised setting with model uncertainty flags might be viable below that bar.","UTQA's misconception-based distractors could double as a diagnostic instrument for categorizing LLM errors, such as regime misclassification versus entropy bookkeeping versus area orientation.","If diagram-binding is the bottleneck, progress in visually-grounded reasoning could be tracked directly on this benchmark before general benchmark improvements appear."],"forward_implications":["No model currently reaches the 95% threshold, so unsupervised LLM tutoring in undergraduate thermodynamics is not supported by this evidence.","Diagram-based items are the main bottleneck; even models that parse axes and curves fail to bind them to thermodynamic meaning.","Prompt phrasing, including chain-of-thought and persona variants, shifts accuracy by a small amount but does not close the gap; elimination-style prompts hurt.","Clause count, a simple proxy for linguistic complexity, shows no significant correlation with accuracy over the observed 1–20 clause range.","Improvements in multimodal binding and finite-rate process reasoning are the most likely path to crossing the threshold."],"supporting_citations":[{"why":"Supplies the 95% accuracy threshold for unsupervised instructional use, imported from tutoring-effectiveness research.","marker":"[17]"},{"why":"Establishes the existing benchmark gap: GPQA contains little entropy or reversibility coverage, motivating UTQA.","marker":"[13]"},{"why":"Provides the closest prior college-level science benchmark, whose thermodynamics items are limited to quantitative end-answer calculations.","marker":"[15]"},{"why":"Gives the educational-measurement guidance used to design items that isolate single constructs and minimize extraneous load.","marker":"[22]"},{"why":"Underpins the claim that diagrams give humans a perceptual advantage, framing why diagram-binding failure is significant.","marker":"[39]"},{"why":"Supports the interpretation that chain-of-thought explanations may not faithfully reflect underlying model computation.","marker":"[34]"},{"why":"Supports the observation that models can recruit internal reasoning without explicit reasoning scaffolds, matching minimal-prompt performance.","marker":"[36]"}],"fun_headline_variants":["Thermodynamics benchmark: best LLM hits 82%, not 95%","LLMs choke on diagram-based thermodynamics questions","AI tutors can't grasp irreversible processes in new test","Top LLM fails 18% of undergraduate thermodynamics items","No LLM reaches tutor-ready 95% on thermodynamics"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim depends on treating 95% accuracy on UTQA's 50 items as the right bar for unsupervised tutoring; if the items overstate difficulty or a lower accuracy is acceptable with safeguards, the suitability conclusion could change even though the raw scores stand.","fun_headline_variants_meta":{"raw":{"variants":["Thermodynamics benchmark: best LLM hits 82%, not 95%","LLMs choke on diagram-based thermodynamics questions","AI tutors can't grasp irreversible processes in new test","Top LLM fails 18% of undergraduate thermodynamics items","No LLM reaches tutor-ready 95% on thermodynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000687,"raw_usage":{"total_tokens":2924,"prompt_tokens":691,"completion_tokens":2233,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":2151}},"tokens_in":435,"tokens_out":2233,"duration_ms":17294,"temperature":1.0,"reasoning_tokens":2151,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:17:20.660987+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifier: run UTQA on a future or untested model with the same protocol; if it scores above 95% on both text and diagram subsets, the central suitability claim is directly refuted. A more targeted check: take items the models miss and present the same diagrams with axis labels added or rotated; if accuracy jumps sharply, the bottleneck is label-pairing rather than conceptual diagram binding, weakening the paper's interpretation.","supporting_citations":[{"cited_title":"The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems","cited_arxiv_id":null,"evidence_quote":"Supplies the 95% accuracy threshold for unsupervised instructional use, imported from tutoring-effectiveness research."},{"cited_title":"Extracting Blockchain Concepts from Text","cited_arxiv_id":"2305.10408","evidence_quote":"Establishes the existing benchmark gap: GPQA contains little entropy or reversibility coverage, motivating UTQA."},{"cited_title":"R.; Zhang, S.; Sun, Y.; Wang, W","cited_arxiv_id":null,"evidence_quote":"Provides the closest prior college-level science benchmark, whose thermodynamics items are limited to quantitative end-answer calculations."},{"cited_title":"M.; Rodriguez, M","cited_arxiv_id":null,"evidence_quote":"Gives the educational-measurement guidance used to design items that isolate single constructs and minimize extraneous load."},{"cited_title":"H.; Simon, H","cited_arxiv_id":null,"evidence_quote":"Underpins the claim that diagrams give humans a perceptual advantage, framing why diagram-binding failure is significant."}],"review_version":1}