{"id":"0bf96a47-2b24-4202-b3fa-7998831ddf06","arxiv_id":"2412.01312","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4 can pass most written and coding assessments in a UK physics degree and would get a 2:1, but it fails because it cannot complete laboratory work or an oral project defense.","lead":"The authors tested whether ChatGPT (GPT-4) could pass an entire UK physics degree by feeding it exams and coursework from every module at the University of Hull. It earned high marks on many written and coding tasks but failed the degree because it cannot do compulsory laboratory work or an in-person oral exam.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 65% / 2:1 estimate rests on author-assigned marks with no published raw outputs or independent grading; if those marks are systematically generous, the central quantitative claim weakens.","rationale":"The reader identified the same load-bearing weakness: the exact module marks were assigned by the authors, not by independent or original assessors, and no raw outputs were published. That is exactly where the central quantitative claim is least secure. The paper is otherwise a valuable, readable case study: it covers the full degree, uses a clearly described 'maximal cheating' protocol, provides concrete examples of GPT-4 failures and successes, and its qualitative findings align with prior work showing strong performance on coding and simple calculations but weaker performance on multi-step and diagram-heavy problems. The conclusion that ChatGPT cannot pass the degree because it cannot physically complete compulsory laboratory work or an oral viva is robust even if every author-assigned mark were shifted downward by several points. This is why the concern does not justify rejection or a change from the reader's conditional acceptance. The condition should be that the quantitative grade claim be treated as provisional until independent grading or release of raw outputs is available. The proposed test directly targets that condition by comparing blind independent marks against the authors' Table 1 marks on a representative sample of assessments. If the independent marks match, the 65% estimate gains real support; if they do not, the paper's quantitative claim should be softened to a range while preserving its stronger qualitative message.","tokens_in":15019,"tokens_out":4442,"duration_ms":44581,"concrete_test":"Release de-identified ChatGPT outputs and the relevant mark schemes for a stratified sample of at least eight assessments spanning Levels 4-6, examinations and coursework, plus the two computational project components. Have two independent physics instructors, blind to the authors' Table 1 marks, grade the same outputs against the published criteria. Compute the mean absolute difference and the 95% confidence interval between the independent marks and the authors' marks; if the interval excludes zero or the mean absolute difference exceeds 5 percentage points, the 65% headline should be replaced by a range or re-derived with conservative marks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative payload is the claim that ChatGPT would obtain a weighted average of 65%, an upper second class, if compulsory laboratory and viva elements were not barriers. That payload depends entirely on the module marks in Table 1, yet every mark was assigned by the authors themselves. Section 3 states that 'assessments were marked by the authors in accordance with the published marking criteria/mark schemes' and concedes that future work should have the original assessors grade the outputs, but this was not done. Section 4 further states that the authors 'explicitly refrain from publishing the entirety of our results and grading due to colleagues' wishes', so an external reader cannot re-grade or audit the individual marks. The 'maximal intelligent cheating' protocol also permits the authors—who are expert physicists—to split questions, clarify prompts, expand answers, and add references, and then to grade the resulting outputs. This creates a conflict of interest: the protocol is optimized by the same people who decide the marks. If the self-assigned marks are systematically generous, the 65% headline may be too high and the 2:1 conclusion could become a 2:2 or lower. The qualitative conclusion that ChatGPT performs well on many written and coding tasks but cannot complete in-person laboratory work or a viva is likely robust to modest marking shifts, but the precise grade claim is not. Since the paper's abstract and conclusions present the 65% figure as the main result, the grading concern is load-bearing. A second, related gap is that the final-project report was never scored; the overall average includes a zero for the project because of the viva, so the 'if these were no issue' counterfactual is not actually tested. The independent-grading check below would settle the primary concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a whole-degree case study in which GPT-4 (ChatGPT) was prompted, under a self-described 'maximal intelligent cheating' protocol, on all summative assessments of the University of Hull BSc Physics programme. The authors assign module marks from the published marking criteria and find that GPT-4 would fail the degree because it cannot complete compulsory laboratory elements and the final project viva, but would otherwise obtain a weighted average of 65%, equivalent to a UK upper second class (2:1). The qualitative findings are that GPT-4 performs very well on coding and single-step problems, moderately on short prose and two-step problems, and poorly on multi-step, diagram-heavy, and interdisciplinary tasks. The paper argues for urgent assessment reform, recommending invigilated in-person examinations, vivas, and laboratory skills testing, or alternatively embedding AI literacy into disciplinary assessment.","tokens_in":15264,"tokens_out":3685,"duration_ms":35768,"significance":"If the quantitative result were independently verifiable, the paper would be an important, unusually comprehensive data point for physics education policy: rather than testing a single course or question bank, it covers an entire three-year degree and shows that written and coding assessments are broadly vulnerable to LLM assistance, while physical and oral components are not. The task-type summary in Table 2 is a useful and falsifiable classification, and the authors are candid about the 'maximal cheating' framing and its inherent human-expert advantage. The main limitation is that the headline 65% figure rests entirely on module marks assigned by the authors themselves, with no published raw outputs, no inter-rater reliability check, and a protocol that is not fully specified; therefore the precise grade claim is not independently auditable, even though the qualitative conclusion is likely robust.","major_comments":[{"comment":"The central quantitative claim, that GPT-4 would obtain 65% and a 2:1, rests on module marks that were all assigned by the authors themselves. Section 3 states that 'assessments were marked by the authors in accordance with the published marking criteria/mark schemes' and concedes that future work should have the original assessors grade the outputs, but this was not done. Section 4 further states that 'we explicitly refrain from publishing the entirety of our results and grading due to colleagues' wishes', so an external reader cannot re-grade or audit the individual marks. Because the 65% figure is presented in the abstract and conclusions as the paper's main payload, this grading reliability issue is load-bearing. The authors should at minimum provide the raw outputs or a marked sample, and ideally have a second independent marker grade a subset of the assessments, reporting inter-rater agreement.","section":"Section 3 and Table 1"},{"comment":"The prompt protocol is under-specified to the point that the study is not reproducible. Items (a)–(e) permit modifying questions for clarity, splitting questions, expanding answers, obtaining references, and using plugins/custom instructions 'where we deemed appropriate', but the paper does not report which actions were taken for which assessment, how many iterations were run, what system prompts or custom instructions were used, or the exact model version and date for each module. This ambiguity matters because the final submitted answers are not purely ChatGPT output: they include human expert expansion and reference augmentation, so the statement that 'ChatGPT would pass' conflates the model with a human-assisted workflow. The authors should publish full prompt logs and transcripts, or clearly mark each module as 'ChatGPT alone' versus 'human-assisted ChatGPT'; without that, the headline grade cannot be independently reproduced.","section":"Section 3, items (a)–(e)"}],"minor_comments":[{"comment":"The arithmetic underlying the final weighted average should be corrected: the paper reports a 76% average over five passed third-year modules, then says awarding zero for the project gives a 64% third-year average, but (76×5 + 0)/6 = 63.3, not 64. The final 2:1 classification is unchanged by this rounding, but the inconsistency should be fixed and the calculation shown explicitly.","section":"Section 5"},{"comment":"The sentence 'ChatGPT was prompted over a period of November 2023 through to February 2024. Image analysis and processing was not possible or implemented in ChatGPT during this period' appears twice, once at the start of Section 4 and again in the fourth paragraph; one occurrence should be removed.","section":"Section 4, first and fourth paragraphs"},{"comment":"The row 'Invigilated examinations' is not a GPT-4 task type but a testing condition that prevents unauthorised use; listing it alongside 'Vivas' and 'Laboratory/practical skills' in a table of GPT-4 performances is conceptually confusing. Either reword the row as 'Assessment modes where GPT-4 cannot be used undetected' or move it to the discussion.","section":"Table 2"},{"comment":"The paper says that due to COVID-19, 'some examinations were taking place online at the time' but does not specify which of the modules in Table 1 were originally invigilated versus online open-book. Since the later argument distinguishes invigilated examinations from take-home ones, this information should be provided for each module.","section":"Section 2"},{"comment":"The passage that asks GPT-4 'what it considers to be problems it would struggle with' is presented as evidence for the model's limitations, but it is a self-report by the model and should be clearly framed as illustrative rather than evidential, or removed.","section":"Section 6"},{"comment":"The in-text citation 'Bubeck et al. 2013' appears to be a typo for Bubeck et al. (2023), as the reference list entry is dated 2023; please correct the year.","section":"Section 6, reference list"},{"comment":"There is a typo in the sentence 'The ability to detect AI in coursework may (or may not) already by impossible' — 'by' should presumably be 'be'.","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of the journal and the qualitative conclusions are consistent with prior work, but the lack of raw data and independent marking is a serious reproducibility concern for the headline quantitative claim. If institutional constraints prevent full data release, the authors should at least provide a marked sample and a clear separation of ChatGPT-only versus human-assisted outputs. The paper would also benefit from a sensitivity analysis showing how the final degree classification changes under plausible marking variations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this if you care about assessment reform. It's the first study I know that runs GPT-4 against every summative assessment in a full BSc physics degree and converts the results into a UK honours classification. The whole-programme framing is a real step beyond earlier single-course or single-exam tests, and the task-type grade table (Table 2) is a genuinely useful summary that matches what others have found: coding and single-step problems are strong, multi-step and diagram-dependent questions are weak, lab skills and vivas are impossible for a text-only model.\n\nThe paper is also honest about its method. Section 3 describes the 'maximal intelligent cheating' protocol in enough detail to see that the authors were careful: they used published marking criteria, allowed themselves to split and clarify questions as a clever student would, and explicitly note the advantage this gives. They also disclose that GPT-4 was used to help write the introduction, which is more transparency than most.\n\nThe soft spot is exactly where the stress-test says it is. Every module mark in Table 1 was assigned by the authors, who also designed the prompt protocol that produced the answers. That is a conflict of interest, not in the malicious sense but in the practical sense: the same people optimised the cheating and then graded it. The authors acknowledge that having original assessors re-grade would improve the study, but they also explicitly withhold the raw outputs because colleagues asked them not to publish. So an external reader cannot audit the 65%. The final-project report was never scored at all; the zero is there because of the viva, which means the 'if these were no issue' counterfactual wasn't actually tested. If the self-grading was generous, a 65% 2:1 could easily become a 2:2.\n\nThat said, the qualitative burden probably survives. The pattern of performance by task type is consistent with several earlier papers, and the lab/viva failure is structural, not a marking artefact. The recommendation that only invigilated in-person exams, vivas, lab skills and presentations are robust is the weakest inference: the study didn't test presentations, and 'only' is doing a lot of work. But as a prompt for departmental debate, it's serviceable.\n\nBottom line: referee it, but ask for the per-assessment grades and ideally independent marking, or at minimum a much narrower claim in the abstract. The paper earns a revision, not a desk rejection.","headline":"A genuinely useful whole-degree case study undermined at the quantitative margin by self-graded marks; referee it, but make them show the data.","tokens_in":15845,"tokens_out":2373,"would_cite":false,"duration_ms":22925,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under a 'maximal cheating' protocol, GPT-4 would score a 65 percent weighted average on the University of Hull BSc Physics—an upper second class—but cannot pass because it fails compulsory laboratory modules and the final project viva.","keywords":["ChatGPT","GPT-4","physics education","assessment reform","academic integrity","large language models","laboratory assessment","viva"],"falsifier":"Ask the original module examiners to independently grade the same GPT-4 outputs without knowing they are AI, then recompute the weighted average; the central claim would fail if the re-marked average falls below the 40 percent pass threshold or if written-only modules that the paper marks as passes become fails. A simpler check is to run the same maximal-cheating protocol on a comparable degree at another institution and see whether the written components still clear 40 percent.","tokens_in":14821,"feed_emoji":"🎓","tokens_out":7124,"duration_ms":57961,"temperature":0.7,"pith_summary":"The paper asks whether a large language model could pass a whole UK physics degree if used as an unauthorised 'maximally intelligent cheat.' It reports that GPT-4, given every advantage the authors could engineer through prompt modification, would score a weighted average of 65 percent across the University of Hull BSc Physics—an upper second class—if the compulsory laboratory modules and the final project viva were ignored. Because those elements are compulsory and cannot be completed by the model, GPT-4 does not actually pass the degree. The authors conclude that most written coursework, take-home exams, and coding assignments in physics are now vulnerable to AI, and that only invigilated in-person examinations, vivas, laboratory skills tests, and presentations remain dependable. The stakes are practical: assessment practice, not just detection, needs to change.","feed_headline":"GPT-4 scores 65 percent on a physics degree—but fails the lab","feed_subtitle":"Only compulsory labs and the final viva stop GPT-4 from earning an upper second class.","key_machinery":"The mechanism carrying the argument is the 'maximal cheating' testing protocol applied module by module to a 360-credit degree. The protocol treats the model as a clever student: questions are clarified and split into sub-components, answers are expanded, references are requested, and plugins and custom instructions are used to optimise output; outputs are then accepted as-is and graded against published mark schemes. The second load-bearing structure is the degree's compulsory-element rules: modules with lab skills and the final viva fail regardless of written marks, and that is what turns a 65 percent performance into a formal 'no.' This combination produces a task-by-task map of what GPT-4 can and cannot do.","core_discovery":"The central claim is that GPT-4 can perform at a 2:1 level on the written and computational assessments of a complete BSc physics degree, while failing the degree as a whole because of two compulsory in-person hurdles: laboratory skills and the project viva. Under the 'maximal cheating' protocol, the authors allowed themselves to rephrase questions, split multi-part problems, ask for expanded prose and references, and use plugins and custom instructions; they then marked the outputs against the published criteria. By this route GPT-4 scores 55 percent in year one, 69 percent in year two, 64 percent in year three (with the failed project counted as zero), for a rounded weighted average of 65 percent—an upper second class in UK terms. The model's best performances are in generic coding and single-step problems; its failures concentrate in multi-step reasoning, diagram-heavy questions, and interdisciplinary or insight-dependent problems. The paper's answer to its title question is therefore 'no' for GPT-4 alone, but the near-miss is the finding: the written degree is effectively open to AI.","pith_inferences":["If the same protocol were run on prose-heavy degrees, the pass rate would probably be higher, since the paper shows GPT-4's weakest tasks are interdisciplinary and insight-heavy prose rather than factual writing.","The diagram-parsing weakness is likely to erode quickly as multimodal models improve, which would raise scores in modules where graphical questions currently cost marks.","Real-time AI copiloting during remote vivas could eventually remove the last in-person barrier the paper identifies, so the protection it offers may have a limited shelf life.","A testable extension is to run the protocol on a second institution's physics degree with original examiners marking, which would show whether the 65 percent is specific to Hull's modules or general."],"forward_implications":["If the 65 percent result holds, any UK physics programme that relies on coursework, take-home exams, or open-book online exams should treat those components as undefendable against unauthorised GPT-4 use.","Compulsory in-person components—laboratory skills, vivas, and invigilated examinations—become the necessary backbone of secure assessment, not an optional extra.","Coding-heavy modules are particularly exposed, with ChatGPT scoring near-perfect or first-class marks on generic programming tasks.","The authors' recommendation is urgent assessment reform: either redesign tasks to be robust to AI or embed AI use explicitly as a graduate skill; 'business as usual' is not viable."],"supporting_citations":[{"why":"Supplies the prior result that an AI agent could not pass an introductory physics course, the baseline this degree-level study extends.","marker":"Kortemeyer 2023"},{"why":"Provides prior evidence that physics exam performance is weak when questions move beyond fact recall, which the paper's task-level findings echo.","marker":"Yeadon & Halliday 2023"},{"why":"Prior finding that short-form physics essays are vulnerable, underpinning the paper's essay and prose results.","marker":"Yeadon et al. 2023"},{"why":"Scoping review of ChatGPT multiple-choice performance used to align the paper's task-type grading.","marker":"Newton & Xiromeriti 2023"},{"why":"Offers a framework for how large language model design explains strengths and weaknesses in physics output.","marker":"Polverini & Gregorcic 2024"},{"why":"Supports the recommendation that high-stakes supervised examinations are required to preserve academic integrity.","marker":"Kortemeyer & Bauer 2024"},{"why":"Documents bias in AI text detectors against non-native English writers, supporting the paper's claim that detection is unreliable.","marker":"Liang et al. 2023"}],"fun_headline_variants":["GPT-4 nearly passes a physics degree—except for labs and viva","AI scores 65% on physics degree, but fails labs and viva","GPT-4 gets a 2:1 in physics if you skip the lab and viva","Physics degree: GPT-4 aces written parts, fails the hands-on","ChatGPT fails physics degree, but only because of labs and viva"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole calculation rests on the authors' own grades being fair; if the original examiners would have marked the same answers lower, the 65 percent result is too high.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 nearly passes a physics degree—except for labs and viva","AI scores 65% on physics degree, but fails labs and viva","GPT-4 gets a 2:1 in physics if you skip the lab and viva","Physics degree: GPT-4 aces written parts, fails the hands-on","ChatGPT fails physics degree, but only because of labs and viva"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000482,"raw_usage":{"total_tokens":2431,"prompt_tokens":1041,"completion_tokens":1390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":1286}},"tokens_in":657,"tokens_out":1390,"duration_ms":9099,"temperature":1.0,"reasoning_tokens":1286,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:28:46.145063+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask the original module examiners to independently grade the same GPT-4 outputs without knowing they are AI, then recompute the weighted average; the central claim would fail if the re-marked average falls below the 40 percent pass threshold or if written-only modules that the paper marks as passes become fails. A simpler check is to run the same maximal-cheating protocol on a comparable degree at another institution and see whether the written components still clear 40 percent.","supporting_citations":[],"review_version":1}