Pith. sign in

REVIEW 2 major objections 7 minor 24 references

Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees

T0 review · 2 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Under a 'maximal cheating' protocol, GPT-4 would score a 65 percent weighted average on the University of Hull BSc Physics—an upper second class—but cannot pass because it fails compulsory laboratory modules and the final project viva.

desk verdict A genuinely useful whole-degree case study undermined at the quantitative margin by self-graded marks; referee it, but make them show the data. read the letter →

arxiv 2412.01312 v1 pith:65V3XA7Z submitted 2024-12-02 physics.ed-ph

classification physics.ed-ph
keywords ChatGPTGPT-4physicseducationassessmentreformacademicintegritylargelanguagemodelslaboratoryviva
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether a large language model could pass a whole UK physics degree if used as an unauthorised 'maximally intelligent cheat.' It reports that GPT-4, given every advantage the authors could engineer through prompt modification, would score a weighted average of 65 percent across the University of Hull BSc Physics—an upper second class—if the compulsory laboratory modules and the final project viva were ignored. Because those elements are compulsory and cannot be completed by the model, GPT-4 does not actually pass the degree. The authors conclude that most written coursework, take-home exams, and coding assignments in physics are now vulnerable to AI, and that only invigilated in-person examinations, vivas, laboratory skills tests, and presentations remain dependable. The stakes are practical: assessment practice, not just detection, needs to change.

What carries the argument

The mechanism carrying the argument is the 'maximal cheating' testing protocol applied module by module to a 360-credit degree. The protocol treats the model as a clever student: questions are clarified and split into sub-components, answers are expanded, references are requested, and plugins and custom instructions are used to optimise output; outputs are then accepted as-is and graded against published mark schemes. The second load-bearing structure is the degree's compulsory-element rules: modules with lab skills and the final viva fail regardless of written marks, and that is what turns a 65 percent performance into a formal 'no.' This combination produces a task-by-task map of what GPT-4 can and cannot do.

What would settle it

Ask the original module examiners to independently grade the same GPT-4 outputs without knowing they are AI, then recompute the weighted average; the central claim would fail if the re-marked average falls below the 40 percent pass threshold or if written-only modules that the paper marks as passes become fails. A simpler check is to run the same maximal-cheating protocol on a comparable degree at another institution and see whether the written components still clear 40 percent.

Watch

Extended reading notes

Core claim

The central claim is that GPT-4 can perform at a 2:1 level on the written and computational assessments of a complete BSc physics degree, while failing the degree as a whole because of two compulsory in-person hurdles: laboratory skills and the project viva. Under the 'maximal cheating' protocol, the authors allowed themselves to rephrase questions, split multi-part problems, ask for expanded prose and references, and use plugins and custom instructions; they then marked the outputs against the published criteria. By this route GPT-4 scores 55 percent in year one, 69 percent in year two, 64 percent in year three (with the failed project counted as zero), for a rounded weighted average of 65 percent—an upper second class in UK terms. The model's best performances are in generic coding and single-step problems; its failures concentrate in multi-step reasoning, diagram-heavy questions, and interdisciplinary or insight-dependent problems. The paper's answer to its title question is therefore 'no' for GPT-4 alone, but the near-miss is the finding: the written degree is effectively open to AI.

Load-bearing premise

The whole calculation rests on the authors' own grades being fair; if the original examiners would have marked the same answers lower, the 65 percent result is too high.

Editorial extensions

If this is right

  • If the 65 percent result holds, any UK physics programme that relies on coursework, take-home exams, or open-book online exams should treat those components as undefendable against unauthorised GPT-4 use.
  • Compulsory in-person components—laboratory skills, vivas, and invigilated examinations—become the necessary backbone of secure assessment, not an optional extra.
  • Coding-heavy modules are particularly exposed, with ChatGPT scoring near-perfect or first-class marks on generic programming tasks.
  • The authors' recommendation is urgent assessment reform: either redesign tasks to be robust to AI or embed AI use explicitly as a graduate skill; 'business as usual' is not viable.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the same protocol were run on prose-heavy degrees, the pass rate would probably be higher, since the paper shows GPT-4's weakest tasks are interdisciplinary and insight-heavy prose rather than factual writing.
  • The diagram-parsing weakness is likely to erode quickly as multimodal models improve, which would raise scores in modules where graphical questions currently cost marks.
  • Real-time AI copiloting during remote vivas could eventually remove the last in-person barrier the paper identifies, so the protection it offers may have a limited shelf life.
  • A testable extension is to run the protocol on a second institution's physics degree with original examiners marking, which would show whether the 65 percent is specific to Hull's modules or general.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. The paper reports a whole-degree case study in which GPT-4 (ChatGPT) was prompted, under a self-described 'maximal intelligent cheating' protocol, on all summative assessments of the University of Hull BSc Physics programme. The authors assign module marks from the published marking criteria and find that GPT-4 would fail the degree because it cannot complete compulsory laboratory elements and the final project viva, but would otherwise obtain a weighted average of 65%, equivalent to a UK upper second class (2:1). The qualitative findings are that GPT-4 performs very well on coding and single-step problems, moderately on short prose and two-step problems, and poorly on multi-step, diagram-heavy, and interdisciplinary tasks. The paper argues for urgent assessment reform, recommending invigilated in-person examinations, vivas, and laboratory skills testing, or alternatively embedding AI literacy into disciplinary assessment.

Significance. If the quantitative result were independently verifiable, the paper would be an important, unusually comprehensive data point for physics education policy: rather than testing a single course or question bank, it covers an entire three-year degree and shows that written and coding assessments are broadly vulnerable to LLM assistance, while physical and oral components are not. The task-type summary in Table 2 is a useful and falsifiable classification, and the authors are candid about the 'maximal cheating' framing and its inherent human-expert advantage. The main limitation is that the headline 65% figure rests entirely on module marks assigned by the authors themselves, with no published raw outputs, no inter-rater reliability check, and a protocol that is not fully specified; therefore the precise grade claim is not independently auditable, even though the qualitative conclusion is likely robust.

major comments (2)
  1. [Section 3 and Table 1] The central quantitative claim, that GPT-4 would obtain 65% and a 2:1, rests on module marks that were all assigned by the authors themselves. Section 3 states that 'assessments were marked by the authors in accordance with the published marking criteria/mark schemes' and concedes that future work should have the original assessors grade the outputs, but this was not done. Section 4 further states that 'we explicitly refrain from publishing the entirety of our results and grading due to colleagues' wishes', so an external reader cannot re-grade or audit the individual marks. Because the 65% figure is presented in the abstract and conclusions as the paper's main payload, this grading reliability issue is load-bearing. The authors should at minimum provide the raw outputs or a marked sample, and ideally have a second independent marker grade a subset of the assessments, reporting inter-rater agreement.
  2. [Section 3, items (a)–(e)] The prompt protocol is under-specified to the point that the study is not reproducible. Items (a)–(e) permit modifying questions for clarity, splitting questions, expanding answers, obtaining references, and using plugins/custom instructions 'where we deemed appropriate', but the paper does not report which actions were taken for which assessment, how many iterations were run, what system prompts or custom instructions were used, or the exact model version and date for each module. This ambiguity matters because the final submitted answers are not purely ChatGPT output: they include human expert expansion and reference augmentation, so the statement that 'ChatGPT would pass' conflates the model with a human-assisted workflow. The authors should publish full prompt logs and transcripts, or clearly mark each module as 'ChatGPT alone' versus 'human-assisted ChatGPT'; without that, the headline grade cannot be independently reproduced.
minor comments (7)
  1. [Section 5] The arithmetic underlying the final weighted average should be corrected: the paper reports a 76% average over five passed third-year modules, then says awarding zero for the project gives a 64% third-year average, but (76×5 + 0)/6 = 63.3, not 64. The final 2:1 classification is unchanged by this rounding, but the inconsistency should be fixed and the calculation shown explicitly.
  2. [Section 4, first and fourth paragraphs] The sentence 'ChatGPT was prompted over a period of November 2023 through to February 2024. Image analysis and processing was not possible or implemented in ChatGPT during this period' appears twice, once at the start of Section 4 and again in the fourth paragraph; one occurrence should be removed.
  3. [Table 2] The row 'Invigilated examinations' is not a GPT-4 task type but a testing condition that prevents unauthorised use; listing it alongside 'Vivas' and 'Laboratory/practical skills' in a table of GPT-4 performances is conceptually confusing. Either reword the row as 'Assessment modes where GPT-4 cannot be used undetected' or move it to the discussion.
  4. [Section 2] The paper says that due to COVID-19, 'some examinations were taking place online at the time' but does not specify which of the modules in Table 1 were originally invigilated versus online open-book. Since the later argument distinguishes invigilated examinations from take-home ones, this information should be provided for each module.
  5. [Section 6] The passage that asks GPT-4 'what it considers to be problems it would struggle with' is presented as evidence for the model's limitations, but it is a self-report by the model and should be clearly framed as illustrative rather than evidential, or removed.
  6. [Section 6, reference list] The in-text citation 'Bubeck et al. 2013' appears to be a typo for Bubeck et al. (2023), as the reference list entry is dated 2023; please correct the year.
  7. [Section 7] There is a typo in the sentence 'The ability to detect AI in coursework may (or may not) already by impossible' — 'by' should presumably be 'be'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the 65% degree-classification estimate is a transparent weighted average of module marks graded against published external criteria, with no fitted parameters and no load-bearing self-citations.

full rationale

The paper contains no circular derivation. The target quantity, passing the BSc Physics degree, is defined externally by University of Hull regulations (40% aggregate pass mark, compulsory laboratory and viva hurdles, 30/70 weightings for years 2 and 3, UK degree-classification boundaries) and the QAA credit framework, not by any quantity the authors construct. Each module mark in Table 1 is a measurement of GPT-4 output against published marking criteria, and the paper reports numerous failing grades (exam scores of 36%, 38%, and 20%, laboratory marks of 0%, 30%, 15%, and 18%, and 0% for the viva), which is inconsistent with any attempt to manufacture the headline figure. The final 65% / 2:1 claim is transparent arithmetic over the Table 1 values (year-2 average 69%, year-3 average 64%, weighted 0.3*69 + 0.7*64 = 65.5), leaving the classification robust even if individual marks shifted several points. No parameter was fitted to a subset of data and then renamed a prediction; the 'maximal cheating' protocol (Section 3, items a-e) is defined procedurally and independently of outcomes. There are no load-bearing self-citations, no imported uniqueness theorem, and no ansatz smuggled in via citation; all external comparisons (Kortemeyer 2023; Yeadon et al. 2023; Kumar and Kats 2023) come from independent groups. The acknowledged limitation that 'assessments were marked by the authors in accordance with the published marking criteria' (Section 3), with raw outputs withheld 'due to colleagues' wishes,' is a grading-reliability and auditability risk, explicitly conceded and proposed for future work; it is not circularity, because the marks are contestable measurements against an external rubric rather than outputs equivalent to their inputs by construction.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters or invented entities appear; the analysis uses institutional credit weightings and pass marks as fixed inputs. The main additional load-bearing assumptions are about grading validity, prompt-protocol representativeness, undetectability, and the transferability of a single-institution case study.

assumptions (5)
  • domain assumption Observed module marks are valid proxies for the marks original assessors would give.
    Section 3 states assessments were marked by the authors, not the original module assessors, so the central grade depends on this assumption.
  • domain assumption The maximal-cheating prompting protocol represents an upper bound for undetected student use.
    Section 3 acknowledges the authors' teaching-staff insight gives an in-built advantage, so real student results could be lower.
  • domain assumption Undetected use is possible, i.e., ChatGPT output cannot be reliably identified as AI.
    Section 3 explicitly assumes detection is bypassed; if detectors work on GPT-4 output, then the cheating scenario is less realistic.
  • domain assumption University of Hull BSc Physics assessment is representative enough for policy conclusions about physics and other disciplines.
    Section 2 selects the programme by convenience; Section 6 generalizes to other disciplines without comparative data.
  • domain assumption In-person invigilation, vivas, and lab tests block unauthorized AI use.
    Abstract and Section 6 recommend these formats as non-vulnerable, but no live invigilated condition was tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees." pith.science (2026). https://pith.science/paper/65V3XA7Z

@misc{pith2026241201312,
  author       = {Pith},
  title        = {Pith review of: Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/65V3XA7Z}},
  note         = {Machine review of arXiv:2412.01312}
}
read the original abstract

The emergence of conversational natural language processing models presents a significant challenge for Higher Education. In this work, we use the entirety of a UK physics undergraduate (BSc with Honours) degree including all examinations and coursework to test if ChatGPT (GPT-4) can pass a degree. We adopt a "maximal cheating" approach wherein we permit ourselves to modify questions for clarity, split questions up into smaller sub-components, expand on answers given - especially for long form written responses, obtaining references, and use of advanced coaching, plug-ins and custom instructions to optimize outputs. In general, there are only certain parts of the degree in question where GPT-4 fails. Explicitly these include compulsory laboratory elements, and the final project which is assessed by a viva. If these were no issue, then GPT-4 would pass with a grade of an upper second class overall. In general, coding tasks are performed exceptionally well, along with simple single-step solution problems. Multiple step problems and longer prose are generally poorer along with interdisciplinary problems. We strongly suggest that there is now a necessity to urgently re-think and revise assessment practice in physics - and other disciplines - due to the existence of AI such as GPT-4. We recommend close scrutiny of assessment tasks: only invigilated in-person examinations, vivas, laboratory skills testing (or "performances" in other disciplines), and presentations are not vulnerable to GPT-4, and urge consideration of how AI can be embedded within the disciplinary context.

Figures

Figures reproduced from arXiv: 2412.01312 by the authors.

Figure 1
Figure 1. Output from ChatGPT when prompted to give a description on how to sketch a FCC [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. ChatGPT’s response to sketching the Hall Effect. [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Diagrammatical output sketch from ChatGPT to explain 3 and 4 level lasers. [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 21 canonical work pages

  1. [1]

    Education as we know it may well be dead

    Introduction. Education as we know it may well be dead. The release of ChatGPT heralds a new era in education that means the old ways of assessment must adapt or potentially become worthless or at best mistrusted if AI can successfully undertake such assessments (cf., Mahligawati et al. 2023; Polverini & Gregorcic 2024; Susnjak & McIntosh 2024; Yeadon & H...

  2. [2]

    good degree

    GPT-4 performances. Percentage grades are given for the equivalent UK classifications. We note that a “good degree” is regularly taken to mean an upper second class degree or better within the UK and such grades are generally required for access to “graduate jobs”. Percentage grade UK Classification Task types 85-100% High First Class and near-perfect sco...

  3. [3]

    outsource

    and can help to improve their critical thinking and problem-solving skills (Palomba & Banta, 1999). A significant fraction of the above – perhaps the dominant component depending on the exact domain in question – may be now at risk. ChatGPT (https://chat.openai.com/chat) is a state-of-the-art natural language processing (NLP) model developed by OpenAI. It...

  4. [4]

    We describe each module in turn herein and give some examples that we found interesting from the outputs obtained from ChatGPT for illustrative purposes

    Results. We describe each module in turn herein and give some examples that we found interesting from the outputs obtained from ChatGPT for illustrative purposes. We explicitly refrain from publishing the entirety of our results and grading due to colleagues’ wishes. We note here that ChatGPT was prompted over a period of November 2023 through to February

  5. [5]

    maximal intelligent cheating

    Approach. Our approach and philosophy to evaluate the effectiveness of ChatGPT is twofold. Firstly, we adopt a “maximal intelligent cheating” approach. To us, this means using ChatGPT in an intelligent way to extract the maximum possible benefit from it in order to answer questions and assignments that might arise during a degree course where its use is p...

  6. [7]

    performances

    Conclusions. We have presented work that shows GPT-4 is capable of tackling a physics degree. We note explicitly that it fails due to very specific circumstance: the inability to tackle in-person laboratory assessment, and in-person viva assessment. We conclude that there is now a necessity to re-think and revise assessment practice in physics – and other...

  7. [9]

    2023; Yeadon & Hardy 2024; see also Zollman et al., 2024)

    and it must be noted represents a strong advance over GPT-3 models (e.g., Gregorcic & Pendrill 2023; Ramkorun 2023; Tong et al., 2023; West 2023; Yeadon et al. 2023; Yeadon & Hardy 2024; see also Zollman et al., 2024). Table

  8. [11]

    Core to a strategy where AI tools are embedded is the need to support students with critical literacies around their use (Ewen

    supports this ‘embrace and adapt’ strategy (JISC 2023). Core to a strategy where AI tools are embedded is the need to support students with critical literacies around their use (Ewen

Show all 24 references
  1. [15]

    Walters, W.H. (2023). The Effectiveness of Software Designed to Detect AI-Generated Writing: A Comparison of 16 AI Text Detectors. Open Information Science, vol. 7, no. 1, 2023, pp. 20220158. Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., ...

  2. [16]

    Journal of Applied Learning and Teaching, 6(2)

    Detecting AI content in responses generated by ChatGPT, YouChat, and Chatsonic: The case of five AI content detection tools. Journal of Applied Learning and Teaching, 6(2). Davies, P. (2019). The role of assessment in supporting student learning. In J. H. F. Meyer & R. Land (E...

  3. [17]

    International Journal for Educational Integrity, 19(1), p.17

    Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. International Journal for Educational Integrity, 19(1), p.17. Ewen, M. (2023). Mapping the potential of AI in the age of competence based higher education WonkHE blog....

  4. [18]

    Physical Review Physics Education Research, 20(1), p.010145

    Cheat sites and artificial intelligence usage in online introductory physics courses: What is the extent and what effect does it have on assessments?. Physical Review Physics Education Research, 20(1), p.010145. Küchemann, S., Steinert, S., Kuhn, J., Avila, K. and Ruzika, S. (...

  5. [19]

    American Journal of Physics, 91(12), pp.955-956

    ChatGPT-4 with Code Interpreter can be used to solve introductory college-level vector calculus and electromagnetism problems. American Journal of Physics, 91(12), pp.955-956. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. (2023). GPT detectors are biased against non-nat...

  6. [20]

    and Rezende, M.F

    López-Simó, V. and Rezende, M.F. (2024). Challenging ChatGPT with Different Types of Physics Education Questions. The Physics Teacher, 62(4), pp.290-294. MacIsaac, D. (2023). Chatbots attempt physics homework—chatgpt: Chat generative pre-trained transformer. The Physics Teache...

  7. [21]

    Journal of Radiological Protection, 44(1), p.013502

    Artificial intelligence model GPT4 narrowly fails simulated radiological protection exam. Journal of Radiological Protection, 44(1), p.013502. Sluijsmans, D., Dochy, F., & Janssens, S. (2004). The complexity of assessment: From notion to implementation. Educational Research Re...

  8. [22]

    and McIntosh, T.R

    Susnjak, T. and McIntosh, T.R. (2024). ChatGPT: The end of online exam integrity?. Education Sciences, 14(6), p.656. Tong, D., Tao, Y., Zhang, K., Dong, X., Hu, Y., Pan, S. and Liu, Q. (2023). Investigating ChatGPT-4’s performance in solving physics problems and its potential ...

  9. [24]

    International Journal for Educational Integrity, 19(1), p.26

    Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(1), p.26. West, C.G. (2023). AI and the FCI: Can ChatGPT project an understanding of introductory physics?. arXiv preprint arXiv:2303.01067. Wulff, P. (2024). Physics language...

  10. [55]

    maximally intelligent cheat

    This results in an average of 55%. However, GPT-4 has already failed the degree since as we note above, it cannot pass the laboratory hurdle, and so would not progress to the next stage of the degree programme. Our “maximally intelligent cheat” might therefore focus on the hur...

  11. [2001]

    and allow educators to provide more targeted support and feedback (Angelo & Cross, 1993). In terms of best practice, the literature strongly supports the use of authentic assessments (Palomba & Banta, 1999), which are tasks that closely resemble the challenges students will fa...

  12. [2022]

    maximally intelligent cheat

    ensuring that students do not become uncritically reliant on them. This too presents a need to rethink what authentic assessment is, and how it might be adapted to reflect emerging workplace trends. As such, the use of AI tools becomes a crucial graduate skill (Beckingham et a...

  13. [2023]

    It also goes further than previous work that use either previous versions of GPT or introductory courses (e.g., Kortemeyer 2023; Tong et al., 2023)

    suggest a potentially weak performance in physics examinations – especially those that move away from fact-based recall and incorporated more open book responses with humans retaining some competitiveness (see also Susnjak & McIntosh 2024), but stronger performance in coding. ...

  14. [2024]

    In undertaking the above, we acknowledge that we have a significant in-built advantage

    Image analysis and processing was not possible or implemented in ChatGPT during this period. In undertaking the above, we acknowledge that we have a significant in-built advantage. By definition, as university teaching staff we possess insight that a typical student might othe...

  15. [2026]

    Low quality might take until 2030 to fully harvest, and imaging data

  16. [2060]

    We need to reform how we undertake assessment with celerity

    When all human output (to date) has been harvested, do we consider ourselves superior any longer? This says nothing of the possibility of artificial general intelligence either. We need to reform how we undertake assessment with celerity. Acknowledgments. KAP thanks colleagues...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.