REVIEW 4 major objections 5 minor 5 references
Distinguishing Fact from Fiction: Student Traits, Attitudes, and AI Hallucination Detection in Business School Assessment
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read In graded coursework with ample time, only one in five students caught a ChatGPT error — and targeted feedback erased the gap by the exam.
desk verdict Real classroom evidence that only 20% of students spot a genuine AI hallucination, but the predictor claims are undermined by measuring attitudes after the outcome. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the study is the assessment instrument itself: a ChatGPT 3.5-generated answer to an econometrics question on selection bias, embedded in a summative coursework worth 20% of the module grade. The hallucination is the AI's conflation of two concepts the course deliberately separated — selection bias as a threat to internal validity, and sample selection bias as a sampling problem — and students had to name that error to pass the question. A second moving part is the 100-word reflective paragraph students wrote after evaluating the AI answer; the same paragraph is mined for two supposed predictors, writing quality (via human marks and standard readability formulas) and AI sentiment (via lexicon-based sentiment analysis). A third is the matched follow-up: the same concept appears in an optional exam question, letting the authors ask whether detection transfers to a higher-stakes setting, with detailed feedback as the intervening treatment.
What would settle it
Run the same detection task on a new cohort with AI sentiment measured before the task, for instance by having students write the reflection before reading the AI answer; if the negative-sentiment coefficient vanishes or reverses, the paper's claim that scepticism drives detection is refuted. A direct check: compare reflections from students who evaluated the AI answer first with reflections from a control group who wrote first — if the evaluation-first group is systematically more negative, the sentiment measure is contaminated by the detection experience.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that AI hallucination detection in a high-stakes but unhurried assessment is rare, patterned, and teachable. Only 20% of the cohort correctly identified that the ChatGPT 3.5 response conflated selection bias (a threat to internal validity in causal inference) with sample selection bias (non-representativeness of the sample), the one error the marking scheme required students to name to score above the pass threshold. Detection was positively associated with marks for the statistics component of the coursework, with year-one statistics performance, with an interpretive question requiring synthesis and evaluation, with human-marked writing quality, with readability measures of the students' own reflective paragraphs, and with a more negative, sceptical sentiment toward AI; it was negatively associated with performance on a procedural, instruction-following question, and female students were more likely to detect. When a related optional question appeared in the final exam, detectors were neither more likely to attempt it nor to score higher — if anything, high coursework AI marks predicted slightly lower exam marks on the topic — which the authors read as evidence that specific, timely feedback had already equalised understanding across the cohort before the exam.
Load-bearing premise
The load-bearing premise is that AI sentiment and writing quality, both read from the 100-word reflective paragraph students wrote after completing the detection task, capture stable attitudes and skills that pre-date the task, rather than being coloured by the experience of having just found or missed the planted error; the claim that AI scepticism causes detection depends on this.
Editorial extensions
If this is right
- If only one in five students can catch a hallucination even with ample time and high stakes, unaided vigilance in faster, lower-stakes settings will be rarer still; AI literacy has to be taught, not assumed.
- Rote procedural competence did not protect against AI misinformation — in this data it predicted worse detection — so curricula that train only rule-following leave students exposed.
- Critical evaluation of AI output behaved like an epistemic skill — a capacity to evaluate, justify, and revise knowledge claims — with statistics knowledge, interpretive reasoning, writing fluency, and a sceptical stance each contributing independently to detection.
- Because targeted feedback equalised exam performance between detectors and non-detectors, a failed detection in a coursework exercise does not compound into later disadvantage.
- The assessment design — an AI error planted in a graded task, followed by a matched exam item — offers a replicable template for measuring and training AI literacy in other disciplines.
Reading between the lines
- A testable extension the paper does not run: measure AI sentiment before the detection task, because the reflective paragraph was written after students had or had not just found the planted error and the negative-sentiment coefficient may partly reflect that experience rather than a stable attitude causing detection.
- The negative association between procedural skill and detection hints at a substitution effect worth probing experimentally: students trained to trust step-by-step procedures may extend that trust to the AI's well-structured, authoritative answer.
- Transferred to professional settings, the results imply that domain experts themselves, not only lay users, can be the weak link in verifying AI output, and that organisations should pair AI tools with verification protocols rather than rely on employee scepticism.
- If the feedback-equalization result generalises, it predicts that formative AI-evaluation exercises will narrow rather than widen achievement gaps, because the students who gain most are precisely those who initially fail to detect errors.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports on a naturalistic assessment in a Year 2 econometrics course at a UK business school, where 211 students evaluated a ChatGPT-generated answer that contained a conceptual error (conflating selection bias with sample selection bias). The main descriptive finding is that only 20% of students detected the error in this high-stakes coursework. The paper then examines predictors of detection, including academic performance, statistics background, interpretive versus procedural skills, text readability, AI sentiment extracted from a post-task reflective paragraph, writing marks from that same paragraph, and gender. A second analysis tracks whether detection predicts choosing and performing on a related optional exam question. The authors find little evidence of an effect and attribute this to detailed feedback equalizing performance. The paper interprets the results through epistemic cognition, cognitive bias, and transfer-of-learning theories.
Significance. If the descriptive detection rate is taken at face value, the study offers a useful and credible measurement: a realistic, high-stakes setting with a verifiable error, a clear marking rubric, and a sample of motivated students. The assessment design is transparent and reproducible, and the paper is honest in its limitations section about the proxy nature of AI sentiment. However, the paper's central predictive claims for writing proficiency and AI scepticism are not yet supported because those measures are derived from the same post-task reflection as the outcome. The feedback-equalization claim similarly lacks a control group. These limitations are fixable with additional data collection or by reframing the analysis as exploratory, which is why the paper is a borderline case rather than a clear accept or reject.
major comments (4)
- [§3.2, §3.4.2 b, Table 3] AI Sentiment and Marks for Writing are both extracted from Q6ii, the 100-word reflective paragraph that students write immediately after completing the detection task. Because detection itself requires articulating the correct error (Section 3.3), a student who has just succeeded has a concrete mistake to describe, making a more negative and more coherent reflection likely; a student who found nothing has no such content. The negative coefficients on AI Sentiment and the positive coefficients on Marks for Writing in Table 3 are therefore consistent with reverse causality and do not provide evidence for Hypotheses 1d and 1b as statements about stable attitudes or skills. The paper's own Section 5.5 concedes that AI sentiment is a proxy and that direct measures would be more valuable. To support the abstract and Section 5 claims that 'AI scepticism' and writing proficiency predict detection, the authors need a design in which attitudes and writing are measured before the detection task, or at least a robustness check showing that reflection content is not driven by the detection experience.
- [§4.3, Table 4] The claim that 'detailed feedback and structured learning support appears to bridge initial performance gaps' is not supported by the design. Both detectors and non-detectors received the same detailed feedback and model solutions; there is no comparison group that did not receive feedback. The null (or slightly negative) association between Detect and Exam AI Mark in Table 4 is equally consistent with the initial detection measure being noisy or unrelated to later performance, or with selection into answering the optional exam question (only 174 of 211 students). The feedback-based interpretation is post hoc and should be substantially weakened or tested with a control cohort that did not receive the feedback intervention.
- [§4.2, paragraph following Figure 3] The text states that 'Coefficients for AI Sentiment are negative and statistically insignificant at 10% in (2), (3) and (9)', but Table 3 reports -0.012* (0.007), -0.010* (0.006), and -0.034* (0.018) in those columns, all of which are significant at the 10% level; only column 5 (-0.009, SE 0.006) is insignificant. This contradiction directly affects the support claimed for Hypothesis 1d, and as written the abstract's assertion that AI scepticism is a key predictor rests on a misreading of the table. The text should be corrected to match the reported estimates.
- [§3.3, §4.1] The binary Detect variable is defined by scoring above 5 out of 10 on Q6i, a threshold that requires identifying the core misclassification to pass 40% of the marks. The headline finding that only 20% of students detected the hallucination is sensitive to this arbitrary cutoff. The authors should report detection rates under alternative thresholds (e.g., ≥7, ≥8) or explicitly justify the 40% pass mark; otherwise the central descriptive claim is hard to calibrate against other studies.
minor comments (5)
- [§4.1] The sentence 'Figure 3A presents the estimation results ... with full details in Table 2' is incorrect: Figure 3 presents coefficient plots while the full estimation results are in Table 3, not Table 2.
- [§3.4.2 b, Table 3] The labels 'Lexical Complexity' and 'Lexical & Syntactic Complexity' are ambiguous because Readability_FKR increases with readability (easier text) while Readability_CDale decreases with readability. The table should state which index is used in each row and what a positive coefficient means in terms of readability.
- [Abstract, §2.3, §5.1] There are several typos and language issues: 'On the context of management education' in the abstract, 'Additionally,, Pennycook et al.' in Section 2.3, and 'lets students' in Section 5.1; these should be corrected.
- [§3.4.2 b] The phrase 'we relegate the full calculations in Appendix A' should be 'we relegate the full calculations to Appendix A'.
- [Table 3 notes] The note that 'The results remain qualitatively similar (statistical significance) if we used any of the academic skill variables' is not verifiable from the reported tables; the authors should show these specifications in an appendix or omit the claim.
Circularity Check
AI Sentiment and Marks for Writing are measured from the post-detection reflection, so the H1b/H1d 'predictors' may capture the outcome; the 20% detection-rate finding is unaffected.
-
other
[Sections 3.2, 3.4.2(b), 4.2 (Table 3); limitations in 5.5]
"A follow-up question requires a 100-word reflection on AI, serving two purposes: (1) capturing students' immediate sentiment towards GenAI, and (2) measuring English writing proficiency... we construct an AI Sentiment score based on their reflective paragraphs... we complement the readability scores with Marks for Writing Quality, the score given to Question 6(ii) by our markers."
AI Sentiment and Marks for Writing are both derived from Q6ii, the reflective paragraph written immediately after the Q6i detection task, while Detect is scored on Q6i. The reflection prompt asks what students learned from the exercise, so a student who just identified the hallucination has a concrete error (sample selection bias vs. selection bias) to describe, plausibly producing a more negative and more coherent reflection. The Table 3 coefficients for AI Sentiment (negative) and Marks for Writing (positive) as 'predictors' of Detect may therefore be consequences of the outcome rather than evidence for Hypotheses 1d and 1b. The paper's Section 5.5 concedes that AI sentiment is only a proxy and that direct measures would give more valuable insights.
full rationale
This is an empirical study, not a derivation, and there are no load-bearing self-citations, fitted parameters renamed as predictions, or imported uniqueness theorems. The only substantive circularity concern is that two headline predictors (AI Sentiment and Marks for Writing) are constructed from the same 100-word reflection that students wrote immediately after completing the AI hallucination detection task. Because detection is defined by success on that task, the reflection is temporally downstream of the outcome, so the measured 'attitudes' and 'writing quality' may be shaped by whether the student just detected the error. The paper's own limitation statement in Section 5.5 acknowledges that AI sentiment is a proxy and that direct measures would be more valuable. The core descriptive result (only 20% detected the hallucination) and the exam-transfer/feedback results are independent of this measurement-order problem and stand on their own. Accordingly, the circularity score is moderate: the predictor claims for H1b and H1d are compromised by construction, but the study's main empirical contributions are not.
Assumptions & free parameters
free parameters (1)
- OLS regression coefficients in Eq. (1) and (2) =
Statistics Background 0.013-0.020; Procedural Skills -0.026 to -0.038; Gender -0.094 to -0.100; Marks for Writing 0.09…
assumptions (4)
- domain assumption The OLS models assume no omitted variable bias after controlling for the listed predictors and marker fixed effects.
- domain assumption Students' gender inferred from first names via genderize.io and manual verification is accurate.
- ad hoc to paper The Syuzhet sentiment score from the reflective paragraph reflects pre-existing AI scepticism, not the experience of the detection task.
- domain assumption Markers scoring the reflective paragraph were not influenced by their knowledge of the student's detection performance.
Cite this review
Pith. "Pith review of Distinguishing Fact from Fiction: Student Traits, Attitudes, and AI Hallucination Detection in Business School Assessment." pith.science (2026). https://pith.science/paper/HR5DY7NK
@misc{pith2026250600050,
author = {Pith},
title = {Pith review of: Distinguishing Fact from Fiction: Student Traits, Attitudes, and AI Hallucination Detection in Business School Assessment},
year = {2026},
howpublished = {\url{https://pith.science/paper/HR5DY7NK}},
note = {Machine review of arXiv:2506.00050}
}
read the original abstract
As artificial intelligence (AI) becomes integral to the society, the ability to critically evaluate AI-generated content is increasingly vital. On the context of management education, we examine how academic skills, cognitive traits, and AI scepticism influence students' ability to detect factually incorrect AI-generated responses (hallucinations) in a high-stakes assessment at a UK business school (n=211, Year 2 economics and management students). We find that only 20% successfully identified the hallucination, with strong academic performance, interpretive skills thinking, writing proficiency, and AI scepticism emerging as key predictors. In contrast, rote knowledge application proved less effective, and gender differences in detection ability were observed. Beyond identifying predictors of AI hallucination detection, we tie the theories of epistemic cognition, cognitive bias, and transfer of learning with new empirical evidence by demonstrating how AI literacy could enhance long-term analytical performance in high-stakes settings. We advocate for an innovative and practical framework for AI-integrated assessments, showing that structured feedback mitigates initial disparities in detection ability. These findings provide actionable insights for educators designing AI-aware curricula that foster critical reasoning, epistemic vigilance, and responsible AI engagement in management education. Our study contributes to the broader discussion on the evolution of knowledge evaluation in AI-enhanced learning environments.
Reference graph
Works this paper leans on
-
[1]
The estimated effect of the variable of interest becomes confounded by the selection process
Inaccurate Coefficient Estimates: Selection bias can lead to coefficient estimates that are biased and do not accurately reflect the true relationships between variables. The estimated effect of the variable of interest becomes confounded by the selection process
-
[2]
This affects the precision of parameter estimates and the ability to detect true relationships
Inefficient Estimation: Biased estimates resulting from selection bias are also likely to be inefficient, leading to wider standard errors. This affects the precision of parameter estimates and the ability to detect true relationships
-
[3]
inherent difference between the treatment and the control group even in the absence of the treatment
RESEARCH DESIGN We discuss how the data is collected and the context of the empirical analysis, followed by a discussion of how we measure different variables of interest. 3.1. The Year 2 Econometrics Course in our data Our data come from a Year 2 economics coursework assessment at a UK business research-led school. UK economics degrees typically span thr...
work page 2022
-
[4]
Incorrect Inference: If selection bias is not appropriately accounted for, the estimated relationships may not be generalizable to the broader population, leading to incorrect policy recommendations or business decisions
-
[5]
Invalid Hypothesis Testing: Selection bias can invalidate hypothesis tests, leading to incorrect conclusions about the statistical significance of relation- ships. Page 3 of 3 The final exam, worth 80% of the course, takes place two months later and includes an optional short-answer section, where students choose four out of six questions (each worth 5%). ...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.