REVIEW 3 major objections 4 minor 2 references
The Impact of AI on Educational Assessment: A Framework for Constructive Alignment
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper argues that whether AI undermines an assessment depends on the Bloom level of the objective, and that AI permissions should be set per component and match between formative and summative tests.
desk verdict A useful practitioner framework for AI policies in assessment, with a real gap in its central proposition and thin empirical support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the pairing of Bloom's taxonomy with a four-level AI-usage scale, attached to Constructive Alignment's requirement that assessment match learning objectives. Bloom levels rank objectives from Remember to Create; the AI scale ranks assistance from 0 (no AI) to 3 (AI produces output the student copies without checking). The argument runs through the observation that assessments are made of components: each component can be assigned its own acceptable AI level, and the paper's Proposition 1 requires that level to be identical in formative and summative contexts.
What would settle it
A controlled comparison would settle it: take one learning objective, assess students under an AI-forbidden and an AI-allowed condition, then probe mastery with a short oral or written exam; if low-level objectives lose no validity when AI is allowed, or if high-level objectives lose validity when AI is allowed, the Bloom-level alignment claim fails.
Extended reading notes
Core claim
The central claim is a validity condition: student assessment remains valid in the age of AI only if the permitted AI level is matched to the Bloom level of each learning objective being assessed, and if the same permission applies to formative and summative assessment. The paper defines four levels of AI use — no AI, AI feedback, AI completes with student verification, AI completes without verification — and argues that the unverified level is never acceptable. Higher Bloom levels are claimed to tolerate more AI assistance; lower levels (Remember, Understand) demand AI-free, controlled assessment. The accompanying survey is used to show that lecturers disagree about enforcement and that their personal AI use correlates with their policy preferences, implying that individual discretion is biased and that university-level guidelines are needed.
Load-bearing premise
The whole framework hinges on the claim that higher Bloom levels genuinely warrant more AI assistance, which the paper supports only with a small, self-selected survey of lecturers' opinions and manual Bloom-level coding, not with evidence about what students actually learn or how AI changes their work.
Editorial extensions
If this is right
- For Remember and Understand objectives, the paper concludes that AI should be forbidden in summative assessment and enforced through controlled, closed-book examinations.
- For Apply, Analyze, Evaluate, and Create objectives, AI policy should be set at the component level by the course coordinator, with oral or written verification available when take-home work is involved.
- AI level 3, where AI produces the work and the student does not verify it, is unacceptable in all cases because students remain responsible for their own output.
- If AI is allowed in a summative assessment, it must also be allowed in the corresponding formative assessment, so that practice and final conditions match.
- Because lecturers' own AI use biases their rules, universities and faculties should issue structured guidelines and train staff on AI capabilities and limits rather than leaving policy to individual lecturers.
Reading between the lines
- A direct consequence not drawn by the paper: the component-level AI scale gives a template for standard AI-use labels on syllabi, making AI policy as transparent as grading criteria.
- The bias finding implies that decentralizing AI decisions to individual instructors creates inequity across sections of the same course; institution-level guidelines are therefore not just administrative convenience but a fairness mechanism.
- Proposition 1 could be tested directly: courses that allow AI in formative work but ban it in summative work should produce larger mismatches between practice and final performance than courses with aligned policies.
- The paper's Bloom-level dependence is a hypothesis about learning, not just policy; an observational study tracking student outcomes across courses that adopt different AI levels would separate lecturer perception from actual validity loss.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a framework, grounded in Constructive Alignment theory and Bloom's taxonomy, for deciding whether and at what level AI use should be permitted in student assessments. It introduces a four-level AI-usage scale for assessment components, argues that the acceptable AI level increases with Bloom level, and states Proposition 1, which aligns AI policy between formative and summative assessment. The framework is illustrated with examples and a small survey of TU Delft course coordinators (n=24; 16 responses used for the Bloom-level analysis). The paper concludes with recommendations for university-wide guidelines and teacher training.
Significance. If the framework is accepted, it offers a practical decision-support tool for an urgent problem in higher education and gives educators a shared vocabulary for AI policy. The component-level decomposition and the four-level scale are clear and actionable, and the paper honestly acknowledges the limitations of the survey. The main contributions are conceptual rather than empirical: the Bloom-level argument is coherent, but the reported survey evidence is too thin and too weakly analyzed to independently establish the central empirical claims. The paper would be substantially strengthened by repairing the alignment proposition and by adding appropriate statistical support or reframing the survey as illustrative.
major comments (3)
- [Section 2.1 and Section 2.3] Proposition 1 is ill-defined for the multi-component assessments that the framework itself introduces in Section 2.3. Section 2.1 states the proposition about 'assessment' as a whole, but Section 2.3 allows different AI levels for different components of one assessment. For an assessment with one allowed and one forbidden component, the sentence 'AI is allowed in summative assessment' has no determinate truth value, so the biconditional cannot be applied. The per-component version the framework needs, e.g., 'for every component C, the permitted AI level in formative assessment equals the permitted AI level in summative assessment', is never stated or defended. This is a load-bearing gap because the proposition is the central alignment rule.
- [Section 3.1] The claim that 'the responses to the first statement have a significant positive correlation with statement 5 and 6 and a significant negative correlation with statement 3' is reported without correlation coefficients, p-values, confidence intervals, or the number of observations behind each test. With n=24 self-selected respondents, these statements cannot be evaluated, and the conclusion that 'lecturers are generally biased based on their own usage of AI' is not supported by the reported evidence. Please report the actual statistics or soften the claim to a qualitative observation.
- [Section 3.2 and Figure 4] The key empirical claim that higher Bloom levels are associated with higher acceptable AI levels rests on 16 manually classified responses, with Bloom levels filled in by the authors when not reported by the lecturer. No inter-rater reliability, coding protocol, or statistical test is provided for the Bloom-level classifications, and the association is inferred visually from Figure 4. This is insufficient to support the conclusion that 'lecturers indicate that the desired AI level is different depending on the Bloom level'. If the survey is meant to validate the theoretical framework, this analysis needs either proper statistical testing with effect sizes or an explicit downgrade to an illustrative pilot.
minor comments (4)
- [Abstract] The phrase 'assessment has to be adopted accordingly' should be 'assessment has to be adapted accordingly'.
- [Section 3.2] The sentence 'Furthermore, some lecturers believe that when questions become more “specific”, “contextualized” or consider a student’s “reflection”' is grammatically incomplete and should be finished or rephrased.
- [Figures 3 and 4] The box plots do not identify the number of respondents per statement or per Bloom-level group, and the axes and statement labels are hard to read at the current resolution. Please add sample sizes and improve figure readability.
- [Section 2.3] The four AI levels are introduced as an 'increasing reliance' ordering, but the manuscript never states whether the scale is intended to be interval-level or ordinal; this should be clarified because the survey analysis appears to treat it as numerical.
Circularity Check
No significant circularity: the framework is argued from Constructive Alignment, Bloom's taxonomy, and LLM capabilities; the survey is independent evidence rather than a fitted input.
full rationale
This paper is a position piece with an accompanying small survey; walking its derivation chain reveals no step where an output is defined in terms of an input or where a fitted quantity is relabeled a prediction. Proposition 1 (Section 2.1) is a normative alignment rule argued from Bloom et al. (1971) and by analogy with calculators and dictionaries; it is not derived from the survey or from any parameter. The Bloom-level recommendations (Section 2.2) are reasoned directly from external facts about LLM capabilities — e.g., 'Questions typically involve paraphrasing, summarizing or giving examples, which are all exercises that LLMs are highly capable of' — and from Krathwohl's taxonomy, with the higher levels explicitly left to coordinator judgment. The four-level AI-use scale (Section 2.3) is a measurement instrument ordered by increasing reliance on AI, not a quantity fitted to data. The survey (Section 3) supplies independent, self-reported evidence; Section 3.2 reports the Bloom/AI-level association as 'in line with' the theory rather than as the theory's source. The paper cites no prior work by its author, so no self-citation chain exists. The manual Bloom coding step ('we filled in the Bloom level manually in case it was easy to deduce') is a possible coding-bias risk, and Section 4 itself cautions that findings rest on self-assessments subject to bias; these are evidence-quality limitations, not definitional reductions. Finally, the skeptic's objection that Proposition 1 is ill-defined for multi-component assessments (Section 2.1 quantifies over whole assessments while Section 2.3 admits per-component AI levels) is a consistency gap in the rule's formulation, not a circularity: no equation or claim reduces to its own input. Accordingly the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Bloom's taxonomy is a valid framework for classifying learning objectives.
- domain assumption Constructive Alignment is a valid educational design framework.
- ad hoc to paper The four-level AI usage scale is a complete and meaningful categorization of AI use in assessment.
- domain assumption Survey respondents accurately self-report their AI usage and assessment practices.
Cite this review
Pith. "Pith review of The Impact of AI on Educational Assessment: A Framework for Constructive Alignment." pith.science (2026). https://pith.science/paper/5IE66OLW
@misc{pith2026250623815,
author = {Pith},
title = {Pith review of: The Impact of AI on Educational Assessment: A Framework for Constructive Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/5IE66OLW}},
note = {Machine review of arXiv:2506.23815}
}
read the original abstract
The influence of Artificial Intelligence (AI), and specifically Large Language Models (LLM), on education is continuously increasing. These models are frequently used by students, giving rise to the question whether current forms of assessment are still a valid way to evaluate student performance and comprehension. The theoretical framework developed in this paper is grounded in Constructive Alignment (CA) theory and Bloom's taxonomy for defining learning objectives. We argue that AI influences learning objectives of different Bloom levels in a different way, and assessment has to be adopted accordingly. Furthermore, in line with Bloom's vision, formative and summative assessment should be aligned on whether the use of AI is permitted or not. Although lecturers tend to agree that education and assessment need to be adapted to the presence of AI, a strong bias exists on the extent to which lecturers want to allow for AI in assessment. This bias is caused by a lecturer's familiarity with AI and specifically whether they use it themselves. To avoid this bias, we propose structured guidelines on a university or faculty level, to foster alignment among the staff. Besides that, we argue that teaching staff should be trained on the capabilities and limitations of AI tools. In this way, they are better able to adapt their assessment methods.
Figures
Reference graph
Works this paper leans on
-
[1]
Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher education , 32(3):347–364. Bin-Nashwan, S. A., Sadallah, M., and Bouteraa, M. (2023). Use of chatgpt in academia: Academic integrity hangs in the balance. Technology in Society, 75:102370. Bloom, B. S. et al. (1971). Handbook on formative and summative evaluation of student learni...
work page 1996
-
[2020]
Education and Information Technologies, 28(7):8445–8501. Nysom, L. (2023). Ai generated feedback for students’ assignment submissions. OpenAI (2022). https://openai.com/index/chatgpt/, Accessed: 11-11-2024. Oregon State University (2024). https://ecampus.oregonstate.edu/faculty/artificial- intelligence-tools/blooms-taxonomy-revisited/, Accessed: 11-11-202...
arXiv 2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.