Pith. sign in

REVIEW 3 major objections 4 minor 2 references

The Impact of AI on Educational Assessment: A Framework for Constructive Alignment

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that whether AI undermines an assessment depends on the Bloom level of the objective, and that AI permissions should be set per component and match between formative and summative tests.

desk verdict A useful practitioner framework for AI policies in assessment, with a real gap in its central proposition and thin empirical support. read the letter →

arxiv 2506.23815 v2 pith:5IE66OLW submitted 2025-06-30 cs.HC cs.AI

classification cs.HCcs.AI
keywords ArtificialIntelligenceAssessmentBloom'sTaxonomyConstructiveAlignmentLargeLanguageModelsformativesummativelecturerbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

With students routinely using large language models, the paper asks when an assessment still measures what it is supposed to measure. It argues that the answer depends on the Bloom level of the learning objective: for remembering and understanding, AI use should be forbidden; for applying, analyzing, evaluating, and creating, course coordinators may allow AI at specified levels per component. The paper further claims that formative and summative assessment must allow the same degree of AI, since otherwise practice does not prepare students for the final test. It also reports a 24-lecturer survey suggesting that lecturers' own AI habits bias how much AI they permit, which motivates institution-level guidelines and staff training rather than individual discretion.

What carries the argument

The central machinery is the pairing of Bloom's taxonomy with a four-level AI-usage scale, attached to Constructive Alignment's requirement that assessment match learning objectives. Bloom levels rank objectives from Remember to Create; the AI scale ranks assistance from 0 (no AI) to 3 (AI produces output the student copies without checking). The argument runs through the observation that assessments are made of components: each component can be assigned its own acceptable AI level, and the paper's Proposition 1 requires that level to be identical in formative and summative contexts.

What would settle it

A controlled comparison would settle it: take one learning objective, assess students under an AI-forbidden and an AI-allowed condition, then probe mastery with a short oral or written exam; if low-level objectives lose no validity when AI is allowed, or if high-level objectives lose validity when AI is allowed, the Bloom-level alignment claim fails.

Watch

Extended reading notes

Core claim

The central claim is a validity condition: student assessment remains valid in the age of AI only if the permitted AI level is matched to the Bloom level of each learning objective being assessed, and if the same permission applies to formative and summative assessment. The paper defines four levels of AI use — no AI, AI feedback, AI completes with student verification, AI completes without verification — and argues that the unverified level is never acceptable. Higher Bloom levels are claimed to tolerate more AI assistance; lower levels (Remember, Understand) demand AI-free, controlled assessment. The accompanying survey is used to show that lecturers disagree about enforcement and that their personal AI use correlates with their policy preferences, implying that individual discretion is biased and that university-level guidelines are needed.

Load-bearing premise

The whole framework hinges on the claim that higher Bloom levels genuinely warrant more AI assistance, which the paper supports only with a small, self-selected survey of lecturers' opinions and manual Bloom-level coding, not with evidence about what students actually learn or how AI changes their work.

Editorial extensions

If this is right

  • For Remember and Understand objectives, the paper concludes that AI should be forbidden in summative assessment and enforced through controlled, closed-book examinations.
  • For Apply, Analyze, Evaluate, and Create objectives, AI policy should be set at the component level by the course coordinator, with oral or written verification available when take-home work is involved.
  • AI level 3, where AI produces the work and the student does not verify it, is unacceptable in all cases because students remain responsible for their own output.
  • If AI is allowed in a summative assessment, it must also be allowed in the corresponding formative assessment, so that practice and final conditions match.
  • Because lecturers' own AI use biases their rules, universities and faculties should issue structured guidelines and train staff on AI capabilities and limits rather than leaving policy to individual lecturers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct consequence not drawn by the paper: the component-level AI scale gives a template for standard AI-use labels on syllabi, making AI policy as transparent as grading criteria.
  • The bias finding implies that decentralizing AI decisions to individual instructors creates inequity across sections of the same course; institution-level guidelines are therefore not just administrative convenience but a fairness mechanism.
  • Proposition 1 could be tested directly: courses that allow AI in formative work but ban it in summative work should produce larger mismatches between practice and final performance than courses with aligned policies.
  • The paper's Bloom-level dependence is a hypothesis about learning, not just policy; an observational study tracking student outcomes across courses that adopt different AI levels would separate lecturer perception from actual validity loss.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper proposes a framework, grounded in Constructive Alignment theory and Bloom's taxonomy, for deciding whether and at what level AI use should be permitted in student assessments. It introduces a four-level AI-usage scale for assessment components, argues that the acceptable AI level increases with Bloom level, and states Proposition 1, which aligns AI policy between formative and summative assessment. The framework is illustrated with examples and a small survey of TU Delft course coordinators (n=24; 16 responses used for the Bloom-level analysis). The paper concludes with recommendations for university-wide guidelines and teacher training.

Significance. If the framework is accepted, it offers a practical decision-support tool for an urgent problem in higher education and gives educators a shared vocabulary for AI policy. The component-level decomposition and the four-level scale are clear and actionable, and the paper honestly acknowledges the limitations of the survey. The main contributions are conceptual rather than empirical: the Bloom-level argument is coherent, but the reported survey evidence is too thin and too weakly analyzed to independently establish the central empirical claims. The paper would be substantially strengthened by repairing the alignment proposition and by adding appropriate statistical support or reframing the survey as illustrative.

major comments (3)
  1. [Section 2.1 and Section 2.3] Proposition 1 is ill-defined for the multi-component assessments that the framework itself introduces in Section 2.3. Section 2.1 states the proposition about 'assessment' as a whole, but Section 2.3 allows different AI levels for different components of one assessment. For an assessment with one allowed and one forbidden component, the sentence 'AI is allowed in summative assessment' has no determinate truth value, so the biconditional cannot be applied. The per-component version the framework needs, e.g., 'for every component C, the permitted AI level in formative assessment equals the permitted AI level in summative assessment', is never stated or defended. This is a load-bearing gap because the proposition is the central alignment rule.
  2. [Section 3.1] The claim that 'the responses to the first statement have a significant positive correlation with statement 5 and 6 and a significant negative correlation with statement 3' is reported without correlation coefficients, p-values, confidence intervals, or the number of observations behind each test. With n=24 self-selected respondents, these statements cannot be evaluated, and the conclusion that 'lecturers are generally biased based on their own usage of AI' is not supported by the reported evidence. Please report the actual statistics or soften the claim to a qualitative observation.
  3. [Section 3.2 and Figure 4] The key empirical claim that higher Bloom levels are associated with higher acceptable AI levels rests on 16 manually classified responses, with Bloom levels filled in by the authors when not reported by the lecturer. No inter-rater reliability, coding protocol, or statistical test is provided for the Bloom-level classifications, and the association is inferred visually from Figure 4. This is insufficient to support the conclusion that 'lecturers indicate that the desired AI level is different depending on the Bloom level'. If the survey is meant to validate the theoretical framework, this analysis needs either proper statistical testing with effect sizes or an explicit downgrade to an illustrative pilot.
minor comments (4)
  1. [Abstract] The phrase 'assessment has to be adopted accordingly' should be 'assessment has to be adapted accordingly'.
  2. [Section 3.2] The sentence 'Furthermore, some lecturers believe that when questions become more “specific”, “contextualized” or consider a student’s “reflection”' is grammatically incomplete and should be finished or rephrased.
  3. [Figures 3 and 4] The box plots do not identify the number of respondents per statement or per Bloom-level group, and the axes and statement labels are hard to read at the current resolution. Please add sample sizes and improve figure readability.
  4. [Section 2.3] The four AI levels are introduced as an 'increasing reliance' ordering, but the manuscript never states whether the scale is intended to be interval-level or ordinal; this should be clarified because the survey analysis appears to treat it as numerical.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the framework is argued from Constructive Alignment, Bloom's taxonomy, and LLM capabilities; the survey is independent evidence rather than a fitted input.

full rationale

This paper is a position piece with an accompanying small survey; walking its derivation chain reveals no step where an output is defined in terms of an input or where a fitted quantity is relabeled a prediction. Proposition 1 (Section 2.1) is a normative alignment rule argued from Bloom et al. (1971) and by analogy with calculators and dictionaries; it is not derived from the survey or from any parameter. The Bloom-level recommendations (Section 2.2) are reasoned directly from external facts about LLM capabilities — e.g., 'Questions typically involve paraphrasing, summarizing or giving examples, which are all exercises that LLMs are highly capable of' — and from Krathwohl's taxonomy, with the higher levels explicitly left to coordinator judgment. The four-level AI-use scale (Section 2.3) is a measurement instrument ordered by increasing reliance on AI, not a quantity fitted to data. The survey (Section 3) supplies independent, self-reported evidence; Section 3.2 reports the Bloom/AI-level association as 'in line with' the theory rather than as the theory's source. The paper cites no prior work by its author, so no self-citation chain exists. The manual Bloom coding step ('we filled in the Bloom level manually in case it was easy to deduce') is a possible coding-bias risk, and Section 4 itself cautions that findings rest on self-assessments subject to bias; these are evidence-quality limitations, not definitional reductions. Finally, the skeptic's objection that Proposition 1 is ill-defined for multi-component assessments (Section 2.1 quantifies over whole assessments while Section 2.3 admits per-component AI levels) is a consistency gap in the rule's formulation, not a circularity: no equation or claim reduces to its own input. Accordingly the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters are fitted. The framework's premises are the validity of Bloom's taxonomy and Constructive Alignment as classification schemes, the completeness of the proposed four-level AI scale, and the reliability of the self-reported survey data. These are domain assumptions rather than validated facts.

assumptions (4)
  • domain assumption Bloom's taxonomy is a valid framework for classifying learning objectives.
    Section 2 uses the six Bloom levels as the basis for analyzing AI's impact on assessment validity; the paper does not question this classification.
  • domain assumption Constructive Alignment is a valid educational design framework.
    The paper's foundation is Biggs (1996); the triangle diagram (Figure 2) is taken as given.
  • ad hoc to paper The four-level AI usage scale is a complete and meaningful categorization of AI use in assessment.
    Section 2.3 introduces levels 0-3 and uses them in the survey without demonstrating exhaustiveness or mutual exclusivity.
  • domain assumption Survey respondents accurately self-report their AI usage and assessment practices.
    The survey results in Section 3 rely on self-reported answers; the paper acknowledges this limitation in the conclusion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Impact of AI on Educational Assessment: A Framework for Constructive Alignment." pith.science (2026). https://pith.science/paper/5IE66OLW

@misc{pith2026250623815,
  author       = {Pith},
  title        = {Pith review of: The Impact of AI on Educational Assessment: A Framework for Constructive Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5IE66OLW}},
  note         = {Machine review of arXiv:2506.23815}
}
read the original abstract

The influence of Artificial Intelligence (AI), and specifically Large Language Models (LLM), on education is continuously increasing. These models are frequently used by students, giving rise to the question whether current forms of assessment are still a valid way to evaluate student performance and comprehension. The theoretical framework developed in this paper is grounded in Constructive Alignment (CA) theory and Bloom's taxonomy for defining learning objectives. We argue that AI influences learning objectives of different Bloom levels in a different way, and assessment has to be adopted accordingly. Furthermore, in line with Bloom's vision, formative and summative assessment should be aligned on whether the use of AI is permitted or not. Although lecturers tend to agree that education and assessment need to be adapted to the presence of AI, a strong bias exists on the extent to which lecturers want to allow for AI in assessment. This bias is caused by a lecturer's familiarity with AI and specifically whether they use it themselves. To avoid this bias, we propose structured guidelines on a university or faculty level, to foster alignment among the staff. Besides that, we argue that teaching staff should be trained on the capabilities and limitations of AI tools. In this way, they are better able to adapt their assessment methods.

Figures

Figures reproduced from arXiv: 2506.23815 by the authors.

Figure 1
Figure 1. Relative interest over time for “Artificial Intelligence” according to Google Trends [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Constructive alignment triangle (TU Delft, 2024) [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Perception of lecturers 3.2 Bloom level, AI level and implementation difficulty We asked the lecturers to comment on their learning objectives, assignments, and the de￾sired AI levels. Lecturers were encouraged to decompose every assignment into components as described in Section 2.3, but some of them only reported a single AI level for the entire assignment. Lecturers were also asked to report the Bloom level of th… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: AI level and difficulty to enforce for varying Bloom levels [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

2 extracted references · 1 canonical work pages

  1. [1]

    Artificial Intelligence

    Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher education , 32(3):347–364. Bin-Nashwan, S. A., Sadallah, M., and Bouteraa, M. (2023). Use of chatgpt in academia: Academic integrity hangs in the balance. Technology in Society, 75:102370. Bloom, B. S. et al. (1971). Handbook on formative and summative evaluation of student learni...

  2. [2020]

    Nysom, L

    Education and Information Technologies, 28(7):8445–8501. Nysom, L. (2023). Ai generated feedback for students’ assignment submissions. OpenAI (2022). https://openai.com/index/chatgpt/, Accessed: 11-11-2024. Oregon State University (2024). https://ecampus.oregonstate.edu/faculty/artificial- intelligence-tools/blooms-taxonomy-revisited/, Accessed: 11-11-202...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.