Pith. sign in

REVIEW 3 major objections 4 minor 17 references

Assessment Twins: A Protocol for AI-Vulnerable Summative Assessment

T0 review · 3 major / 4 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read The paper proposes 'assessment twins'—pairing a GenAI-vulnerable task with a second, less vulnerable task that tests the same learning outcomes—to preserve the pedagogical value of essays while recovering assessment validity.

desk verdict A clearly written, honest conceptual proposal for pairing AI-vulnerable and low-vulnerability assessments—the central diagnostic premise is untested, but it is a useful protocol worth serious peer review. read the letter →

arxiv 2510.02929 v1 pith:Y4EIBZDC submitted 2025-10-03 cs.CY

classification cs.CY
keywords assessmenttwinsgenerativeAIvaliditysummativeacademicintegrityredesignhighereducationcross-verification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish a practical protocol for keeping summative assessments trustworthy in the age of generative AI, without abandoning the formats educators value. Its central claim is that any assessment a student could plausibly complete with GenAI—such as a take-home essay or pre-prepared presentation—should be paired with a second, less vulnerable assessment that targets the same intended learning outcomes through a different mode of evidence, such as an oral defence, group discussion, or in-class demonstration, and is scheduled so the two performances can be cross-checked. It argues that this 'assessment twin' design strengthens validity along six distinct strands because the two pieces of evidence triangulate the same constructs. Why a sympathetic reader should care: the protocol offers a middle path between giving up rich assessment tasks and relying on unreliable AI-detection tools, and it explicitly frames validity rather than surveillance as the goal. The paper also states plainly that no empirical data yet exists and that its central interpretive premise—that discrepancies between twins signal inauthentic learning—could be confounded by legitimate factors like anxiety or format discomfort.

What carries the argument

The central object is the assessment twin itself—a paired assessment design in which two components share the same learning outcomes, use different modes of evidence, and are timed for cross-verification. The mechanism doing the work is the discrepancy check: a large inconsistency between a student's performance in the vulnerable component and the less vulnerable component is interpreted as a flag that the first score may not reflect the learner's own competence. The protocol rests on this diagnostic reading of performance gaps, and it is operationalised through a three-step design process (identify vulnerability, align outcomes and select a complementary task, build an interdependent markin

What would settle it

Run a pilot in which students are randomly assigned to complete the vulnerable component with or without GenAI assistance, and all students take the twin; if the distribution of discrepancies between the two groups overlaps heavily after controlling for test anxiety and format preference, then the cross-verification mechanism cannot distinguish inauthentic completion from ordinary variation.

Watch

Extended reading notes

Core claim

An assessment twin is defined as two deliberately designed, interdependent assessment components that (a) address the same intended learning outcomes, (b) require different modes of evidence or production, and (c) are scheduled so performance on each component can be cross-checked to mitigate a known vulnerability, such as GenAI completion or impersonation. The paper maps generative AI threats onto a six-strand validity framework and argues that the twin design counters each threat: a take-home essay paired with an oral defence confirms that the learner actually holds the knowledge the essay claims; a timed in-class test or practical demonstration verifies that performance is not dependent o

Load-bearing premise

The load-bearing premise is that performance discrepancies between the two twin components can be read as evidence about the authenticity of learning; the paper itself concedes that such inconsistencies could just as easily reflect legitimate factors such as anxiety, uneven skill development, or differences in comfort with assessment formats.

Editorial extensions

If this is right

  • Institutions can keep pedagogically rich but AI-susceptible assignments like essays and research projects, pairing them with in-person oral defences or demonstrations rather than replacing them.
  • Reliance on AI-detection and surveillance tools can be reduced, since verification comes from a second human-observed performance instead of probabilistic text analysis.
  • Twinning scales by cohort: individual vivas for small classes, group discussions or poster presentations for medium classes, and random-sampling verification for large cohorts, with the caveat that sampling dilutes assurance.
  • Marking schemes can make the two components interdependent, for example by requiring a minimum score on the twin before the vulnerable component's score is accepted, thereby addressing score reliability.
  • The framework connects assessment integrity to construct validity, shifting the debate from catching cheating toward designing assessments that generate trustworthy evidence of learning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the twin protocol would be testable as a diagnostic instrument; a controlled study comparing discrepancy rates between students who deliberately used GenAI and those who did not would measure whether the cross-check has signal, and the paper's own limitation note implies such a test is needed.
  • Beyond the paper: if the twin components are meant to measure the same construct, factor-analytic or correlational studies should be run to confirm that the two modes converge; if they diverge, the design may be capturing different skills rather than verifying the same ones.
  • Beyond the paper: the twin concept could be extended to formative contexts, using the second component not as a check on authenticity but as a diagnostic of which aspects of an outcome students have actually mastered.
  • Beyond the paper: the random-sampling suggestion for large cohorts undermines per-student verification; an alternative inference is that the framework is best understood as a cohort-level assurance tool, not an individual-level detection tool, unless the sampling design is justified statistically.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces 'assessment twins,' a protocol for redesigning GenAI-vulnerable summative assessments in higher education. An assessment twin pairs a task susceptible to GenAI completion (e.g., a take-home essay) with a second, less vulnerable task (e.g., an oral viva, group discussion, or practical demonstration) that addresses the same intended learning outcomes through a different mode of evidence. The two components are scheduled closely and cross-checked, with interdependent marking, to triangulate evidence and thereby, the authors argue, enhance validity relative to either component alone. The paper uses Messick's unified validity framework, as operationalized by Shaw and Crisp (2012), to map GenAI threats to content, substantive, structural, consequential, generalisability, and external validity, and proposes a three-step design process (identify vulnerability, align outcomes and choose a twin, develop a marking framework). Implementation strategies for small, medium, and large cohorts are discussed, along with limitations including resource intensity, equity concerns, and the absence of empirical evidence.

Significance. If the central claim is correct, assessment twins offer a practical middle path between abandoning traditional written assessments and relying on surveillance technologies, preserving pedagogical value while adding authenticity checks. The paper's systematic mapping of GenAI threats to six validity strands is a useful conceptual contribution, and the explicit, honest acknowledgment of the lack of empirical support—and the call for controlled pilots, mixed-methods research, and longitudinal studies—is a strength. The paper is also clearly written and provides actionable design steps. However, the load-bearing premise—that performance discrepancies between twin components can be interpreted as evidence about the authenticity of learning—is not established, and the paper's own limitations concede that legitimate factors (anxiety, uneven skill development, format comfort) can produce the same discrepancies. As a protocol proposal, the paper is valuable; as a demonstration that twins 'enhance validity,' it is logically underdetermined. The contribution is therefore a promising framework needing both theoretical tightening and empirical testing.

major comments (3)
  1. [How Do Assessment Twins Enhance Validity? (Table 2)] The central claim that twins 'enhance validity' presupposes that discrepancies between twin components are interpretable as evidence of inauthenticity or GenAI misuse. The paper's own Limitations section acknowledges that 'such inconsistencies could just as easily reflect legitimate factors such as anxiety, uneven skill development, or differences in comfort with assessment formats' (citing Struyven et al., 2005). This concession directly undermines the diagnostic link in Table 2's Structural Validity row, where cross-verification is claimed to 'maintain score reliability.' The paper provides no decision rule linking discrepancy size to a judgment of inauthenticity, and no sensitivity/specificity estimates. Without such a rule, a twin design may introduce construct-irrelevant variance and penalize honest students, damaging the structural and consequential validity it aims to protect.
  2. [Step 3: Develop a Marking Framework] The marking framework example explicitly 'caps scores on a written task if oral explanation is poor.' This operationalizes the problematic assumption that the oral twin is a valid gold standard for the same construct as the written task. If honest students exhibit legitimate cross-component variability (e.g., strong writers with oral communication anxiety, or students with disabilities or language barriers), this rule systematically disadvantages them. The paper acknowledges equity concerns in general terms but does not address the fundamental measurement issue: the less vulnerable component must measure the same true ability with negligible format-specific variance for the cap to be fair. The paper offers no evidence or theoretical argument for this condition, making the marking framework unfair as specified.
  3. [Definition of assessment twins (Introduction)] The definition claims twins provide 'enhancing validity compared to either component considered alone,' but the paper does not resolve the following dilemma: if the less vulnerable twin measures nearly the same construct as the vulnerable task, then the vulnerable task is largely redundant and could be replaced by a single secure assessment; if the twin measures a distinct skill or format, then performance discrepancies become uninterpretable as authenticity signals. The paper asserts that twins use 'different modes of evidence' yet assess 'the same intended learning outcomes,' but does not explain how two tasks can differ in mode while being sufficiently construct-equivalent to make cross-verification diagnostic. This is a logical gap at the core of the proposed framework; it needs a validity argument (or empirical data) addressing construct representation and format effects.
minor comments (4)
  1. [Abstract] The abstract lists 'a three-step design process: identifying vulnerabilities, aligning outcomes, selecting complementary tasks, and developing interdependent marking schemes' — this is four items. The body correctly describes three steps. Please align the abstract with the body.
  2. [General] The repeated 'A PREPRINT' header on each page and the duplicated 'Assessment Twins: A Protocol for AI-Vulnerable Summative Assessment: A PREPRINT' on the first page is a formatting artifact; remove the duplicate.
  3. [Table 3] Table 3 is introduced as 'When (not) to use twins,' but it is not referenced in the text before the 'Strategies for Implementing' section. Consider citing it earlier (e.g., in the Practical Design section) to help readers navigate the boundary conditions.
  4. [Limitations and Future Research] The statement 'no empirical data exists on the effectiveness of this approach' is appropriately honest, but the subsequent phrase 'we believe assessment twins have utility' is presented as a conclusion rather than a hypothesis. Phrasing this as a testable conjecture would strengthen the scientific framing.

Circularity Check

1 steps flagged · score 4.0 of 10

The paper's core validity claim is partly built into its definition of 'assessment twin,' but the framework has independent content and no fitted predictions or load-bearing self-citations.

  1. self definitional [Section 'Assessment Twins' (definition paragraph); applied in 'How Do Assessment Twins Enhance Validity?' and Table 2]
    "Consequently, we define assessment twins as two deliberately designed, interdependent assessment components that (a) address the same intended learning outcomes, (b) require different modes of evidence or production, and (c) are scheduled so that performance on each component can be cross-checked to mitigate a known vulnerability (e.g. GenAI completion, impersonation), thereby enhancing validity compared to either component considered alone."

    The paper's central assertion—that the twin approach enhances validity—is contained in the definiendum: an 'assessment twin' is defined as a pairing that is cross-checked 'thereby enhancing validity.' Later claims (e.g., 'We argue that the twin approach helps mitigate validity threats') restate this stipulation rather than deriving it from Messick's framework or from empirical evidence. The paper concedes 'no empirical data exists on the effectiveness of this approach,' so the enhancement is accepted by definition, not demonstrated. Table 2's 'Validity is Enhanced by' column repackages this stipulated property as a conclusion.

full rationale

The paper is a conceptual protocol, not a fitted or empirical derivation: there are no equations, no parameters fitted to data, and no predictions benchmarked against external results. The validity argument is anchored in external frameworks (Messick 1989/1993; Shaw & Crisp 2012) that are applied rather than assumed, so the use of those frameworks is not circular. The main circularity concern is definitional: the concept is defined so that the desired property ('enhancing validity compared to either component considered alone') is part of the term itself, and the later validity-enhancement argument unpacks that definition. The paper's own Empirical Limitations section concedes the load-bearing cross-verification premise is untested ('such inconsistencies could just as easily reflect legitimate factors such as anxiety, uneven skill development, or differences in comfort with assessment formats'), which underscores that the mechanism is assumed rather than derived. Self-citations (Roe & Perkins; Perkins et al.; Giray et al.) are background or alternative-framework references and are not load-bearing for the twin concept; no uniqueness theorem or ansatz is imported from the authors' prior work. Accordingly, the circularity is mild and localized to the definitional encoding, not a forced chain of self-citation or a renamed fitted result.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The paper relies on domain assumptions about GenAI threats and the security of in-person tasks, plus the untested ad hoc assumption that mismatches between twin components are diagnostic of unearned credit. No free parameters or fitted values appear, as this is a conceptual protocol. The invented entity 'assessment twin' is a policy/design construct with no independent empirical evidence beyond the paper's argument.

assumptions (5)
  • domain assumption Messick's unified validity framework, as operationalized by Shaw and Crisp's six strands, is the correct lens for judging assessment quality.
    Adopted without critique as the paper's evaluative backbone; the entire claim that twins enhance validity is framed as an improvement on these strands (Section 'We frame our understanding of validity through Messick's...').
  • domain assumption Current GenAI models can produce high-quality, human-indistinguishable output for common summative tasks (essays, reports, presentations), threatening validity.
    Given as the motivating premise (Introduction, Literature); supported by citation to prior work but not demonstrated here.
  • domain assumption In-person, oral, or supervised performance tasks are materially less vulnerable to GenAI completion than written take-home tasks.
    Assumed throughout the design discussion (Assessment Twins; Step 2); no empirical support is offered, and the paper acknowledges deepfake/technological manipulation cases for video formats.
  • ad hoc to paper Performance discrepancies between the two twin components can be used as evidence about the authenticity of learning.
    This is the core inference of the cross-verification mechanism; the authors explicitly concede in Limitations that inconsistencies could instead reflect anxiety, uneven skill development, or format comfort, which would break the inference.
  • ad hoc to paper Interdependent marking schemes (e.g., capping written scores if oral performance is poor) are fair and do not introduce construct-irrelevant variance.
    Proposed as a key design element (Step 3, Table 2) without pilot evidence; grading interdependence could penalize students for non-target skills.
invented entities (1)
  • assessment twin
    purpose: Paired summative assessment components addressing the same learning outcomes through different evidence modes, designed for cross-verification against GenAI vulnerability.
    A newly named protocol; no validation data yet. The paper proposes future pilot studies, so the concept currently has no falsifiable handle outside the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Assessment Twins: A Protocol for AI-Vulnerable Summative Assessment." pith.science (2026). https://pith.science/paper/Y4EIBZDC

@misc{pith2026251002929,
  author       = {Pith},
  title        = {Pith review of: Assessment Twins: A Protocol for AI-Vulnerable Summative Assessment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y4EIBZDC}},
  note         = {Machine review of arXiv:2510.02929}
}
read the original abstract

Generative Artificial Intelligence (GenAI) is reshaping higher education and raising pressing concerns about the integrity and validity of higher education assessment. While assessment redesign is increasingly seen as a necessity, there is a relative lack of literature detailing what such redesign may entail. In this paper, we introduce assessment twins as an accessible approach for redesigning assessment tasks to enhance validity. We use Messick's unified validity framework to systematically map the ways in which GenAI threaten content, structural, consequential, generalisability, and external validity. Following this, we define assessment twins as two deliberately linked components that address the same learning outcomes through different modes of evidence, scheduled closely together to allow for cross-verification and assurance of learning. We argue that the twin approach helps mitigate validity threats by triangulating evidence across complementary formats, such as pairing essays with oral defences, group discussions, or practical demonstrations. We highlight several advantages: preservation of established assessment formats, reduction of reliance on surveillance technologies, and flexible use across cohort sizes. To guide implementation, we propose a three-step design process: identifying vulnerabilities, aligning outcomes, selecting complementary tasks, and developing interdependent marking schemes. We also acknowledge the challenges, including resource intensity, equity concerns, and the need for empirical validation. Nonetheless, we contend that assessment twins represent a validity-focused response to GenAI that prioritises pedagogy while supporting meaningful student learning outcomes.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

17 extracted references · 1 canonical work pages

  1. [1]

    While assessment redesign is increasingly seen as a necessity, there is a relative lack of literature detailing what such redesign may entail

    Assessment Twins: A Protocol for AI-Vulnerable Summative Assessment: A PREPRINT 1 ASSESSMENT TWINS: A PROTOCOL FOR AI-VULNERABLE SUMMATIVE ASSESSMENT A PREPRINT Jasper Roe 1*, Mike Perkins 2, Louie Giray 3, 1Durham University, United Kingdom 2 British University Vietnam, Vietnam 3 Mapúa University, Philippines * Corresponding Author: jasper.j.roe@durham.a...

  2. [3]

    For example, if an existing assessment is pedagogically valuable yet GenAI vulnerable, this is the most important criterion for implementing a twin strategy. Even assessments potentially vulnerable to GenAI may still retain pedagogical value, and instructors may still wish to retain these as part of a formal assessment task, rather than changing them to a...

  3. [4]

    Assessment Twins: A Protocol for AI-Vulnerable Summative Assessment: A PREPRINT 4 Table 1: Shaw and Crisp’s (2012) Six-Strands of Validity, based on Messick (1989,

  4. [5]

    Validity Strand Definition Content Relates to the representativeness and relevance of the content. Substantive Relates to the justifications and theoretical basis for the consistency of the assessment, how comparable the underlying cognitive processes are vis-à-vis the assessment and performance in practice. Structural Relates to how reliable the procedur...

  5. [6]

    https://doi.org/10.53761/q3azde36 Perkins, M., Jasper, R., & Furze, L. (2025). Reimagining the artificial intelligence assessment scale: A refined framework for educational assessment. Journal of University Teaching and Learning Practice. https://doi.org/10.53761/rrm4y757 Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat,...

  6. [9]

    Just as importantly, we must explore whether twin performance links meaningfully to real-world competencies, giving us a stronger basis for claims of predictive validity

    or AI-integrated pedagogies (Foung et al., 2024). Just as importantly, we must explore whether twin performance links meaningfully to real-world competencies, giving us a stronger basis for claims of predictive validity. By pursuing these lines of inquiry, we not only fill the current evidence gap but also build a more credible and resilient framework. Co...

  7. [10]

    https://doi.org/10.37074/jalt.2025.8.1.20 Bennett, R. E. (2011). Formative assessment: A critical review. Assessment in Education: Principles, Policy & Practice, 18(1), 5–25. https://doi.org/10.1080/0969594X.2010.513678 Boud, D., Dawson, P., Bearman, M., Bennett, S., Joughin, G., & Molloy, E. (2018). Reframing assessment research: Through a practice persp...

  8. [11]

    https://doi.org/10.1007/s40979-018-0036-7 Roe, J., & Perkins, M. (2022). What are automated paraphrasing tools and how do we address them? A review of a growing threat to academic integrity. International Journal for Educational Integrity, 18(1), Article

Show all 17 references
  1. [15]

    https://doi.org/10.1007/s40979-022-00109-w Roe, J., & Perkins, M. (2024). Generative AI and agency in education: A critical scoping review and thematic analysis. arXiv. https://doi.org/10.48550/arXiv.2411.00631 Rudolph, J., Tan, S., & Tan, S. (2023). ChatGPT: Bullshit spewer o...

  2. [16]

    https://doi.org/10.37074/jalt.2023.6.1.9 Runyon, N. (2025). Deepfakes on trial: How judges are navigating AI evidence authentication. Thomson Reuters Institute. https://www.thomsonreuters.com/en-us/posts/ai-in-courts/deepfakes-evidence-authentication/ Shaw, S., & Crisp, V. (20...

  3. [26]

    https://doi.org/10.1007/s40979-023-00146-z

  4. [53]

    M., & Kinden, C

    https://doi.org/10.1186/s41239-024-00487-w Prentice, F. M., & Kinden, C. E. (2018). Paraphrasing tools, language translation tools and plagiarism: An exploratory study. International Journal for Educational Integrity, 14(1), Article

  5. [72]

    Don’t let Grammarly overwrite your style and voice

    https://doi.org/10.1007/s44217-025-00461-2 Corbin, T., Bearman, M., Boud, D., & Dawson, P. (2025). The wicked problem of AI and assessment. Assessment & Evaluation in Higher Education, 0(0), 1–17. https://doi.org/10.1080/02602938.2025.2553340 Corbin, T., Dawson, P., & Liu, D. ...

  6. [1993]

    Traditional conceptions of assessment validity are classified into three types: content, criterion-related, and construct validity

    work on validity. Traditional conceptions of assessment validity are classified into three types: content, criterion-related, and construct validity. Messick (1989,

  7. [2012]

    Creating an assessment twin is distinct from the simple process of pairing a written essay with a traditional oral viva voce

    or certify achievement (Craddock & Mathias, 2009). Creating an assessment twin is distinct from the simple process of pairing a written essay with a traditional oral viva voce. Notably, a twin task for a GenAI-susceptible assessment could take multiple formats, including a gro...

  8. [2023]

    trustworthy judgements about student learning

    and the use of multiple, inclusive, and contextualised methods of assessment to form “trustworthy judgements about student learning” (TEQSA, 2025, p1). The remainder of this paper is organised as follows: We begin by reviewing the current literature on GenAI and assessment val...

  9. [2025]

    yet still could provide another significant data point in collaboration with other, more significantly GenAI-vulnerable tasks (such as a take-home, written assessment). In-person examinations of course remain an important and secure form of assessment, and could be part of a t...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.