REVIEW 2 major objections 5 minor 1 cited by
Validity Arguments For Constructed Response Scoring Using Generative Artificial Intelligence Applications
T0 review · 2 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Scoring essays with generative AI demands more validity evidence than traditional automated scoring.
desk verdict Useful extension of validity theory to generative AI scoring, but the empirical support for the abstract's combined-score claim is weak outside TOEFL. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the validity-evidence matrix in Table 1, which takes the five evidence types from the 2014 testing Standards (internal structure, relations to external variables, response processes, test content, and consequences of use, plus fairness) and adds a column specifying what must be documented when scores come from a generative model. The load-bearing additions are the observable traces of an otherwise opaque model: the prompt wording and task order, chain-of-thought reasoning, fine-tuning and in-context-learning examples, temperature and API version, and the results of consistency and prompt-injection tests.
What would settle it
Repeatedly score the same response set with a temperature-0, version-locked generative model across several weeks and attempt standardized prompt-injection attacks; perfectly stable scores with no successful injections would weaken the paper's claim that generative AI introduces distinctive consistency and gaming hazards that require extra evidence.
Extended reading notes
Core claim
The central discovery is that the traditional validity chain for automated scoring does not hold for generative AI: an LLM does not predict human ratings through designed, inspectable features, so scores lack the built-in link from rubric to construct that feature-based engines have. The paper therefore expands the standard validity framework with a distinct generative-AI column, adding five categories of evidence: documentation of the LLM choice and rationale; a complete record of prompting and in-context-learning strategy; a review of chain-of-thought output against expert annotations; reproducibility and consistency checks across time, temperature settings, and model or API versions; and fairness checks that cover pretraining data and prompt-injection attempts. The demonstrative study with GPT-4 on 1,581 responses shows that standard evaluation tools—quadratic weighted kappa, partial correlations, and contributory composites—transfer to generative AI, but the evidence collected was not sufficient to justify operational use: GPT-4 underperformed e-rater on GRE and Praxis agreement, and only a small TOEFL subsample showed that combining GPT-4 and e-rater could reach or exceed human-human agreement.
Load-bearing premise
Human ratings are treated as the gold standard for what a score should mean, and a second human rating is the benchmark for composite score quality; if human raters systematically overlook construct-relevant aspects that an LLM captures, the proposed validity evidence would miss those aspects.
Editorial extensions
If this is right
- Testing programs adopting generative AI scoring will need to collect and document prompt, fine-tuning, and chain-of-thought evidence in addition to the concordance and fairness evidence already required for feature-based engines.
- Concordance with high-quality human ratings remains the primary evidence, so LLM scores should be validated against expert or operational human ratings before they are reported.
- Combining scores from multiple AI engines and, where needed, human raters can improve construct coverage and reliability; the paper shows one small sample where two AI scores together match or exceed human-human agreement.
- Off-the-shelf LLM scoring may be less cost-effective than it appears, because assembling the required validity evidence involves expert reviews, annotation studies, consistency monitoring, and fairness analyses.
Reading between the lines
- I infer that the transparency gap may close as models become version-locked, self-hosted, and deterministic, but the paper sets no threshold at which generative-AI evidence requirements shrink to the feature-based level.
- The sample where two AI scores beat human-human agreement hints at a future in which human raters audit LLM outputs instead of producing the primary scores, a step the paper explicitly declines to endorse.
- A testable extension is to apply the same evidence-curation table to an open-source, fine-tuned LLM with known weights and training data; this would isolate how much of the added evidence burden is caused by opacity rather than by generative architecture.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that constructed-response scoring systems based on generative AI require a broader and different validity-evidence base than human-rater or feature-based NLP scoring. Drawing on the 2014 Standards (AERA/APA/NCME) and prior ETS best practices, it proposes a validity-evidence framework organized by evidence type and illustrates it with an exploratory study in which GPT-4 scored responses from GRE, Praxis, and TOEFL writing tasks. The paper also discusses contributory scoring, in which multiple AI scores are combined, possibly without human ratings, and suggests that such combinations may cover more of the construct. The authors candidly identify missing evidence—fairness, reproducibility, expert annotations—and caution that their results are not sufficient for operational use.
Significance. The paper's principal contribution is a practical, standards-anchored checklist for validity evidence when LLM-based scoring is contemplated. The distinction among construct-feature models, general linguistic/embedding models, and prompted generative models, with corresponding transparency implications, is useful and should influence assessment practice. The demonstration in Table 6 of how to organize and identify gaps in a validity argument is a strength. However, the empirical part is a small exploratory illustration, and the abstract's statement about construct coverage in the absence of human ratings goes beyond what the data can support. The paper is nevertheless valuable for its framework and for specifying the missing evidence, rather than for establishing that GPT-4 scoring is valid.
major comments (2)
- [Abstract and Section 'Combining AI Scores for Reported Score Computation' (Table 5)] The abstract claims that a contributory scoring approach combining multiple AI scores 'will cover more of the construct in the absence of human ratings,' but the analysis does not test construct coverage beyond human ratings. The criterion throughout Table 5 is H2, a second human rating; all composite-score correlations are correlations with a human judgment, not with an external construct measure. Moreover, the observed gains are inconsistent: pooled rm(E+G),H2 = .76 versus rEH2 = .75; for Praxis the AI-only composite (.75) is worse than e-rater alone (.80); for GRE it is .87 versus rH1H2 = .90; and the only clear improvement is TOEFL (.84 versus .68, N=123). The authors' own caveat that 'these results do not demonstrate that the reasons the scores agree so strongly are appropriate' directly undermines the abstract's wording. Please soften the claim to a hypothesis or provide criterion evidence independent of human ratings.
- [Section 'Demonstrative Study Using GPT4 for Scoring' (Tables 2, 4, 5, 6)] The sample-size accounting is internally inconsistent. The text states 'In total, we scored N=1,581 responses to 14 items,' but Table 2 sums to 1,172 (569+357+246); Table 4 has N=1,261; and Table 5 has N=667. The paper never explains which responses are excluded at each stage (e.g., missing e-rater scores, missing second human rating). Please reconcile the totals or state the missing-data rules. Additionally, Table 6 reports QWK ranges and e-rater/GPT4 QWKs that do not match Table 2: Table 6 says GPT4-human QWKs ranged .60-.77, but Table 2 shows .55, .67, and .76; Table 6 lists e-rater vs. GPT4 QWKs as .57, .60, and .74 for TOEFL, Praxis, and GRE, while Table 2 reports .54, .62, and .76. These discrepancies need to be corrected before the demonstration can be credited.
minor comments (5)
- [Section 'Validity Evidence for CR Scoring Systems Using Generative AI'] 'Hoffman at al. (2018, 2023)' should read 'Hoffman et al.'
- [References] The reference list contains duplicated and inconsistently authored entries for Liu et al. (2023): 'Liu, Z., Xu, P., Liu, F., & Song, H.' and 'Liu, F., Xu, P., Li, Z., Feng, Y., & Song, H.' appear to be the same work; please unify.
- [Appendix: Prompt for Feedback and Score] The prompt asks for output in JSON format, but the example shows an unquoted 'score' field; use valid JSON notation ('"score"') to avoid ambiguity in the prompt.
- [Figure 2] The footnote says the TOEFL sample was very small, but the figure itself gives no sample sizes; please state the per-test Ns in the caption or text.
- [Table 2] The column 'No. Raters' shows 10 for all three programs; please clarify whether the same 10 raters scored all items or this is a placeholder, since the text elsewhere describes different scoring operations.
Circularity Check
No significant circularity: the paper's normative validity framework and empirical concordance analyses are self-contained; self-citations are present but not load-bearing in a circular way.
full rationale
The paper does not derive a quantitative result from fitted parameters and then relabel it as a prediction. Its central claim is a normative validity-argument framework: generative AI scoring requires more extensive validity evidence because LLMs are less transparent and less deterministic than feature-based engines. That claim is supported by a logical contrast between scoring approaches and by reference to external standards (AERA/APA/NCME 2014, Williamson et al. 2012), not by an equation that reduces to its own inputs. The empirical demonstration uses standard psychometric statistics (QWK, semi-partial correlations, correlations of composite scores with a second human rating) computed directly from observed scores; no parameter is estimated from a subset of data and then used to predict that same subset. The paper's self-citations (McCaffrey et al. 2022 Best Practices, Johnson et al. 2022 fairness methods, McCaffrey et al. 2024 PRMSE) provide background methodology and are externally published; they are not invoked as an unverified uniqueness theorem or as the sole justification for the central claim. The abstract's statement that a contributory combination of AI scores 'will cover more of the construct in the absence of human ratings' is not actually established by the data, since all composite-score evaluations use a second human rating (H2) as the criterion. That is an evidentiary overreach or correctness concern, not a circular reduction: the claim is not used as an input to the analysis that supposedly supports it. Accordingly, no specific circular step can be quoted, and the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Human ratings are treated as the gold standard benchmark for evaluating AI scores.
- domain assumption The 2014 Standards validity framework is the appropriate organizing structure for generative AI scores.
- domain assumption The sample responses are representative enough for the illustrative validity evidence demonstration.
Cite this review
Pith. "Pith review of Validity Arguments For Constructed Response Scoring Using Generative Artificial Intelligence Applications." pith.science (2026). https://pith.science/paper/JYEWXFTI
@misc{pith2026250102334,
author = {Pith},
title = {Pith review of: Validity Arguments For Constructed Response Scoring Using Generative Artificial Intelligence Applications},
year = {2026},
howpublished = {\url{https://pith.science/paper/JYEWXFTI}},
note = {Machine review of arXiv:2501.02334}
}
read the original abstract
The rapid advancements in large language models and generative artificial intelligence (AI) capabilities are making their broad application in the high-stakes testing context more likely. Use of generative AI in the scoring of constructed responses is particularly appealing because it reduces the effort required for handcrafting features in traditional AI scoring and might even outperform those methods. The purpose of this paper is to highlight the differences in the feature-based and generative AI applications in constructed response scoring systems and propose a set of best practices for the collection of validity evidence to support the use and interpretation of constructed response scores from scoring systems using generative AI. We compare the validity evidence needed in scoring systems using human ratings, feature-based natural language processing AI scoring engines, and generative AI. The evidence needed in the generative AI context is more extensive than in the feature-based NLP scoring context because of the lack of transparency and other concerns unique to generative AI such as consistency. Constructed response score data from standardized tests demonstrate the collection of validity evidence for different types of scoring systems and highlights the numerous complexities and considerations when making a validity argument for these scores. In addition, we discuss how the evaluation of AI scores might include a consideration of how a contributory scoring approach combining multiple AI scores (from different sources) will cover more of the construct in the absence of human ratings.
Forward citations
Cited by 1 Pith paper
-
Using LLMs to identify features of personal and professional skills in an open-response situational judgment test
Zero-shot LLMs identify several construct-relevant features in open-response situational judgment test answers with moderate agreement with human raters, below human-level reliability for most features.
Reference graph
Works this paper leans on
-
[1]
American Educational Research Association (AERA), American Psychological Association (APA), & National Council on Measurement in Education (NCME). (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.testingstandards.net/open-access-files.html Atari, M., Xue, M. J., Park, P. S., Blasi, D., & H...
work page 2014
-
[2]
The Rise of Artificial Intelligence in Educational Measurement: Opportunities and Ethical Challenges
The Journal of Technology, Learning and Assessment, 4(3). Bennett, R. E., & Zhang, M. (2015). Validity and automated scoring. In Technology and testing (pp. 142-173). Routledge. Bommasani, R., Liang, P., & Lee, T. (2023). Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525(1), 140-146. Breyer, F. J., Rupp, A. A., & Bri...
work page Pith review arXiv 2015
-
[5]
ETS Research Report Series, 2008(2), i-
work page 2008
-
[102]
Yao, L., Haberman, S. J., & Zhang, M. (2019a). Penalized best linear prediction of true test scores. Psychometrika, 84(1), 186-211. Yao, L., Haberman, S. J., & Zhang, M. (2019b). Prediction of writing true scores in automated scoring of essays by best linear predictors and penalized best linear predictors. ETS Research Report Series, 2019(1), 1-27. Zhang,...
work page 2019
-
[142]
Kwon, S. McCaffrey, D. F., Jewsbury, P. & Casabianca, J. M. (n.d.). Empirical Bayes estimation for evaluating subgroup biases in automated scoring. White paper. Leacock, C., & Chodorow, M. (2003). C-rater: Automated scoring of short-answer questions. Computers and the Humanities, 37, 389-405. Lee, G. G., Latif, E., Wu, X., Liu, N., & Zhai, X. (2024). Appl...
arXiv 2003
-
[258]
Tan, X., Kim, S., Paek, I., & Xiang, B. (2009). An alternative to the trend scoring method for adjusting scoring shifts in mixed-format tests. In Annual Meeting of the National Council on Measurement in Education, San Diego, CA. Retrieved from http://www. etsliteracy. org/Media/Conferences_and_Events/AERA_2009_pdfs/AERA_NCME_2009_Tan. pdf. Wei, J., Wang, ...
work page 2009
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.