{"id":"d5779873-2f92-4749-abed-5cbe56e0b6e8","arxiv_id":"2501.02334","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Generative AI scoring needs extra validity evidence for transparency and consistency, and a small demonstration shows GPT-4 trails e-rater in matching human ratings, though a combined score helps.","lead":"This paper proposes a set of best practices for proving that generative AI essay scoring is valid enough for high-stakes tests, and illustrates them on three standardized exams. It finds that GPT-4 matched human scores less consistently than the older e-rater engine, but combining the two AI scorers can rival human-human agreement in some cases.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's contributory-scoring claim is tested only against a second human rating; the pooled gain is negligible and the pattern is inconsistent across tests.","rationale":"The reader's weakest assumption is that human ratings are treated as the gold standard. My concern is related but more specific: the paper not only relies on human ratings as the benchmark, it also claims that combining AI scores can cover more of the construct in the absence of human ratings—a claim that the human-ratings benchmark cannot support. This is an internal tension rather than merely an external assumption. The central framework—that generative AI scoring needs additional validity evidence (prompt documentation, reproducibility, chain-of-thought review, prompt-injection handling, fairness analyses)—remains credible because it follows from the LLM's opacity and stochasticity, and the paper is appropriately cautious in its operational conclusions (e.g., recommending retention of e-rater). The empirical demonstration is explicitly exploratory, and the paper flags small samples and missing evidence. Therefore the concern does not change the reader's conditional verdict; it reinforces the conditionality by identifying a specific overreach in the abstract's contributory-scoring claim. A targeted re-analysis with confidence intervals and a non-human criterion would settle whether the claim can be salvaged.","tokens_in":20800,"tokens_out":3904,"duration_ms":41526,"concrete_test":"Recompute the pooled and per-test comparisons in Table 5 with bootstrap 95% confidence intervals for the difference r_m(E+G),H2 − r_E,H2. If the pooled interval includes zero—and especially if the Praxis difference is significantly negative—the 'cover more of the construct' claim loses empirical support. Additionally, test construct coverage beyond human ratings by evaluating whether the combined AI scores improve prediction of a non-human criterion (e.g., subsequent course performance, known-group contrasts, or expert content-quality ratings independent of the rubric); without such a criterion, agreement with H2 cannot distinguish construct coverage from shared method variance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central normative claim—that generative AI scoring requires more extensive validity evidence—is reasonable and does not hinge on the empirical results. However, the abstract also asserts that a contributory scoring approach combining multiple AI scores 'will cover more of the construct in the absence of human ratings.' This is an empirical, load-bearing sub-claim, and the evidence in Table 5 does not support it. First, the criterion used throughout the composite-score analysis is H2, a second human rating; this measures concordance with human judgment, not construct coverage beyond human ratings. The phrase 'in the absence of human ratings' is therefore not actually tested. Second, the gains from adding GPT4 to e-rater are tiny or negative: pooled r_m(E+G),H2 = .76 vs r_E,H2 = .75; for Praxis the composite is worse (.75) than e-rater alone (.80); only TOEFL shows a meaningful improvement (.84 vs .68, but N=123). Table 4's semi-partial correlations show GPT4 contributes some unique variance in predicting H1, but that is still human-referenced variance. The paper itself cautions (Section 'Combining AI Scores') that 'these results do not demonstrate that the reasons the scores agree so strongly are appropriate.' Thus the abstract overstates the empirical basis for construct coverage beyond human ratings, and the claim should be treated as speculative rather than demonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that constructed-response scoring systems based on generative AI require a broader and different validity-evidence base than human-rater or feature-based NLP scoring. Drawing on the 2014 Standards (AERA/APA/NCME) and prior ETS best practices, it proposes a validity-evidence framework organized by evidence type and illustrates it with an exploratory study in which GPT-4 scored responses from GRE, Praxis, and TOEFL writing tasks. The paper also discusses contributory scoring, in which multiple AI scores are combined, possibly without human ratings, and suggests that such combinations may cover more of the construct. The authors candidly identify missing evidence—fairness, reproducibility, expert annotations—and caution that their results are not sufficient for operational use.","tokens_in":21082,"tokens_out":7270,"duration_ms":68502,"significance":"The paper's principal contribution is a practical, standards-anchored checklist for validity evidence when LLM-based scoring is contemplated. The distinction among construct-feature models, general linguistic/embedding models, and prompted generative models, with corresponding transparency implications, is useful and should influence assessment practice. The demonstration in Table 6 of how to organize and identify gaps in a validity argument is a strength. However, the empirical part is a small exploratory illustration, and the abstract's statement about construct coverage in the absence of human ratings goes beyond what the data can support. The paper is nevertheless valuable for its framework and for specifying the missing evidence, rather than for establishing that GPT-4 scoring is valid.","major_comments":[{"comment":"The abstract claims that a contributory scoring approach combining multiple AI scores 'will cover more of the construct in the absence of human ratings,' but the analysis does not test construct coverage beyond human ratings. The criterion throughout Table 5 is H2, a second human rating; all composite-score correlations are correlations with a human judgment, not with an external construct measure. Moreover, the observed gains are inconsistent: pooled rm(E+G),H2 = .76 versus rEH2 = .75; for Praxis the AI-only composite (.75) is worse than e-rater alone (.80); for GRE it is .87 versus rH1H2 = .90; and the only clear improvement is TOEFL (.84 versus .68, N=123). The authors' own caveat that 'these results do not demonstrate that the reasons the scores agree so strongly are appropriate' directly undermines the abstract's wording. Please soften the claim to a hypothesis or provide criterion evidence independent of human ratings.","section":"Abstract and Section 'Combining AI Scores for Reported Score Computation' (Table 5)"},{"comment":"The sample-size accounting is internally inconsistent. The text states 'In total, we scored N=1,581 responses to 14 items,' but Table 2 sums to 1,172 (569+357+246); Table 4 has N=1,261; and Table 5 has N=667. The paper never explains which responses are excluded at each stage (e.g., missing e-rater scores, missing second human rating). Please reconcile the totals or state the missing-data rules. Additionally, Table 6 reports QWK ranges and e-rater/GPT4 QWKs that do not match Table 2: Table 6 says GPT4-human QWKs ranged .60-.77, but Table 2 shows .55, .67, and .76; Table 6 lists e-rater vs. GPT4 QWKs as .57, .60, and .74 for TOEFL, Praxis, and GRE, while Table 2 reports .54, .62, and .76. These discrepancies need to be corrected before the demonstration can be credited.","section":"Section 'Demonstrative Study Using GPT4 for Scoring' (Tables 2, 4, 5, 6)"}],"minor_comments":[{"comment":"'Hoffman at al. (2018, 2023)' should read 'Hoffman et al.'","section":"Section 'Validity Evidence for CR Scoring Systems Using Generative AI'"},{"comment":"The reference list contains duplicated and inconsistently authored entries for Liu et al. (2023): 'Liu, Z., Xu, P., Liu, F., & Song, H.' and 'Liu, F., Xu, P., Li, Z., Feng, Y., & Song, H.' appear to be the same work; please unify.","section":"References"},{"comment":"The prompt asks for output in JSON format, but the example shows an unquoted 'score' field; use valid JSON notation ('\"score\"') to avoid ambiguity in the prompt.","section":"Appendix: Prompt for Feedback and Score"},{"comment":"The footnote says the TOEFL sample was very small, but the figure itself gives no sample sizes; please state the per-test Ns in the caption or text.","section":"Figure 2"},{"comment":"The column 'No. Raters' shows 10 for all three programs; please clarify whether the same 10 raters scored all items or this is a placeholder, since the text elsewhere describes different scoring operations.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"This is a useful standards/validity paper, but it is not primarily a computational-linguistics contribution; it is a psychometric and assessment-practice paper. If the journal's scope is strictly computational linguistics, the fit may be limited, though for assessment-focused venues it is appropriate. The empirical inconsistencies described in the major comments are fixable within the scope of a revision, but the abstract's overclaim should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the framework is the real contribution, and the empirical analysis is an illustration, not proof of construct coverage beyond human ratings. The abstract's combined-score claim goes further than the data.\n\nWhat's new: the paper extends the Standards-based validity framework to generative AI scoring in a practical way. The Table 1 comparison across human, feature-based AI, and generative AI evidence types is genuinely useful, and the new evidence categories—prompt documentation, consistency/reliability over time, chain-of-thought review, prompt-injection handling—are a real extension beyond the prior feature-based NLP validity work. The paper is also honest about what it didn't collect: no fairness subgroup analyses, no reproducibility experiments, no expert annotations beyond a small chain-of-thought check. That honesty helps the reader see the framework in action rather than as a marketing exercise.\n\nWhere it's soft: the stress-test note lands. The abstract says a contributory scoring approach combining multiple AI scores 'will cover more of the construct in the absence of human ratings.' That's an empirical claim, and the evidence in Table 5 doesn't support it across the board. The criterion throughout is H2, a second human rating, so the comparison is still concordance with human judgment, not construct coverage beyond humans. The pooled correlation for e-rater + GPT4 is .76 vs .75 for e-rater alone—basically no gain. For Praxis the combination is worse (.75 vs .80). Only TOEFL shows a real improvement (.84 vs .68), with N=123. The paper itself includes the caveat that these results don't demonstrate the reasons for agreement are appropriate, but the abstract oversells it. That's a real but minor flaw because the central framework claim doesn't depend on that result.\n\nAlso minor: the sample sizes are small and there are no error bars, but for a methods/framework paper that's acceptable if the claims are appropriately bounded.\n\nWho it's for: psychometricians and NLP engineers who need a shared vocabulary for validating LLM scoring. It could genuinely change practice if testing organizations adopt this evidence checklist.\n\nBottom line: send it out for review. The framework is worth refereeing, and the empirical section can be revised to match the abstract's claims more carefully. This deserves a serious referee despite the soft empirical spot.","headline":"Useful extension of validity theory to generative AI scoring, but the empirical support for the abstract's combined-score claim is weak outside TOEFL.","tokens_in":21583,"tokens_out":2935,"would_cite":true,"duration_ms":26146,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Scoring essays with generative AI demands more validity evidence than traditional automated scoring.","keywords":["constructed response scoring","generative AI","large language models","validity evidence","automated scoring","human ratings","fairness","contributory scoring"],"falsifier":"Repeatedly score the same response set with a temperature-0, version-locked generative model across several weeks and attempt standardized prompt-injection attacks; perfectly stable scores with no successful injections would weaken the paper's claim that generative AI introduces distinctive consistency and gaming hazards that require extra evidence.","tokens_in":20646,"feed_emoji":"🤖","tokens_out":7887,"duration_ms":68092,"temperature":0.7,"pith_summary":"This paper argues that scoring constructed responses with generative AI models such as GPT-4 requires a broader and more demanding validity evidence base than the feature-based natural language processing engines that already automate this task. Because the large language model is opaque, probabilistic, and sensitive to prompt wording, the authors propose that testing programs document the choice of model, the exact prompt and in-context-learning examples, chain-of-thought outputs, fine-tuning data, consistency over time, and prompt-injection resistance, and that they run expanded fairness analyses. They demonstrate this evidence-curation exercise on essay responses from the GRE, Praxis, and TOEFL writing tests, where GPT-4's agreement with human raters was generally lower than that of the e-rater engine. They also show that averaging scores from two different AI engines can cover more of the construct and, in one small sample, match or exceed the correlation between two human raters. If correct, the paper gives testing programs a concrete checklist for deciding when generative AI scores can support high-stakes score interpretation.","feed_headline":"GPT-4 essay scores need extra validity checks","feed_subtitle":"Prompt logs, consistency tests, and fairness data are required before high-stakes use.","key_machinery":"The central mechanism is the validity-evidence matrix in Table 1, which takes the five evidence types from the 2014 testing Standards (internal structure, relations to external variables, response processes, test content, and consequences of use, plus fairness) and adds a column specifying what must be documented when scores come from a generative model. The load-bearing additions are the observable traces of an otherwise opaque model: the prompt wording and task order, chain-of-thought reasoning, fine-tuning and in-context-learning examples, temperature and API version, and the results of consistency and prompt-injection tests.","core_discovery":"The central discovery is that the traditional validity chain for automated scoring does not hold for generative AI: an LLM does not predict human ratings through designed, inspectable features, so scores lack the built-in link from rubric to construct that feature-based engines have. The paper therefore expands the standard validity framework with a distinct generative-AI column, adding five categories of evidence: documentation of the LLM choice and rationale; a complete record of prompting and in-context-learning strategy; a review of chain-of-thought output against expert annotations; reproducibility and consistency checks across time, temperature settings, and model or API versions; and fairness checks that cover pretraining data and prompt-injection attempts. The demonstrative study with GPT-4 on 1,581 responses shows that standard evaluation tools—quadratic weighted kappa, partial correlations, and contributory composites—transfer to generative AI, but the evidence collected was not sufficient to justify operational use: GPT-4 underperformed e-rater on GRE and Praxis agreement, and only a small TOEFL subsample showed that combining GPT-4 and e-rater could reach or exceed human-human agreement.","pith_inferences":["I infer that the transparency gap may close as models become version-locked, self-hosted, and deterministic, but the paper sets no threshold at which generative-AI evidence requirements shrink to the feature-based level.","The sample where two AI scores beat human-human agreement hints at a future in which human raters audit LLM outputs instead of producing the primary scores, a step the paper explicitly declines to endorse.","A testable extension is to apply the same evidence-curation table to an open-source, fine-tuned LLM with known weights and training data; this would isolate how much of the added evidence burden is caused by opacity rather than by generative architecture."],"forward_implications":["Testing programs adopting generative AI scoring will need to collect and document prompt, fine-tuning, and chain-of-thought evidence in addition to the concordance and fairness evidence already required for feature-based engines.","Concordance with high-quality human ratings remains the primary evidence, so LLM scores should be validated against expert or operational human ratings before they are reported.","Combining scores from multiple AI engines and, where needed, human raters can improve construct coverage and reliability; the paper shows one small sample where two AI scores together match or exceed human-human agreement.","Off-the-shelf LLM scoring may be less cost-effective than it appears, because assembling the required validity evidence involves expert reviews, annotation studies, consistency monitoring, and fairness analyses."],"supporting_citations":[{"why":"Defines the five validity-evidence types that the paper extends to generative AI.","marker":"AERA, APA, & NCME (2014)"},{"why":"Supplies the existing best-practice validity framework for CR scoring that the paper expands.","marker":"McCaffrey et al. (2022)"},{"why":"Establishes the evaluation metrics and thresholds for automated scoring agreement used in the study.","marker":"Williamson et al. (2012)"},{"why":"Describes e-rater, the feature-based scoring engine that GPT-4 is compared against.","marker":"Attali & Burstein (2006)"},{"why":"Introduces the contributory scoring approach for combining scores that the paper applies to multiple AI scores.","marker":"Breyer et al. (2017)"},{"why":"Shows chain-of-thought prompting, which the paper proposes to analyze as validity evidence.","marker":"Wei et al. (2022)"},{"why":"Provides the expanded fairness definitions the paper recommends for LLM scores.","marker":"Johnson & McCaffrey (2023)"},{"why":"Demonstrates chain-of-thought and in-context learning for essay scoring, supporting the prompting documentation requirements.","marker":"Lee et al. (2024)"}],"fun_headline_variants":["LLM essay scoring needs extra validity checks","For AI essay scores, LLM transparency gaps force extra evidence","GPT-4 underperforms e-rater on GRE and Praxis scoring","LLM scoring: more validity evidence, not less, than feature-based AI","Generative AI essay scores lack rubric-to-construct link"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Human ratings are treated as the gold standard for what a score should mean, and a second human rating is the benchmark for composite score quality; if human raters systematically overlook construct-relevant aspects that an LLM captures, the proposed validity evidence would miss those aspects.","fun_headline_variants_meta":{"raw":{"variants":["LLM essay scoring needs extra validity checks","For AI essay scores, LLM transparency gaps force extra evidence","GPT-4 underperforms e-rater on GRE and Praxis scoring","LLM scoring: more validity evidence, not less, than feature-based AI","Generative AI essay scores lack rubric-to-construct link"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1626,"prompt_tokens":991,"completion_tokens":635,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":549}},"tokens_in":607,"tokens_out":635,"duration_ms":5784,"temperature":1.0,"reasoning_tokens":549,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:13:08.595467+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeatedly score the same response set with a temperature-0, version-locked generative model across several weeks and attempt standardized prompt-injection attacks; perfectly stable scores with no successful injections would weaken the paper's claim that generative AI introduces distinctive consistency and gaming hazards that require extra evidence.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the five validity-evidence types that the paper extends to generative AI."}],"review_version":1}