{"id":"4ace9d70-5eac-45f1-9ac9-8b383b36c7d5","arxiv_id":"2504.21330","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4o essay scoring shows larger errors for non-native English writers when the model correctly infers that they are non-native, but no meaningful gender effect.","lead":"This paper tested whether ChatGPT-style AI models can guess a student's gender or language background from an essay, and whether that guess makes scoring less fair. For non-native English writers, the results suggest scoring errors grow when the AI correctly identifies them as non-native.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Correctness is a confounded proxy: non-native essays labeled \"Correct\" are selected on the same surface cues that drive scoring errors, so the overall interaction may not reflect demographic bias.","rationale":"The reader's weakest assumption is that the Correct/Unreliable split is an independent measure of demographic recognition. My reading supports that and locates the exact mechanism: since coverage is nearly complete and \"Unreliable\" is defined by misclassification, the split is a property of the essay under evaluation. Non-native essays that the model labels as non-native are systematically different from those it labels as native—they contain stronger surface-level non-native cues. The scoring outcome and the demographic prediction are both functions of the same text, so without controlling for those cues, the Correctness*Language interaction cannot separate \"the model recognized the student was non-native\" from \"the essay has features that cause both correct classification and larger scoring error.\" I would phrase the paper's actual finding as an associational pattern conditional on the model's own predictions, not as evidence that demographic knowledge causes scoring bias. The concrete re-analysis with essay-set fixed effects and surface controls would settle whether the interaction survives. I agree with CONDITIONAL; I would not escalate to REJECT because the paper is transparent about its per-topic regressions and the descriptive pattern may still hold once the confound is checked. I would not lower the verdict because the claim as stated is currently over-interpreted.","tokens_in":12623,"tokens_out":7693,"duration_ms":89199,"concrete_test":"Re-estimate the weighted regression in Table 4 on the pooled data with essay-set fixed effects and a fixed, pre-specified set of surface-level controls (essay length, grammatical-error rate, lexical diversity, and syntactic complexity), using the same inverse-probability weights. If the Correctness*Language coefficient moves toward zero or loses significance, the association is explained by essay features; if it remains near 0.5 with a tight confidence interval, the demographic-recognition interpretation is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract and Section 4.2 is causal in spirit: scoring error for non-native students increases when GPT-4o correctly identifies them as non-native (Overall Correctness*Language = 0.502, SE 0.074). This requires Correctness to be as-if randomly assigned conditional on covariates, but it is a deterministic function of the essay text. For first-language background, coverage is near 100% (Table 2), so \"Unreliable\" is essentially the set of essays the model misclassifies as native. A non-native essay will be misclassified exactly when its linguistic surface is close to native-like writing; that same closeness is independently related to scoring error. The Correct group is therefore selected on strong non-native cues (grammatical deviations, lexical and syntactic patterns). The regression in Table 4 includes only Language, Correctness, and their interaction; it does not adjust for any essay-level linguistic covariate or essay-set fixed effects in the overall model. If Correctness is a proxy for prototypical non-native language rather than for the model's \"knowledge\" of demographics, the positive interaction is explainable without any bias mechanism. Topic-level results also weaken the generalization: Essay Set 2 has a significant negative interaction (-0.490, SE 0.203), Sets 1 and 4 are non-significant, and the overall estimate is largely driven by Set 6. Without a design that separates demographic recognition from the essay features that cause it, the paper's main causal statement is not identified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript investigates whether a prompt-based large language model (GPT-4o) can infer students' gender and first-language background from their essays, and whether the accuracy of those demographic inferences is associated with bias in automated essay scoring. Using six essay sets from the PERSUADE 2.0 corpus, the authors prompt GPT-4o to predict demographics and to score essays, then compare fairness metrics and fit inverse-probability-weighted regressions with demographic group, correctness of demographic prediction, and their interaction. The main reported findings are that first-language background can be inferred with roughly 79–87% accuracy on covered cases, and that the regression interaction Correctness×Language is positive overall (coefficient 0.502, SE 0.074), interpreted as larger scoring error for non-native English speakers when the model correctly identifies them as non-native. For gender, coverage is very low (4–13%) and no consistent relationship is found.","tokens_in":12828,"tokens_out":5459,"duration_ms":57028,"significance":"If the central claim held, the study would provide an actionable connection between demographic inferability and scoring bias in prompt-based AES, going beyond prior work on fine-tuned models and suggesting concrete debiasing levers for widely used LLM APIs. The paper has notable strengths: a large public corpus, a transparent prompt design, robustness checks via five paraphrased prompts, and clearly specified regression models. The demographic inference results for RQ1 are a solid empirical contribution, especially the finding that first-language background is more detectable than gender. However, the load-bearing claim (iii) is an interpretation of a non-experimental comparison between 'Correct' and 'Unreliable' groups, and the manuscript does not yet rule out essay-level linguistic confounds. The significance of the paper is therefore conditional on either additional covariate-adjusted analysis or a more modest, descriptive framing of the main finding.","major_comments":[{"comment":"The Correct/Unreliable dichotomy does not isolate demographic recognition. As defined, 'Unreliable' merges incorrect predictions and 'Uncertain'; for first-language background the coverage is 97–99% (Table 2), so 'Unreliable' is almost entirely essays the model misclassified. A non-native essay is misclassified as native precisely when its linguistic surface resembles native writing, and that same surface property is likely associated with how the scoring model errs, independent of any demographic 'awareness'. The regression in Table 4 adjusts for Language, Correctness, and their interaction but includes no essay-level covariates (e.g., lexical sophistication, grammatical error density, topic) in the overall model. The positive Correctness×Language coefficient may therefore reflect selection on surface properties rather than a demographic-recognition mechanism. Please add covariates or explicitly report the analysis as a descriptive association, and revise the causal phrasing in the abstract and Section 5 accordingly.","section":"§3.3 and Table 4"},{"comment":"The interaction is not consistent across essay sets. Essay Set 2 shows a significant negative interaction (-0.490, SE 0.203), Sets 1 and 4 are non-significant, and the positive overall estimate appears driven largely by Set 6, the largest set. The abstract's unconditional claim (iii) is not supported by the per-set results. A heterogeneity test, a mixed-effects model with essay-set random slopes, or a meta-analytic summary would clarify whether a general phenomenon exists or whether the effect is topic-specific. Without such an analysis, the overall interaction should be interpreted with caution.","section":"§4.2, Table 4"},{"comment":"The paper does not correct for multiple testing across six essay sets, two demographic attributes, four fairness metrics, and a large family of regression coefficients. For example, the gender interaction terms in Sets 3 and 5, and some of the OSA/CSD entries in Table 3, could be chance findings. Please report adjusted p-values (e.g., Benjamini-Hochberg) or at minimum disclose the total number of tests and argue why the pattern of significance is robust. This is especially important because the main claim rests on a single interaction coefficient in the overall model.","section":"§4.2, Tables 3 and 4"}],"minor_comments":[{"comment":"Please clarify how majority voting was applied for demographic predictions when the five paraphrased prompts produced a tie, and whether the same repeated prompts were used for scoring (which was averaged). Reporting agreement among the five runs would also strengthen the robustness claim.","section":"§3.5"},{"comment":"The inverse probability weighting is not fully specified. State how the weights were computed (e.g., inverse group frequency within each essay set) and indicate whether weighting was applied to the QWK and MAED calculations or only to the regressions and fairness regression-based metrics.","section":"§3.3"},{"comment":"The text says the model achieved 'an accuracy of about 0.79–0.86' for first-language background, but Table 2 shows accuracy values from 0.753 to 0.868. The reported range is inconsistent with the table; please correct the text to match the data.","section":"§4.1"},{"comment":"The sentence 'Each experiment was run five times with paraphrased prompts by other prompt-based LLMs (i.e., Gemini and Claude)' is ambiguous. It is unclear whether the paraphrases were generated by Gemini/Claude or whether the experiments themselves were executed with those models. Please rephrase.","section":"§3.5"},{"comment":"The table would benefit from a note explaining the direction of MAED and OSD (which group is over- or under-scored) and why a negative value indicates larger error for a particular group. As written, the interpretation of negative versus positive values across rows is not transparent to a reader unfamiliar with the metrics.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and practically important question, and the demographic inference results are publishable in principle. The main barrier is the interpretation of the Correctness×Language interaction as evidence of a demographic-recognition mechanism when the comparison groups are formed by the same essay features that plausibly drive scoring errors. I would support acceptance after the authors either add covariate-adjusted analysis (e.g., essay-level linguistic features, topic fixed effects) or substantially soften the causal claim. The per-set heterogeneity in the interaction should also be addressed. The paper's heavy reliance on the authors' prior framework is not itself a problem, but the framing that findings 'align' with prior fine-tuning work should be tempered given the confound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first paper I know that tests the relation between an LLM's ability to infer demographics and scoring bias in the prompt-based paradigm, and it's a competently run study of GPT-4o on PERSUADE. The main regression is well specified, and the positive Correctness*Language interaction (0.502, SE 0.074) is real in the pooled model, with similar positive interactions in essay sets 3, 5, and 6. That is worth taking seriously.\n\nCredit where due: it moves beyond fine-tuned embeddings (Kwako and Ormerod), uses separate sessions for demographic inference and scoring, reports several fairness metrics, and is transparent about limitations. The paper also honestly notes that only one model and two demographic attributes were tested.\n\nThe soft spot is the identification assumption. 'Correct' is not randomly assigned; it is a deterministic function of essay text. For language background, coverage is about 97-99%, so 'Unreliable' is largely the set of non-native essays the model misclassified as native. A non-native essay gets misclassified exactly when its linguistic surface looks native-like, and that same native-likeness is independently related to scoring error. The regression controls for Language, Correctness, and their interaction, but not for any essay-level linguistic covariate or essay-set fixed effect in the pooled model. So the positive interaction is compatible with correctness being a proxy for prototypical non-native writing rather than for the model's 'knowledge' of demographics. The topic-level pattern reinforces this worry: Set 2 is significantly negative, Sets 1 and 4 are null, and the pooled estimate leans heavily on Set 6.\n\nSecondary issues: 'Unreliable' merges wrong predictions with 'Uncertain' (mostly irrelevant for language, but important for gender), there is no correction for multiple comparisons across metrics and essay sets, and no code or data is provided. These are manageable in revision.\n\nIs the paper still useful? Yes. As a descriptive finding—GPT-4o can infer L1 from essays, and its scoring errors correlate with that inference—it is a legitimate empirical observation. The causal phrasing in the abstract ('increases when the LLM correctly identifies') goes beyond what the design supports. A serious revision should either add linguistic controls, justify correctness as as-if random, or soften the claim.\n\nI would send this to peer review rather than desk reject. It deserves a serious referee, and the central question is timely for AES and educational NLP.","headline":"First real test of the demographic-inference/scoring-bias link in prompt-based AES, with a solid pooled interaction for first-language background, but the central causal claim is not identified because 'Correct' is a proxy for essay surface features.","tokens_in":13423,"tokens_out":1947,"would_cite":true,"duration_ms":21427,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4o can infer first-language background from essays and scores non-native writers less fairly when it does","keywords":["automated essay scoring","prompt-based large language model","demographic inference","first-language background","algorithmic fairness","GPT-4o","educational NLP","scoring bias"],"falsifier":"Rerun the weighted regression with an added control for surface grammatical error rate, or score human-corrected transcripts of the non-native essays; if the correctness-by-language coefficient drops to zero, the apparent demographic bias is actually error-pattern bias and the paper's conclusion fails.","tokens_in":12376,"feed_emoji":"📝","tokens_out":11724,"duration_ms":114436,"temperature":0.7,"pith_summary":"This paper asks whether a prompt-based large language model, GPT-4o, can read students' demographics from their essays and whether that recognition changes the fairness of automated scoring. On six independent writing topics from an argumentative-essay corpus, the model inferred first-language background with accuracy around $0.75$–$0.87$ when it gave a definite answer, and its scoring error for non-native English speakers increased when it correctly identified them as non-native. The stakes are concrete because prompt-based essay scoring is accessible to educators with no machine-learning expertise, so a hidden demographic penalty could enter real classrooms unnoticed. The paper also finds little gender bias in this setup, even though the model's gender guesses were accurate on the small share of essays where it committed to a prediction.","feed_headline":"GPT-4o spots non-native writers, then scores them worse","feed_subtitle":"In six essay topics, non-native writers are scored less accurately once GPT-4o recognizes their background.","key_machinery":"The central object is the pair of independent GPT-4o prompts: a demographic-inference prompt (gender or first-language background, each with an 'Uncertain' option) and a rubric-based scoring prompt with few-shot examples and chain-of-thought steps. The argument is carried by splitting demographic predictions into 'Correct' versus 'Unreliable' (wrong or uncertain) and then fitting a weighted multivariate regression of scoring error on demographics, correctness, and the correctness-by-language interaction, with inverse probability weighting to keep the minority group represented. The positive, significant interaction coefficient is the load-bearing quantity: it is what connects successful demographic recognition to larger scoring errors for non-native writers.","core_discovery":"On the paper's own terms, the discovery is that GPT-4o's scoring fairness depends on whether the model can recognize the writer's first-language background. For essays it classified as native or non-native, the model reached near-total coverage and roughly $0.75$–$0.87$ accuracy, and the weighted regression of absolute scoring error on demographics, prediction correctness, and their interaction produced a positive and significant correctness-by-language coefficient of $0.502$ (SE $0.074$) overall. That coefficient means non-native English speakers are scored with bigger errors exactly when the model has correctly labeled them as non-native. Fairness metrics (OSA, OSD, CSD, MAED) show more pronounced and more often statistically significant bias in the correctly-predicted group than in the unreliable group. The same relationship does not appear for gender, which the paper attributes to gender information being less directly accessible through prompts and to the model's low commitment rate on gender judgments.","pith_inferences":["A direct causal reading of the interaction implies that stripping away surface error patterns, for example by editing non-native essays into standard English before scoring, should reduce the bias; the paper does not run that test.","The gender null result may be an artifact of low coverage: with only 4–13 percent of essays receiving a definite gender prediction, the analysis excludes most of the data, so a forced-choice prompt could reveal gender bias that this design misses.","The demographic-inference prompt itself could serve as an audit tool, letting a school or platform measure how accurately any LLM guesses student attributes and use the correctness-by-language coefficient as a pre-deployment fairness check.","If the bias mechanism is demographic recognition rather than essay quality, then telling the model the author is non-native, or removing such cues, should shift the interaction term and give a cheap experimental handle for future studies."],"forward_implications":["If the result holds, prompt-based essay scoring with GPT-4o will systematically disadvantage non-native English writers whenever the essay text reveals their language background.","Because first-language coverage is near 100 percent, this penalty applies to almost all non-native writers in the analyzed essays, not to a small subset.","Debiasing strategies developed for fine-tuned scoring models do not transfer directly to prompt-only tools, so fair use requires new mitigation such as deliberately diverse few-shot examples.","The absence of consistent gender bias shows that demographic fairness must be checked attribute by attribute; a model can be fair on gender while biased on language background, or vice versa."],"supporting_citations":[{"why":"supplies the PERSUADE 2.0 argumentative-essay corpus with demographic labels and human holistic scores, the data for all inference and scoring experiments.","marker":"[6]"},{"why":"established that fine-tuned LLM representations can predict demographics and are linked to scoring bias, the finding this paper extends to prompt-based models.","marker":"[13]"},{"why":"showed that demographic information embedded in LLM text embeddings relates to predictive bias in educational text classification, motivating the research questions.","marker":"[25]"},{"why":"provided effective prompt design strategies for GPT-based essay scoring that the paper's scoring prompts adapt.","marker":"[31]"},{"why":"supplied comparative evidence on prompt-based essay scoring accuracy and prompt-instruction variants used in the design.","marker":"[26]"},{"why":"contributed the fairness metrics (OSA, OSD, CSD) and earlier AES bias findings that the evaluation builds on.","marker":"[34]"},{"why":"is the prior prompt-based GPT scoring fairness study that found little gender or L1 bias, the result this paper's interaction analysis qualifies.","marker":"[33]"},{"why":"defined the multi-dimensional fairness evaluation framework for educational applications that underlies the bias metric selection.","marker":"[18]"}],"fun_headline_variants":["GPT-4o's correct non-native ID worsens essay scoring bias","When GPT-4o spots non-native writers, scoring bias increases","LLM essay bias: recognizing non-native status leads to worse scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 'Correct' versus 'Unreliable' split cleanly measures the model's demographic recognition; if being 'Correct' actually tracks essay traits such as grammatical error patterns or topic, the positive correctness-by-language interaction in Section 4.2 could reflect those traits rather than demographic bias, and the central claim would not follow.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o's correct non-native ID worsens essay scoring bias","When GPT-4o spots non-native writers, scoring bias increases","LLM essay bias: recognizing non-native status leads to worse scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000527,"raw_usage":{"total_tokens":2594,"prompt_tokens":1047,"completion_tokens":1547,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":663,"completion_tokens_details":{"reasoning_tokens":1487}},"tokens_in":663,"tokens_out":1547,"duration_ms":11532,"temperature":1.0,"reasoning_tokens":1487,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:05:37.788910+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the weighted regression with an added control for surface grammatical error rate, or score human-corrected transcripts of the non-native essays; if the correctness-by-language coefficient drops to zero, the apparent demographic bias is actually error-pattern bias and the paper's conclusion fails.","supporting_citations":[{"cited_title":"Assessing Writing 61, 100865 (2024)","cited_arxiv_id":null,"evidence_quote":"supplies the PERSUADE 2.0 argumentative-essay corpus with demographic labels and human holistic scores, the data for all inference and scoring experiments."},{"cited_title":"In: Proceedings of t he 19th Workshop on Innovative Use of NLP for Building Educational Application s (BEA 2024)","cited_arxiv_id":null,"evidence_quote":"established that fine-tuned LLM representations can predict demographics and are linked to scoring bias, the finding this paper extends to prompt-based models."},{"cited_title":"In : International Conference on Computational Linguistics 2022","cited_arxiv_id":null,"evidence_quote":"showed that demographic information embedded in LLM text embeddings relates to predictive bias in educational text classification, motivating the research questions."},{"cited_title":"In: Proceedings of the 15th International Learning Analytics and Knowledge Conference","cited_arxiv_id":null,"evidence_quote":"provided effective prompt design strategies for GPT-based essay scoring that the paper's scoring prompts adapt."},{"cited_title":"In: Proceedings of the AAA I Conference on Artiﬁcial Intelligence","cited_arxiv_id":null,"evidence_quote":"contributed the fairness metrics (OSA, OSD, CSD) and earlier AES bias findings that the evaluation builds on."},{"cited_title":"In: Proceedings of the 18th Workshop o n Innovative Use of NLP for Building Educational Applications (BEA 2023)","cited_arxiv_id":null,"evidence_quote":"is the prior prompt-based GPT scoring fairness study that found little gender or L1 bias, the result this paper's interaction analysis qualifies."},{"cited_title":"In: Proceedings of the fo urteenth workshop on innovative use of NLP for building educational application s","cited_arxiv_id":null,"evidence_quote":"defined the multi-dimensional fairness evaluation framework for educational applications that underlies the bias metric selection."}],"review_version":1}