{"id":"d854b021-8324-4945-9c2e-5b2b67e3df1e","arxiv_id":"2502.08450","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Using grammar-corrected essays as a second input improves cross-prompt trait scoring, with the largest gains on grammar-related traits like Conventions.","lead":"This paper trains an essay scorer to look at both the original essay and a grammar-corrected rewrite, and reports better agreement with human raters on prompts the model has never seen. The approach targets the practical setting of automated essay grading, where new prompts must be scored after training is complete.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported gains over the strongest baseline are within one standard deviation, and GAPS's overall QWK is numerically below ProTACT, so the central claim of notable cross-prompt improvement is not established.","rationale":"The reader's stated weakest assumption is GEC output quality, which the paper acknowledges in its Limitations section. That is a legitimate concern but not the most load-bearing one: even with an accurate GEC model, the reported experimental evidence for the central claim is weak. The reader's rationale does mention that the gains over the strongest prior method are small and not significance-tested, which is the concern I identify as primary. However, because the reader's formal 'weakest_assumption' field names GEC dependency, I mark partial agreement. The verdict itself remains CONDITIONAL: the paper contributes a plausible architecture and informative ablations, but the headline empirical claim requires significance testing and a clearer statement of which specific prompts/traits are improved. No change to the reader's verdict is needed, as the conditional recommendation already captures the need for stronger evidence.","tokens_in":8920,"tokens_out":7017,"duration_ms":68751,"concrete_test":"Recompute the per-prompt and per-trait differences between GAPS [TA+PA+GA] and ProTACT [TA+PA] across the five reported seeds, and run a paired permutation or bootstrap test on the prompt-level QWK differences and on the overall trait-level differences. Report 95% confidence intervals and Holm-corrected p-values. If the CI for the headline overall or averaged QWK difference includes zero, the claim of notable cross-prompt gains should be withdrawn or explicitly downgraded. As a secondary check, replace the corrected essay input with the original essay duplicated (or token-masked) to test whether any residual effect is specific to grammatical correction rather than extra model capacity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing problem is not GEC quality; it is that the reported headline effect may be within run-to-run noise. In Table 2, GAPS [TA+PA+GA] has an overall QWK of 0.670, numerically below ProTACT [TA+PA] at 0.674, despite using two encoders and a cross-attention module. Its advantage is confined to selected traits (Conventions +0.022, Language +0.012, Narrativity +0.011), while Organization and Word Choice are slightly lower. In Table 3, the average cross-prompt gain is 0.005 (0.597 vs 0.592) with per-seed standard deviations of 0.019 and 0.016. The highlighted Prompt 7 gain (+0.023) is offset by a 0.043 loss on Prompt 8 (0.498 vs 0.541). No significance test, confidence interval, or paired comparison is reported. Given overlapping standard deviations, the central claim that grammar awareness delivers 'notable QWK gains' is not supported by the evidence as presented. The GEC-dependency limitation is real but secondary: if the improvement is not statistically reliable, the mechanism is moot.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes GAPS, a grammar-aware cross-prompt multi-trait automated essay scoring method. A T5-based GEC model produces a corrected version of each essay; the scoring architecture uses two hierarchical encoders for the original and corrected texts, a cross-attention knowledge-sharing layer, explicit correction tags, trait attention, and prompt-independent features. The method is evaluated on ASAP/ASAP++ in a leave-one-prompt-out cross-prompt setting. The central claim is that feeding grammar-corrected text together with correction tags improves cross-prompt trait scoring, particularly for prompt-agnostic traits such as Conventions and Sentence Fluency, and that the grammar-aware signal is more useful than prompt-aware information in challenging low-resource prompts.","tokens_in":9111,"tokens_out":3182,"duration_ms":31513,"significance":"If the empirical gains are statistically reliable, the paper would make a useful contribution by showing that an externally obtained grammar-corrected input transfers to unseen prompts without auxiliary training. The strengths of the submission include a properly matched single-encoder baseline, an external GEC component with reported F0.5 scores on CoNLL-2014 and BEA-19, and ablations that isolate the knowledge-sharing layer and the correction tags. The main weakness is that the central headline claim rests on point estimates that are within one standard deviation of the strongest baseline, with no significance tests or confidence intervals. As presented, the evidence does not yet support the claimed 'notable QWK gains' over ProTACT.","major_comments":[{"comment":"The headline claim of 'notable QWK gains' over ProTACT is not supported by the reported numbers. GAPS [TA+PA+GA] has an overall QWK of 0.670, which is below ProTACT [TA+PA] at 0.674 and below Single Encoder at 0.673. The improvements are trait-specific (Conventions +0.022, Language +0.012, Narrativity +0.011), but no significance tests, confidence intervals, or per-trait standard deviations are provided. The conclusion that GA is superior for grammar-related traits requires paired comparisons or effect sizes with uncertainty estimates.","section":"§5, Table 2"},{"comment":"The average cross-prompt advantage over ProTACT is 0.005 (0.597 vs. 0.592), which is within the reported averaged standard deviations of 0.019 and 0.016. The highlighted Prompt 7 gain of +0.023 is offset by a loss of -0.043 on Prompt 8 (0.498 vs. 0.541). In the same table, the statement that GAPS 'consistently outperforms' the Single Encoder model is contradicted by Prompt 8, where GAPS scores 0.498 versus Single Encoder's 0.534. Without per-prompt variance estimates or a paired significance test, the central generalization claim is not established.","section":"§5, Table 3"},{"comment":"The conclusion that grammar-aware information (GA) outperforms prompt-aware information (PA) in low-resource cross-prompt settings is based on a single prompt (P7) and point estimates without variance information. Figure 2 shows no error bars or confidence intervals. This claim should either be supported with statistical evidence across multiple challenging prompts or be explicitly framed as a preliminary observation.","section":"§5, Figure 2 and 'Impact of grammar-aware vs. prompt-aware approaches'"}],"minor_comments":[{"comment":"There is a typo in 'Apendix 1'; it should be 'Appendix 1'.","section":"§3.2"},{"comment":"The caption lists trait abbreviations but does not expand 'Style', which is one of the traits evaluated for P7.","section":"Table 1"},{"comment":"The use of bold text to mark the highest values is not visible in the plain-text rendering; please ensure the final PDF clearly distinguishes bold values and defines what 'AVG' and 'SD' refer to in both tables.","section":"Tables 2 and 3"},{"comment":"The abstract claims 'notable QWK gains in the most challenging cross-prompt scenario', but the body shows that the gains over ProTACT are small and not tested for significance; the wording should be softened to match the evidence.","section":"Abstract and §5"}],"recommendation":"major_revision","confidential_remarks":"The central idea is worth publishing if the authors add proper significance testing and per-prompt/per-trait uncertainty estimates, and if they revise claims about consistent and notable gains. I have no concerns about citation patterns or novelty disclosure; the main issue is statistical support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a plausible new idea in cross-prompt AES—feed the model a GEC-corrected version of the essay as a second input, with explicit correction tags and a cross-attention knowledge-sharing layer. The ablations show those components help, and the writing is clear. But the headline claim of 'notable QWK gains' over the strongest prior system is not supported by the numbers: GAPS's average gain over ProTACT is 0.005 (0.597 vs. 0.592, SDs ~0.019 and ~0.016), and on the Overall trait it's actually 0.670 vs. 0.674. The Prompt 7 gain (+0.023) is offset by a 0.043 drop on Prompt 8. No significance tests or confidence intervals are reported. So the central effect is within run-to-run noise as presented.\n\nWhat's genuinely new: prior work used error counts, auxiliary detection, or hand-crafted features. Here the corrected text itself is the input, with tags marking M/R/U corrections, and a cross-attention layer shares knowledge between original and corrected representations. That's a clean, cheap idea. The evaluation is standard: leave-one-prompt-out on ASAP/ASAP++, five seeds, and the ablations (w/o KS, w/o GCT) consistently show small but plausible contributions from each component. The limitations section is honest about GEC dependency, though that's secondary to the statistical issue.\n\nSoft spots in proportion: The effect sizes are small everywhere. A few individual traits (Conventions +0.022, Language +0.012, Narrativity +0.011) improve, but others (Organization, Word Choice) go slightly backward. With overlapping SDs, these could easily be seed luck. The paper also cherry-picks Prompt 7 as 'most challenging' while ignoring Prompt 8, which is smaller and more variable. And there's no test of sensitivity to GEC quality, despite the acknowledged dependency.\n\nWho should read it: people working on cross-prompt AES will find the GEC-as-input idea worth trying, and the paper is careful enough to be a useful data point. But it doesn't establish a robust gain.\n\nRecommendation: send it to peer review—a decent referee would demand significance testing, error bars across prompts, and a tempered abstract. As submitted, the claims outrun the evidence.","headline":"Plausible new idea—GEC-corrected text as a second input—but the reported cross-prompt gains are small and within noise, so the 'notable' claim isn't supported.","tokens_in":9658,"tokens_out":4174,"would_cite":false,"duration_ms":38471,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Grammar-corrected essay text improves scoring on unseen prompts.","keywords":["automated essay scoring","cross-prompt generalization","grammar error correction","multi-trait scoring","prompt-independent features","knowledge sharing","ASAP++"],"falsifier":"Run the same GAPS pipeline on the same ASAP/ASAP++ splits but replace the grammar-corrected essay with a corrupted version, such as random word deletions or a deliberately poor correction model; if the quadratic weighted kappa gains on Conventions and Sentence Fluency persist, the grammar signal is not the cause, and if they vanish, the claim is supported. A second test is to evaluate on a new prompt whose essays contain error types the correction model handles badly, such as dialectal or heavily non-native writing, and check whether the grammar-trait advantage shrinks with correction accuracy.","tokens_in":8703,"feed_emoji":"📝","tokens_out":6877,"duration_ms":61383,"temperature":0.7,"pith_summary":"This paper proposes GAPS, a method for automated essay scoring that aims to score essays on prompts never seen during training. The idea is to feed the scoring model not only the original essay but also a grammar-corrected version produced by an off-the-shelf grammar error correction model, together with tags marking each inserted, replaced, or deleted token. Because grammar is largely independent of a prompt's content, the paper argues, referring to the corrected text lets the model learn generic, prompt-independent essay features. On the combined ASAP/ASAP++ benchmark, the method improves agreement with human raters in cross-prompt settings, with the largest gains in grammar-related traits such as Conventions and Sentence Fluency, and its best configuration also improves the hardest unseen-prompt case. A sympathetic reader would care because real deployment requires grading new prompts without retraining.","feed_headline":"Corrected essays help graders score unseen prompts","feed_subtitle":"Supplying original and corrected text improves prompt-agnostic trait scoring, with the largest gains in grammar and fluency.","key_machinery":"The method's load-bearing mechanism is a pair of hierarchical essay encoders that process the original and grammar-corrected essays separately, followed by a cross-attention knowledge-sharing layer. The corrected essay's representation is used as the query while the original essay supplies the key and value, letting the model align the two texts before trait-specific scoring. Each correction is annotated in the input with a tag of the form <corr> M: token </corr>, <corr> R: token </corr>, or <corr> U: token </corr>, marking tokens the grammar error correction model inserted, replaced, or deleted, so the network can attend to the exact revision points. The encoders share the same architecture and use part-of-speech embeddings, convolutional and LSTM layers, and attention pooling, and the trait-specific layers include trait attention so the model can relate the different scored dimensions.","core_discovery":"The central discovery is that directly providing a grammar-corrected version of an essay, rather than learning grammar signals through an auxiliary task, makes a multi-trait automated essay scoring model more robust to unseen prompts. The paper shows that the model's best configuration, which combines the grammar-aware input with prompt-aware and trait-aware components, raises prompt-wise average quadratic weighted kappa to 0.597 compared with 0.592 for the prior prompt- and trait-aware system, and lifts the most challenging prompt's score to 0.469 from 0.446. The gains concentrate in the traits that depend least on prompt-specific content: Conventions, Sentence Fluency, Language, and Narrativity. The paper interprets this as evidence that the corrected essay supplies a stable syntactic signal that transfers to prompts the model has not encountered.","pith_inferences":["Because the grammar error correction model is fixed and unsupervised relative to the scoring task, the same mechanism could in principle work with any text-normalization model that produces a stable second text differing from the original in grammar-relevant ways; the paper does not test this.","The dependency on grammar error correction quality flagged in the paper's limitations means the transfer advantage should erode on essay types where the correction model makes systematic errors, such as heavily non-standard or dialectal writing; this is a testable consequence, not a claim the paper makes.","The correction tags act as weak supervision pointing at revision spans; one could probe their contribution by training a scorer on original-plus-corrected text without tags but with token-level loss weights, a variant the paper does not explore."],"forward_implications":["If the central claim holds, cross-prompt automated essay scoring can be improved simply by preprocessing essays with an existing grammar error correction model, without new supervision or auxiliary training objectives.","Grammar-related traits such as Conventions should no longer lag behind semantic traits in cross-prompt evaluation, closing a gap that has persisted across prior systems.","Combining grammar-aware input with prompt-aware and trait-aware methods yields the strongest results, so the mechanism is a complement to, not a replacement for, existing transfer techniques.","The largest gains in the most difficult unseen prompt, a structurally different essay type with only partial trait overlap, suggest the approach is especially useful when training prompts are dissimilar to the target."],"supporting_citations":[{"why":"Supplies the pre-trained T5-based grammar error correction model that produces the corrected essay text.","marker":"Rothe et al. (2021)"},{"why":"Provides the error annotation taxonomy used to classify corrections into Missing, Replacement, and Unnecessary tags.","marker":"Bryant et al. (2017)"},{"why":"Provides the hierarchical trait-wise encoder, prompt-independent features, trait attention, and masking scheme that the method builds on.","marker":"Ridley et al. (2021)"},{"why":"Provides the prompt-aware and trait-aware components that the paper compares against and integrates with its grammar-aware module.","marker":"Do et al. (2023)"},{"why":"Provides the ASAP++ dataset that supplies the multiple trait scores used for training and evaluation.","marker":"Mathias and Bhattacharyya (2018)"},{"why":"Provides the strongest prior baseline, PLAES, that GAPS must beat in the cross-prompt comparison.","marker":"Chen and Li (2024)"}],"fun_headline_variants":["Grammar-fixed essays boost unseen prompt scores","Corrected text sharpens essay scoring on new prompts","Prompt-agnostic scoring gains from grammar correction","Grammar-aware model outperforms on unseen essay prompts","Fixing grammar helps AI grade unfamiliar essays"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes the pre-trained grammar error correction model produces corrections accurate enough on the kinds of student essays being scored to provide a clean, prompt-independent signal; if the corrected text is noisy, the second input could inject errors instead of generic grammar knowledge.","fun_headline_variants_meta":{"raw":{"variants":["Grammar-fixed essays boost unseen prompt scores","Corrected text sharpens essay scoring on new prompts","Prompt-agnostic scoring gains from grammar correction","Grammar-aware model outperforms on unseen essay prompts","Fixing grammar helps AI grade unfamiliar essays"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000606,"raw_usage":{"total_tokens":2776,"prompt_tokens":849,"completion_tokens":1927,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":1857}},"tokens_in":465,"tokens_out":1927,"duration_ms":14599,"temperature":1.0,"reasoning_tokens":1857,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T05:01:08.088173+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same GAPS pipeline on the same ASAP/ASAP++ splits but replace the grammar-corrected essay with a corrupted version, such as random word deletions or a deliberately poor correction model; if the quadratic weighted kappa gains on Conventions and Sentence Fluency persist, the grammar signal is not the cause, and if they vanish, the claim is supported. A second test is to evaluate on a new prompt whose essays contain error types the correction model handles badly, such as dialectal or heavily non-native writing, and check whether the grammar-trait advantage shrinks with correction accuracy.","supporting_citations":[],"review_version":1}