{"id":"a050888c-6850-47f0-95a0-030ef92eadee","arxiv_id":"2505.22771","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"On the PERSUADE corpus, adding generated argument-component tags to essay text raised automated scoring agreement from a QWK of 0.860 to 0.868, while error-only tags lowered it.","lead":"This paper tests whether adding machine-generated tags for argument structure and language errors to essay text helps an automated scorer. It reports a small accuracy gain for argument tags, but the grammar-error tags alone slightly hurt accuracy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The combined-annotations row in Table 6 is internally inconsistent, and error-only annotations do not beat full text; the robust finding is limited to argument-component markup.","rationale":"The reader's weakest assumption identified the lack of direct validation for the GEC-plus-ERRANT error annotations. That concern is real and is explicitly admitted in Section 2.3.2, but it is not the most load-bearing issue: even if the error tags were perfectly accurate, the Error Annotated condition in Table 6 shows no improvement over full text, and the Combined row as printed cannot be interpreted because its average is below its own minimum. The Component Annotated condition is the one clean positive result, so the broad abstract claim that feedback-oriented annotations enhance scoring is only supported for argumentative-component markup. This is likely a typo rather than a sign of fraud, but it blocks verification and should be corrected with raw results disclosed. The reader's conditional verdict already captures the need for additional evidence and code, so my read does not change the verdict; it sharpens the conditions.","tokens_in":10876,"tokens_out":6195,"duration_ms":72287,"concrete_test":"Release the per-run predictions or logs used for Table 6 and recompute the Combined row exactly. If the corrected Combined average is 0.868 with min 0.867, the table typo is confirmed and the claim should be assessed with the corrected value; if the corrected average is 0.866 or the min is 0.864, the combined-condition improvements over full text are not established. Either way, the authors should also run a paired permutation test across the 10 seeded runs to report whether the Component versus Full Text difference is statistically significant.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is operationalized by Table 6. The only fully coherent positive result is Component Annotated (QWK avg 0.868, min-max 0.867-0.870) versus Full Text (0.860, 0.859-0.862). The Error Annotated row (0.858, 0.856-0.859) is not above full text, so the error-annotation stream shows no benefit. The Combined row is printed with avg 0.866 but min 0.867, which is arithmetically impossible; either the average or the minimum is wrong, and without raw per-run values the combined condition cannot be interpreted. Section 2.3.2 concedes 'we have no direct way of evaluating the accuracy of any annotations' for the T5-GEC plus ERRANT tags, and Section 3.1 notes the pipeline 'does not seem to be uncovering as many errors as expected,' so that input is both unmeasured and apparently sparse. The abstract's broad phrasing 'feedback-oriented annotations' therefore overstates what is supported: the convincing effect is specifically the argumentative-component annotation, while the contribution of spelling and grammar annotations is unsupported and the combined result is unreliable as printed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an automated essay scoring (AES) pipeline in which two feedback-oriented annotation streams are inserted into the input text: argumentative component tags (Lead, Position, Claim, etc.) produced by a ModernBERT token classifier, and spelling/grammar error tags produced by a T5 grammatical-error-correction model combined with ERRANT error classification. Using the PERSUADE corpus, the authors fine-tune ModernBERT scorers on full text, component-annotated text, error-annotated text, and combined-annotated text, reporting QWK, exact agreement, and SMD over ten training runs. The headline result is that component-annotated scoring achieves average QWK 0.868 (range 0.867–0.870) versus 0.860 (0.859–0.862) for full text, with all conditions above the 0.745 human baseline. The paper also reports subgroup SMDs and argues that the combined annotations mitigate some bias.","tokens_in":11111,"tokens_out":3951,"duration_ms":44565,"significance":"If the component-annotation result holds, it is a useful empirical contribution: it shows that a high-accuracy argumentative-component annotator can provide a small but reproducible scoring gain over a strong full-text fine-tuned transformer baseline, while also producing interpretable markup. The ten-run min/max reporting is a strength, and the non-overlapping QWK ranges for component-annotated versus full-text models make the main comparison credible. However, the broad abstract claim about 'feedback-oriented annotations' is only supported for the argumentative-component stream, not for the error-annotation stream, and the combined-condition row in Table 6 is internally inconsistent as printed. The error-annotation pipeline also lacks any direct accuracy evaluation, a limitation the paper itself acknowledges in Section 2.3.2. The contribution is therefore narrower than claimed, though the component-level finding is real and fixable.","major_comments":[{"comment":"The Combined Annotations row is arithmetically impossible as printed: it reports average QWK 0.866 while also reporting a minimum QWK of 0.867. An average cannot be below its minimum. Because the paper's own stability argument relies on the min/max ranges, this row cannot be interpreted without the ten per-run values for QWK, exact agreement, and SMD. The combined condition is referenced in the abstract and in the bias discussion, so this inconsistency is load-bearing; it must be corrected with raw per-run data before any claim involving the combined pipeline is accepted.","section":"§3.2, Table 6"},{"comment":"The headline claim that 'incorporating feedback-oriented annotations into the scoring pipeline can enhance the accuracy of automated essay scoring' is not supported for the spelling/grammar error annotations. In Table 6, the Error Annotated condition (QWK 0.858, range 0.856–0.859) does not beat the Full Text baseline (0.860, range 0.859–0.862), and the ranges overlap. Section 2.3.2 explicitly concedes that 'we have no direct way of evaluating the accuracy of any annotations' for the T5-GEC/ERRANT stream, and Section 3.1 reports only 2795 spelling, 1401 grammatical, and 201 punctuation errors, noting that the pipeline 'does not seem to be uncovering as many errors as expected.' The central claim should be narrowed to argumentative-component annotations, or the error-annotation stream should be evaluated directly on PERSUADE.","section":"Abstract; §3.2, Table 6; §2.3.2"},{"comment":"The claim that combined annotations are 'mitigating some of the biases' rests on the Combined column of Table 7, but that condition's Table 6 row is internally inconsistent, so its SMDs are uninterpretable as printed. Moreover, even taken at face value, the Combined SMDs are larger than the Original SMDs for several subgroups (e.g., WC 0.15 vs 0.09, AP 0.52 vs 0.44). The bias discussion should be reframed as exploratory, and the primary bias comparison should be between the Full Text and Component Annotated conditions, which are both valid.","section":"§3.3, Table 7"},{"comment":"The paper justifies the T5-GEC error annotations by citing benchmarks on JFLEG and BEA60k and stating that the model is comparable to ChatGPT and GPT-4. This is an indirect transfer argument, and the paper itself notes the lack of direct evaluation on PERSUADE. Because the error-only condition fails to help, the manuscript should either provide a direct evaluation of error-tag accuracy on this domain or explicitly state that the error-annotation results are inconclusive rather than supportive of the abstract's claim.","section":"§2.3.2"}],"minor_comments":[{"comment":"The label 'ModerBERT' is a typo; it should read 'ModernBERT'.","section":"Figure 2"},{"comment":"There are several typographical slips, including 'pretaining' in Section 2.2.1, 'inter-rated reliability' in Section 2.1.2, and 'Ind..' in the Table 1 caption. These should be corrected.","section":"Throughout"},{"comment":"The parenthetical reference '(see (6))' after the reported human inter-rater reliability κ = 0.745 is misleading, because Equation (6) defines the QWK formula and the 0.745 value comes from the PERSUADE corpus paper. Please cite the source directly.","section":"§2.1.2"},{"comment":"The paper reports no significance tests or confidence intervals. The QWK comparison between Component Annotated and Full Text is supported by non-overlapping min/max ranges, but the exact-agreement and SMD differences would benefit from a paired test or bootstrap intervals, and the error-condition comparison should not be described as a difference without such support.","section":"§3.2, Table 6"},{"comment":"It would be helpful to report the macro-averaged F1 over the seven component classes explicitly, since the text says it is the statistic of interest, rather than leaving the reader to compute it from the TP/FP/FN columns.","section":"§3.1, Table 5"}],"recommendation":"major_revision","confidential_remarks":"The central scoring comparison (component-annotated vs full-text) is reproducible and non-overlapping, so this is not a reject. The main problems are the impossible combined-row statistics and the overbroad claim about error annotations. I would be willing to re-review a revision that narrows the claim, corrects Table 6 with per-run values, and either evaluates or explicitly de-emphasizes the error-annotation stream. I do not see a circularity problem: the annotator is trained on the training split and applied to the test split, which is standard supervised learning."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know about this paper for one reason: the component-annotated scoring result is real. Over 10 runs, the Component Annotated model gets QWK 0.868 (0.867–0.870) against 0.860 (0.859–0.862) for full text on PERSUADE. The ranges don't overlap, so the effect, though small, is stable. That is the paper's contribution.\n\nWhat's new is the specific combination: generated argument-component tags plus ERRANT error tags fed as XML into a ModernBERT scorer. The idea that feedback features can help scoring predates this paper (Uto and Uchida 2020; Crossley et al. 2022), but the concrete comparison on PERSUADE is new. The authors also report subgroup SMDs, which is a plus.\n\nThe soft spots are real but addressable. First, the abstract says 'feedback-oriented annotations' broadly, but only the argument-component stream helps. Error-only annotations get QWK 0.858, below full text, and the combined condition is worse than component-only. Section 3.1 itself admits the error pipeline 'does not seem to be uncovering as many errors as expected.' So the headline should be narrowed to argument-component annotation.\n\nSecond, Table 6's Combined row has avg 0.866 with min 0.867, which is arithmetically impossible. Either the average is wrong or the min is. Without raw per-run numbers, that condition can't be interpreted. This looks like a typo, but it needs fixing.\n\nThird, there are no significance tests, only min-max ranges. Those are a weak form of inference. Fourth, the GEC annotations are not directly validated—Section 2.3.2 says so—so the error-stream results are interpretable only as 'this particular noisy pipeline didn't help.' Fifth, no code or data release, making replication harder.\n\nNone of these sink the component-annotation finding. The authors are honest about the annotation uncertainty, and the argument-component annotator itself is well evaluated (F1 0.818–0.960). The bias analysis is thoughtful, even if the SMDs are noisy for small subgroups.\n\nWho should read this: anyone building AES with PERSUADE or working on interpretable features. It deserves a serious referee, with requested revisions: fix Table 6, add significance tests or per-run values, and rewrite the abstract to claim only what the component-annotation result supports. I'd engage with it.","headline":"Component-annotation scoring on PERSUADE is a real but small effect; the paper overstates it, and Table 6 has an impossible summary row.","tokens_in":11603,"tokens_out":2442,"would_cite":true,"duration_ms":22466,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automated essay scoring improves when the model sees argument-structure and error tags alongside the text.","keywords":["automated essay scoring","argumentative components","grammatical error correction","PERSUADE corpus","ModernBERT","ERRANT","annotation","scoring bias"],"falsifier":"Re-run the scoring experiments replacing the predicted argumentative-component tags with the corpus's gold human annotations; if the gold-annotated model does not outperform the full-text baseline (QWK 0.860), then the reported gains come from annotation artifacts rather than the argument structure itself. Separately, a human-annotated error set for a sample of PERSUADE essays would test whether the T5-ERRANT annotations used here are accurate enough to support the error-only and combined pipelines.","tokens_in":10689,"feed_emoji":"📝","tokens_out":3419,"duration_ms":40796,"temperature":0.7,"pith_summary":"This paper tests whether feeding automated essay scoring (AES) models the same annotations that automated writing feedback systems produce can make scoring more accurate. Using the PERSUADE corpus of persuasive essays, the author augments each essay's text with XML tags marking argumentative components (lead, position, claim, evidence, and so on) and with tags marking spelling, punctuation, and grammar errors. The central claim is that these feedback-oriented annotations help an encoder-based fine-tuned language model agree better with human scores than the same model trained on plain text alone. The best configuration, component-annotated text, reaches a quadratic weighted kappa of 0.868 versus 0.860 for full text, both above the reported human inter-rater agreement of 0.745. A careful reader would care because this offers a path to scoring models that are both more accurate and more aligned with the rubric-based feedback students actually receive.","feed_headline":"Essay tags lift automated scoring accuracy","feed_subtitle":"Marking argument structure and errors in the text pushes model-human agreement from 0.860 to 0.868 QWK.","key_machinery":"The central mechanism is the XML-augmented input: the essay text is wrapped in <Lead>, <Position>, <Claim>, <Counterclaim>, <Rebuttal>, <Evidence>, <Concluding Statement> tags for argumentative components, and <Spelling>, <PunctOrth>, and <Grammar> tags for conventions errors. This markup lets a single long-context encoder model, ModernBERT, consume both the original text and the feedback-oriented structure at once. ModernBERT's rotary positional embeddings allow the scorer to process full essays without truncation, and the annotation model is the same architecture with a token-classification head. The argumentative-component tags and error tags are the load-bearing information added to the scorer's input, and the paper's premise is that these tags make rubric-relevant organization and language conventions explicit to the model.","core_discovery":"The paper demonstrates, on 25,996 PERSUADE essays graded on a 1-6 holistic SAT-style rubric, that adding automated annotations to the scoring input improves agreement with human scores. A ModernBERT token classifier predicts seven argumentative component types, and a T5 grammatical-error-correction model followed by ERRANT produces spelling, punctuation, and grammar error tags; these are inserted into the essay text as XML. Averaged over ten training runs, the component-annotated model scores QWK 0.868 (range 0.867-0.870) and exact agreement 68.9%, compared with QWK 0.860 (range 0.859-0.862) and exact agreement 67.2% for the full-text baseline. Error annotations alone do not improve over full text (QWK 0.858), while combining argument and error annotations gives QWK 0.866. The author also reports standardized mean differences across demographic subgroups, finding that the combined-annotation model shows smaller negative bias for several groups, including English Language Learners, compared with the full-text model.","pith_inferences":["If argument-structure tags are the source of the gain, then other argumentative writing corpora with discourse annotations should show similar improvements when scored with annotated input.","The error-only result may reflect annotation noise (the paper admits it has no direct accuracy measure for error tags) or it may mean conventions information is redundant with what the encoder already infers from text; a direct human-annotated error set would separate these possibilities.","Because the annotator model itself is trained on the same essays used for scoring, the pipeline could be evaluated as a fully automated loop where annotation quality is part of the scoring system, not a fixed external input.","If the bias reductions replicate, feedback-driven annotations could become a practical fairness intervention for AES, particularly for English Language Learners, though the small subgroup sizes make these SMD estimates uncertain."],"forward_implications":["If this holds, scoring engines can be made more accurate simply by feeding them annotations that automated writing feedback systems already generate, reusing existing infrastructure.","The component-annotated model beats the full-text model across all ten training runs, suggesting the gain is stable rather than a lucky seed.","Error annotations alone did not help, so the improvement appears to come from argumentative structure rather than from error flags.","XML-tagged output could let one model both score an essay and produce interpretable feedback for students, linking assessment with instruction.","The combined-annotation pipeline shows smaller negative bias for several demographic subgroups, hinting that annotation-informed scoring may reduce rather than exacerbate disparities."],"supporting_citations":[{"why":"Supplies the PERSUADE corpus, its gold argumentative annotations, the holistic rubric scores, and the 0.745 human inter-rater reliability baseline.","marker":"(Crossley et al., 2022)"},{"why":"Provides the ERRANT tool that classifies spelling, punctuation, and grammar errors into the categories collapsed into the three error tags.","marker":"(Bryant et al., 2017)"},{"why":"Defines the T5 grammatical error correction model used to generate corrected sentences from which ERRANT derives the error annotations.","marker":"(Martynov et al., 2023)"},{"why":"Introduces ModernBERT, the long-context encoder whose rotary positional embeddings enable both annotation and scoring of full essays beyond 512 tokens.","marker":"(Warner et al., 2024)"},{"why":"Establishes the operational evaluation framework, including QWK, exact agreement, and standardized mean difference for subgroup bias, used throughout the paper.","marker":"(Williamson et al., 2012)"},{"why":"Defines the argumentative component annotation schema (lead, position, claim, counterclaim, rebuttal, evidence, concluding statement) that the annotator model learns to predict.","marker":"(Stab and Gurevych, 2014a)"}],"fun_headline_variants":["Argument markers raise automated essay scoring accuracy","Add feedback tags to essays for better scoring","Scoring gains from argument annotations in essays","Tagging argument structure improves essay scoring","Annotations in essays lift scoring agreement with humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automated argumentative-component annotations are accurate enough on PERSUADE essays that the scoring gains reflect genuine rubric-relevant structure, and the error annotations, whose accuracy is never directly evaluated, are reliable enough not to corrupt the combined pipeline.","fun_headline_variants_meta":{"raw":{"variants":["Argument markers raise automated essay scoring accuracy","Add feedback tags to essays for better scoring","Scoring gains from argument annotations in essays","Tagging argument structure improves essay scoring","Annotations in essays lift scoring agreement with humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000295,"raw_usage":{"total_tokens":1683,"prompt_tokens":880,"completion_tokens":803,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":739}},"tokens_in":496,"tokens_out":803,"duration_ms":9909,"temperature":1.0,"reasoning_tokens":739,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:00:18.686205+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the scoring experiments replacing the predicted argumentative-component tags with the corpus's gold human annotations; if the gold-annotated model does not outperform the full-text baseline (QWK 0.860), then the reported gains come from annotation artifacts rather than the argument structure itself. Separately, a human-annotated error set for a sample of PERSUADE essays would test whether the T5-ERRANT annotations used here are accurate enough to support the error-only and combined pipelines.","supporting_citations":[{"cited_title":"A Methodology for Generative Spelling Correction via Natural Spelling Errors Emulation across Multiple Domains and Languages","cited_arxiv_id":"2308.09435","evidence_quote":"Defines the T5 grammatical error correction model used to generate corrected sentences from which ERRANT derives the error annotations."}],"review_version":1}