{"id":"247f95de-4fb3-4563-b45b-eefbbc534712","arxiv_id":"2501.00715","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"eRevise+RF combines automated essay scoring and revision classifiers to give grades 4-8 students feedback on evidence use and revision quality, with pilot results showing draft-over-draft improvement.","lead":"Researchers built and tested an AI system that scores middle-school essays on evidence use and gives feedback on how students revised their essays between drafts. In a pilot with 406 students, essay quality improved across three drafts, but there was no control group to prove the feedback caused the improvement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline-dependent feedback assignment confounds the Section 5.3 causal claim: EF1/EF2 groups start with lower NPE/SPC, so their large gains may be regression to the mean; without significance tests or attrition analysis, helpfulness is not established.","rationale":"The paper has real strengths: the NLP components are evaluated against expert annotations (QWK 0.67/0.82 for NPE/SPC, F1 0.96/0.66/0.70 for classifiers), the system was deployed at scale with diverse schools, and the Limitations section transparently acknowledges the missing control group. The RQ1 assessment claims are supported. The load-bearing weakness is in RQ2's causal claim. Section 5.3 tries to address the no-control problem by showing that improvements align with specific feedback messages. However, feedback assignment is not random: EF1 is given to essays with NPE≤α, EF2 to essays with NPE>α and SPC≤β, EF3 to the rest. This creates systematic baseline differences and unequal room for improvement. The dramatic percentage gains for EF1 (NPE +128%, SPC +108%) and EF2 (SPC +55%) are consistent with regression to the mean from low baselines; the tiny EF2 subsample (N=11) makes the 55% estimate especially fragile. No significance testing, confidence intervals, or baseline-adjusted analyses are provided. In addition, the complete-case sample excludes 22 MVP and 26 SPACE students who missed a session; if absence correlates with lower writing performance, the reported gains are biased upward. The paper acknowledges the control-group limitation but does not acknowledge these specific confounds. A concrete reanalysis of Table 4 with ANCOVA/adjusted gains and an attrition check on draft1 scores would settle whether the differential gains are real. If they are not, the claim that both evidence use feedback and revision feedback are helpful (Sec 5.3) overstates what the data can show; the paper should be framed as a promising pilot with descriptive gains, pending a controlled or baseline-adjusted evaluation.","tokens_in":22137,"tokens_out":7975,"duration_ms":73356,"concrete_test":"Reanalyze the MVP data behind Table 4 with baseline-adjusted gain scores: fit ANCOVA models ΔNPE = EF + NPE_draft1 and ΔSPC = EF + SPC_draft1, reporting adjusted group means with 95% CIs, and test whether EF1's adjusted NPE gain still exceeds EF2/EF3's. Separately, compare draft1 NPE/SPC/WC of the 22 non-completers (194 enrolled vs 172 analyzed) against completers; if dropouts have lower baselines or the adjusted EF differences are not significant, the helpfulness claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.3's evidence for feedback effectiveness is confounded by baseline-dependent assignment. EF levels are selected via thresholds on NPE and SPC (Appendix D, Table 5: EF1 if NPE≤α, EF2 if NPE>α and SPC≤β, EF3 otherwise). Consequently, the EF1 group has the lowest baseline NPE (1.09 vs 2.64 vs 3.03, Table 4), giving it the most room for improvement; the observed 128% NPE gain and 108% SPC gain for EF1 are exactly the pattern expected from regression to the mean and floor effects even if feedback had no specific content. The EF2 group (N=11) is tiny, and its 55% SPC gain is highly unstable. No significance tests, confidence intervals, or baseline-adjusted analyses are reported. Additionally, the analysis sample excludes 22 MVP students (194→172) and 26 SPACE students (176→150) who did not complete all drafts; if attrition correlates with low performance, the completers' gains are biased upward. The paper's Limitations acknowledge the missing control group, but the Section 5.3 comparison was intended to address this; as it stands, the comparison cannot distinguish feedback-specific effects from baseline confounding and attrition bias. Therefore the central claim that 'both evidence use feedback and revision feedback are helpful' (Sec 5.3) is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents eRevise+RF, an automated writing evaluation system that scores student argumentative essays on evidence use (NPE and SPC), extracts and classifies revisions across drafts, and provides expert-designed evidence-use feedback (EF1-EF3) and revision feedback (RF1-RF10). The system was deployed with 6 teachers and 406 students in grades 4-8 across two RTA article conditions; the analysis is based on students who completed all three drafts (172 MVP and 150 SPACE essays). Two research questions are addressed: RQ1 asks whether the NLP components can effectively assess essays and revisions, and RQ2 asks whether the feedback helps students improve. RQ1 is evaluated against human annotations, reporting QWK of 0.67 (NPE) and 0.82 (SPC), and F1 of 0.96 (RC-Content), 0.66 (RC-Evidence), and 0.70 (RC-Success). RQ2 is evaluated by tracking predicted (and for MVP, gold-annotated) NPE/SPC across drafts and by comparing gains for students receiving different EF levels. The paper concludes that both evidence-use feedback and revision feedback help students improve their writing.","tokens_in":22422,"tokens_out":3284,"duration_ms":33166,"significance":"The NLP assessment contribution of RQ1 is solid and well supported: the evaluation uses gold human annotations, reports standard agreement metrics, includes grade-level breakdowns and confusion matrices, and the system source code is publicly released. The integration of revision classification with formative feedback in a deployed, multi-school system is a useful engineering contribution for the AWE community. The causal claim embedded in RQ2, however, is not established by the current analysis: the draft-to-draft comparisons lack a control group, the feedback assignment is baseline-dependent, and no significance or attrition analyses are reported. The paper's own Limitations section acknowledges the missing control group, but the Section 5.3 analysis is intended to address this gap and, as presented, cannot support the inference that feedback, rather than practice or floor effects, drove the observed gains.","major_comments":[{"comment":"The Section 5.3 evidence for feedback effectiveness is confounded by the feedback assignment mechanism. Appendix D (Table 5) assigns EF1 when NPE is at most alpha, EF2 when NPE > alpha and SPC at most beta, and EF3 otherwise. Consequently, the EF1 group starts with the lowest baseline NPE (1.09 vs. 2.64 vs. 3.03 in the annotated MVP rows of Table 4) and has the most room for improvement. The observed EF1 gains of +128% NPE and +108% SPC are exactly the pattern expected from regression to the mean and floor effects even if the feedback content had no specific effect. No significance tests, confidence intervals, or baseline-adjusted analyses are reported. The Section 5.3 conclusion that 'both evidence use feedback and revision feedback are helpful for students to improve their writing' therefore does not follow from these data.","section":"5.3, Table 4, Appendix D"},{"comment":"The analysis sample is the set of students who submitted all three drafts, but the deployment-level counts in Section 4 show substantial attrition: 194 MVP draft-1/draft-2 to 172 draft-3, and 176 SPACE draft-1/draft-2 to 150 draft-3. If attrition correlates with low performance or low engagement, the completers' gains are biased upward. The paper does not report any comparison of baseline (draft-1) NPE/SPC between completers and non-completers, nor does it state the direction of any such difference. Without an attrition analysis, the draft-over-draft improvements in Tables 3 and 4 cannot be interpreted as average gains for the full deployed population.","section":"4, 5.3"},{"comment":"The draft2-to-draft3 comparison intended to support the revision-feedback claim is also not statistically validated. The EF2 group in the MVP annotated analysis has only N=11 students, and its predicted NPE actually decreases by 10.4% (Table 4), so the aggregate patterns are unstable and are not disaggregated by the ten RF types used in Figure 3/Table 6. Additionally, the RF selection relies on RC-Evidence and RC-Success, whose F1 scores are 0.66 and 0.70 (Table 2); the paper does not examine how classifier errors propagate to the RF messages or to the observed NPE/SPC changes. The claim that revision feedback is helpful would require either a comparison against a no-revision-feedback condition or at least a per-RF-type analysis with effect sizes and uncertainty intervals.","section":"5.3, Appendix E"}],"minor_comments":[{"comment":"The text reads 'Glove embedding' and 'GloVe' in the same sentence; the capitalization should be consistent (GloVe).","section":"3.1"},{"comment":"Please state whether the QWK values for NPE and SPC are computed on the 516 MVP essays described in the text, and clarify why the revision-classifier sample sizes (1,525 and 1,024) differ from the 172 re-annotated draft pairs; the current paragraph is clear on the sources but a one-sentence summary would help.","section":"5.1, Table 2"},{"comment":"The paper reports percentage increases (e.g., 24.6% NPE, 38.5% SPC) without confidence intervals. Adding standard errors or effect sizes would help the reader judge the stability of the improvement.","section":"5.2"},{"comment":"The definition of alpha as 'half of the topics discussed' applies to the MVP article (4 topics), yielding alpha=2, but the SPACE article also has 4 topics; if alpha and beta were set identically for both articles, this should be stated explicitly.","section":"Appendix D, Table 5"},{"comment":"The case study describes the student receiving RF after draft2, but the figure caption does not indicate the specific RF message; adding the message text next to the relevant arrow would make the example easier to follow.","section":"5.4, Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a competent systems paper for the BEA/AWE audience, and the RQ1 evaluation is credible. The main risk is that the RQ2 claim is presented too strongly relative to the design. In a revision I would expect the authors to either add baseline-adjusted analyses (or a clear regression-to-the-mean discussion with effect sizes), provide an attrition comparison, and explicitly reframe RQ2 as 'students' revisions are responsive to feedback' rather than 'feedback causes improvement.' The paper's self-citations are frequent but not inappropriate given its lineage; the system integration is genuinely novel."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you the short version: eRevise+RF is a real deployed system that integrates previously published NLP scoring and revision classifiers into a web tool with a new 10-level revision feedback taxonomy, and it reports a two-site pilot with 406 students. The strength is the assessment side: NPE and SPC scoring agree with human annotations at QWK 0.67 and 0.82, and the revision classifiers get F1s from 0.66 to 0.96, with error analysis in the appendix. That part is fine and honestly reported.\n\nThe soft spot is the causal claim. Section 5.3 argues that feedback helps because students who received EF1 and EF2 showed large NPE/SPC gains from draft1 to draft2. But the EF levels are assigned by thresholds on the same NPE/SPC scores, so EF1 students start with the lowest scores (NPE 1.09 vs 2.64 and 3.03) and have the most room to improve. The 128% and 108% gains are exactly the pattern you'd expect from regression to the mean and floor effects. The EF2 group has N=11, so its 55% gain is unstable. On top of that, the analysis only includes students who completed all three drafts (172 of 194 MVP, 150 of 176 SPACE); if lower performers attrited, the completers' gains are biased upward. No significance tests, confidence intervals, or baseline-adjusted models are reported. The Limitations section does acknowledge the missing control group, but says Section 5.3 is meant to start teasing apart the effect; as it stands, that comparison cannot distinguish feedback-specific effects from these confounds.\n\nSo my verdict: the paper is a solid system description with transparent reporting, and the NLP assessment claims (RQ1) are credible. But the abstract's 'confirmed its effectiveness' is too strong for the feedback-helpfulness claim (RQ2). That needs a control condition or at least a baseline-adjusted analysis and attrition check before it can be accepted.\n\nI'd send it to peer review—a serious referee can push for the revision. The system is deployed, the code is on GitHub, and the paper is honest about its limits. It's a useful data point for AWE research, but I wouldn't cite the causal claim yet.","headline":"Solid system description with honest NLP evaluation, but the causal claim that feedback drives writing gains is not supported by the current evidence.","tokens_in":22983,"tokens_out":2316,"would_cite":false,"duration_ms":20808,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automated revision feedback lifts student essay quality","keywords":["automated writing evaluation","revision feedback","argumentative writing","evidence use","formative feedback","essay revision classification","Response-to-Text Assessment","NLP in education"],"falsifier":"A controlled deployment in which comparable classrooms write the same Response-to-Text essay through three drafts without any eRevise+RF feedback would settle it: if their NPE and SPC gains match the feedback group's gains, the system's feedback is not the cause of improvement.","tokens_in":21923,"feed_emoji":"📝","tokens_out":6502,"duration_ms":57782,"temperature":0.7,"pith_summary":"This paper presents eRevise+RF, an automated writing evaluation system that tries to do two things at once: assess whether a student's revision actually responds to the feedback they received, and then give new feedback on the revision itself. The system was deployed with 406 students in grades 4 to 8 across three schools, where students wrote three drafts of an argumentative essay. The paper claims the assessment works: its evidence-use scores agree with human raters, and its revision classifiers identify surface versus content changes, evidence versus reasoning changes, and successful versus unsuccessful revisions. It also claims the feedback works: essay quality, measured by evidence count and specificity, improves from draft to draft, and the pattern of gains lines up with which feedback message each student received. A sympathetic reader would care because the system addresses a known gap: most automated writing evaluation systems score a single draft but do not tell students whether their revision fixed the problem.","feed_headline":"Automated revision feedback lifts student essay quality","feed_subtitle":"A deployed NLP system scored drafts, classified revisions, and guided 406 students in grades 4-8 across three drafts.","key_machinery":"The carrying mechanism is a two-stage pipeline of evidence scoring indicators and revision classifiers. The indicators NPE and SPC are extracted from each draft using sliding windows and keyword similarity to count evidence topics and specific examples; thresholds on them select one of three evidence-feedback messages. For revisions, consecutive drafts are sentence-aligned, and three classifiers label each revision pair: RC-Content (surface or content), RC-Evidence (evidence or reasoning), and RC-Success (successful or unsuccessful). These labels, combined with NPE/SPC changes, navigate a feedback tree with ten revision-feedback levels, from 'no attempt' through 'repeated evidence' to 'successful evidence plus successful reasoning'.","core_discovery":"The central claim is that natural language processing can scaffold the revision process in argumentative writing, not just grade the final product. eRevise+RF extends an earlier evidence-feedback system by adding a revision-assessment layer: it aligns consecutive drafts, classifies each content revision, and maps the classifications plus changes in evidence indicators onto a ten-level revision feedback tree. In the deployment, NPE (number of evidence topics) and SPC (specificity) rose across drafts, with the largest jumps after the first round of evidence-use feedback and smaller but consistent gains after revision feedback; annotated gold scores showed the same direction as system predictions. The paper's summary is that both evidence use feedback and revision feedback are helpful, in ways aligned with the feedback message each student received.","pith_inferences":["Editorial inference: the deployment had no control group, so some of the observed improvement may reflect the simple act of writing three drafts; a randomized or wait-list design would separate the feedback effect from practice effects.","Editorial inference: because RC-Evidence and RC-Success were trained partly on college-level revision corpora and have F1 of 0.66 and 0.70, their mistakes may concentrate on younger students' less explicit reasoning; fine-tuning on the newly collected grade 4-8 revisions could tighten the feedback loop.","Editorial inference: the observed NPE drop of 11.3% when EF2 students revised for specificity hints at a completeness-specificity trade-off that future feedback could address explicitly by telling students to preserve breadth while adding detail.","Editorial inference: the revision feedback tree is a reusable template for other formative tasks in which a learner responds to feedback, such as scientific explanations or short-answer reasoning, with the success labels guiding the next prompt."],"forward_implications":["Students who received evidence-use feedback on their first draft and revision feedback on their second draft produced third drafts with higher predicted NPE and SPC, across both the MVP and SPACE reading passages.","The system can separate superficial edits from meaning-altering changes with F1 of 0.96, so revision feedback can focus on content changes rather than typos.","Because feedback selection is rule-based and transparent, teachers and students can see why a particular message was chosen, unlike a black-box language model.","The revised NPE algorithm reported in the limitations reaches quadratic weighted kappa of 0.87 against human annotations, up from 0.67, suggesting the scoring component can be improved with data-driven tuning."],"supporting_citations":[{"why":"Supplies the base eRevise evidence scoring indicators NPE and SPC and the original evidence feedback messages.","marker":"Zhang et al., 2019"},{"why":"Supplies the RC-Success revision quality model and the Argument Context extraction used to judge whether a revision succeeds.","marker":"Liu et al., 2023"},{"why":"Supplies the AargRewrite v2 college-level revision corpus used to train the RC-Content surface-versus-content classifier.","marker":"Kashefi et al., 2022"},{"why":"Supplies the evidence/reasoning revision annotation scheme and the RER quality labels used to train and evaluate the revision classifiers.","marker":"Afrin et al., 2020"},{"why":"Supplies the evidence topic and specificity keyword lists that the NPE and SPC sliding-window algorithms rely on.","marker":"Rahimi et al., 2017"},{"why":"Defines the Response-to-Text Assessment task that the system and its essay prompts are built around.","marker":"Correnti et al., 2013"},{"why":"Provides the validity argument and annotation procedures used for scoring evidence in the RTA essays.","marker":"Correnti et al., 2022"},{"why":"Documents that students often revise without improving their evidence use, motivating the addition of revision feedback.","marker":"Wang et al., 2020"},{"why":"Provides the Bertalign sentence alignment tool used to extract revision pairs between consecutive drafts.","marker":"Liu and Zhu, 2022"},{"why":"Supports the design choice of expert-crafted feedback over LLM-generated feedback for young student writers.","marker":"Behzad et al., 2024"}],"fun_headline_variants":["Revision-aware feedback system boosts essay quality in 406 students","AI revision feedback lifts essay quality across three drafts","Feedback on revisions sharpens argumentative writing skills","Revision feedback, not just grading, lifts essay quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that feedback, rather than practice, caused the essay gains assumes that students who wrote three drafts without receiving feedback would not have improved as much, because the deployment had no control group.","fun_headline_variants_meta":{"raw":{"variants":["Revision-aware feedback system boosts essay quality in 406 students","AI revision feedback lifts essay quality across three drafts","Feedback on revisions sharpens argumentative writing skills","Revision feedback, not just grading, lifts essay quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001317,"raw_usage":{"total_tokens":5311,"prompt_tokens":842,"completion_tokens":4469,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":4407}},"tokens_in":458,"tokens_out":4469,"duration_ms":35520,"temperature":1.0,"reasoning_tokens":4407,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:43:21.188947+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled deployment in which comparable classrooms write the same Response-to-Text essay through three drafts without any eRevise+RF feedback would settle it: if their NPE and SPC gains match the feedback group's gains, the system's feedback is not the cause of improvement.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the base eRevise evidence scoring indicators NPE and SPC and the original evidence feedback messages."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the evidence/reasoning revision annotation scheme and the RER quality labels used to train and evaluate the revision classifiers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the evidence topic and specificity keyword lists that the NPE and SPC sliding-window algorithms rely on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Response-to-Text Assessment task that the system and its essay prompts are built around."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Bertalign sentence alignment tool used to extract revision pairs between consecutive drafts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the design choice of expert-crafted feedback over LLM-generated feedback for young student writers."}],"review_version":1}