{"id":"dfa6daf4-3d81-4b94-9209-fb1e05e809fa","arxiv_id":"1908.01992","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A classroom pilot of an automated writing feedback system called eRevise found improved text evidence use after revision, but the effect on human rubric scores was only marginally significant and there was no control condition.","lead":"This paper presents eRevise, a web-based writing tool that uses natural language processing to give fifth and sixth grade students automated feedback on how they use text evidence in their essays. In a pilot study with 143 students, text evidence usage improved after students revised their drafts using the tool, though the study lacked a control group.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No control condition means the pilot cannot support the causal claim that eRevise's adaptive feedback—rather than revision itself—improved evidence use; the planned generic-feedback comparison would settle it.","rationale":"The reader's weakest assumption is exactly the load-bearing concern: the absence of a control condition prevents causal attribution of the observed improvements to eRevise's adaptive feedback. The paper is honest about this limitation and has planned a generic-feedback control, so a conditional verdict with a request to soften the causal wording is appropriate. I do not see a more damaging internal inconsistency. The H2 measures are somewhat circular because they are tied to feedback selection, and the human-score improvement is only marginal, but these issues are secondary to the missing control. The differential improvement on Evidence versus the other four RTA dimensions is a useful internal check, yet it cannot fully substitute for a control group. A controlled deployment, ideally with a revise-only arm, would directly test the causal claim. Since the reader already identified this concern and the verdict is conditional, no adjustment to the verdict is needed.","tokens_in":9223,"tokens_out":3326,"duration_ms":41454,"concrete_test":"Run the controlled comparison already planned in Current and Future Directions, with three arms: adaptive eRevise feedback, a generic feedback message, and a revise-only condition with no feedback, using the same RTA MVP article and holding teacher instruction constant across arms. Pre-register the primary outcome as the human RTA Evidence score and secondary outcomes as NPE and SPC Total Merged. If the generic-feedback and revise-only arms show the same pre-post gains as the adaptive arm, the causal role of adaptive feedback is not supported; if the adaptive arm significantly outperforms both, the central claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is causal: eRevise's adaptively selected feedback is said to help students improve text evidence usage. The deployment, however, has no control condition in which students revise without eRevise feedback or with generic feedback. Pre-post gains on the human Evidence rubric (2.62 to 2.72, p ≤ 0.08), NPE (2.61 to 2.81, p ≤ 0.003), and SPC Total Merged (9.65 to 11.15, p ≤ 0.001) could be produced by the opportunity to revise, task familiarity, or teacher scaffolding in the second class period rather than by the adaptive messages. The authors explicitly acknowledge this in Current and Future Directions, where they report adding a generic-feedback control condition. The differential null results on the other four RTA dimensions provide some evidence against a general writing-improvement effect, but they do not rule out revision-specific or evidence-prompt-specific effects. A compounding issue is that H2's outcome features (NPE and SPC Total Merged) are the same features used to select feedback, so increases may reflect literal compliance with 'use more evidence' or copying article phrases rather than improved evidentiary reasoning. The human-score result, which is not circular, is only trending and is driven mainly by the low-scoring subgroup receiving messages 1 and 2 (p = 0.02), a pattern compatible with regression to the mean. Without a control arm, the abstract's wording that eRevise 'empowers students to better revise' overstates what the data can establish.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents eRevise, a web-based automated writing evaluation (AWE) environment for grades 5–6. eRevise uses an existing rubric-based automatic essay scoring (AES) system (the SG word-embedding model) to extract features — Number of Pieces of Evidence (NPE) and a merged specificity count (SPC Total Merged) — and then selects two of four predefined feedback messages via a hand-built lookup table. In a pilot deployment across seven classrooms (143 students), the authors test whether students improve from first to second drafts on human rubric Evidence scores (H1) and on the NLP feature counts (H2). They report a trending improvement on the human rubric (2.62 to 2.72, p ≤ 0.08), significant improvements on NPE (2.61 to 2.81, p ≤ 0.003) and SPC Total Merged (9.65 to 11.15, p ≤ 0.001), and no significant changes on the other four RTA dimensions. The authors acknowledge in the Current and Future Directions section that the next deployment will include a generic-feedback control condition. The central claim is that eRevise's adaptively selected feedback improves text evidence usage.","tokens_in":9699,"tokens_out":2513,"duration_ms":25634,"significance":"If the causal claim were established, eRevise would be a notable contribution to automated writing evaluation for upper elementary students, particularly for its focus on evidence use, its rubric-aligned interpretable features, and its use of NLP-generated feedback rather than scores alone. The paper has clear strengths: it validates the choice of the interpretable SG model over the better-performing CO-ATTN model, it provides a concrete feedback-selection algorithm with pre-specified thresholds, it reports classroom-deployment data with paired pre-post design and human rubric scoring, and it publicly acknowledges the lack of a control condition. The differential null results on non-targeted RTA dimensions provide some useful evidence that the observed gains are not merely a general writing-improvement effect. However, as presented, the evidence supports only a descriptive 'students improved after using eRevise' conclusion; the abstract's causal phrasing overstates what the data can establish.","major_comments":[{"comment":"The study lacks any control condition. The pre-post gains on the Evidence rubric (2.62 to 2.72, p ≤ 0.08), NPE (2.61 to 2.81, p ≤ 0.003), and SPC Total Merged (9.65 to 11.15, p ≤ 0.001) could be produced by the opportunity to revise, by teacher scaffolding in the second class period, or by task familiarity. The authors explicitly state in Current and Future Directions that a generic-feedback control condition has been added for the next deployment, which confirms that the current design cannot support the causal claim 'eRevise helped students improve their text evidence usage.' The abstract and conclusions should be reworded to describe a preliminary feasibility study with descriptive pre-post comparisons, or the authors should provide a compelling post-hoc argument ruling out these alternative explanations (e.g., using the non-targeted RTA dimensions as a control is only partially convincing because those dimensions did not receive feedback).","section":"Experimental Deployment and Results, H2 analysis"},{"comment":"The H2 outcome measures (NPE and SPC Total Merged) are the very features used by the AWE feedback-selection algorithm. Students who receive 'Use more evidence' and 'Provide more details' can increase these counts by literally copying more phrases from the article, which the authors themselves note is a limitation (see the plagiarism discussion in Current and Future Directions). The significant NPE and SPC Total Merged gains are therefore partly mechanical compliance with the feedback, not evidence of improved evidentiary reasoning. The human-rubric result (H1) is not circular, but it is only trending (p ≤ 0.08). The paper should present H1 as the primary test of learning, treat H2 as a manipulation-check or compliance measure, and add an analysis that distinguishes, for example, copied phrases from paraphrased evidence, or that uses the human rubric's Relevance and Specificity subcomponents rather than the NLP counts.","section":"Experimental Deployment and Results"},{"comment":"The only statistically significant human-score improvement appears in the subgroup of 45 students receiving feedback messages 1 and 2 (2.33 to 2.64, p = 0.02), while the other two subgroups (messages 2 and 3; messages 3 and 4) show no improvement. Because this subgroup is selected by the feedback algorithm on the basis of the students' low initial evidence usage, the improvement may be inflated by regression to the mean. The paper should report effect sizes (e.g., Cohen's d) for the overall and subgroup comparisons, and should discuss regression-to-the-mean as a plausible alternative explanation for the subgroup result, especially given the overall p-value of 0.08.","section":"Experimental Deployment and Results"}],"minor_comments":[{"comment":"The column headers 'SCP5' through 'SCP8' appear to be typos for 'SPC5' through 'SPC8'; the text consistently uses 'SPC' elsewhere.","section":"AWE Feedback Selection"},{"comment":"The function 'RN D' is not defined; it should be stated as rounding to the nearest integer.","section":"AWE Feedback Selection"},{"comment":"The text 'RT AM V P' has unusual spacing; it should read 'RTA_MVP' (and similarly for 'RTAspace' in Current and Future Directions).","section":"System Usage and Architecture"}],"recommendation":"major_revision","confidential_remarks":"The paper would be a reasonable fit for the AIED/AWE community, but the central causal claim in the abstract is not supported by the current design. The authors already acknowledge the missing control condition, which suggests the revision path is to reframe the contributions as a feasibility/pilot study with explicit descriptive claims, or to include a control-arm analysis if any data is available. The circularity of H2 is a further load-bearing issue that should be addressed head-on, not merely noted as a future plagiarism concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is an honest pilot of a well-scoped automated writing evaluation system, and the authors are appropriately careful about what their data can and cannot show. The abstract's claim that eRevise 'empowers students to better revise' is causal, and the design—no control condition—cannot support that. The authors know this; they explicitly plan a generic-feedback control for the next deployment.\n\nWhat's actually new: the feedback selection algorithm based on NPE and a new SPC Total Merged feature is a reasonable extension of prior AES work. Using interpretable, rubric-aligned features to trigger pre-authored feedback messages is a sensible idea and the system is well motivated. The pilot deployment in 7 classrooms with 143 students is a real effort, and the authors report the messy details: ceiling effects, a majority of students showing no change, a non-significant human-rubric improvement (p=0.08), and null results on the four other RTA dimensions. That transparency earns credit.\n\nSoft spots: the no-control issue is not hypothetical. The observed pre-post gains on NPE and SPC Total Merged are partly mechanical, since those are exactly the features used to assign feedback; students who follow 'use more evidence' and 'be specific' will move those counts. The human rubric gain is only trending, and the subgroup analysis suggests the effect is concentrated in the weakest essays, which is compatible with regression to the mean. The lookup table was tuned on a development corpus, so its thresholds are fitting, not prediction. The authors acknowledge these limitations in the paper, which is good, but the abstract overstates what the data can establish.\n\nThe feature engineering is straightforward and reproducible, and the citation pattern is appropriate. This is a good pilot paper, not a definitive evaluation. A serious reviewer should send it out, but with a request to soften the causal language and frame the results as descriptive. The controlled follow-up is what will make the claim stick.\n\nI'd bring this to a reading group on learning technologies. It deserves peer review.","headline":"A transparent pilot of a well-designed AWE system, but the abstract's causal claim outruns the no-control data; the planned control condition will settle it.","tokens_in":10097,"tokens_out":2358,"would_cite":true,"duration_ms":26215,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that NLP-selected formative feedback from rubric-based essay scoring helps fifth- and sixth-grade students improve their use of text evidence when revising drafts.","keywords":["eRevise","automated writing evaluation","formative feedback","text evidence","automatic essay scoring","response to text assessment","natural language processing","revision"],"falsifier":"A randomized trial comparing three revision conditions—eRevise adaptive feedback, generic feedback, and no feedback—would settle it. If students who revise with generic feedback or no feedback improve their evidence scores as much as students who receive eRevise's adaptive messages, the central claim fails.","tokens_in":9054,"feed_emoji":"📝","tokens_out":5854,"duration_ms":52681,"temperature":0.7,"pith_summary":"This paper describes eRevise, a web-based writing environment that automatically selects two of four formative feedback messages for upper elementary students based on natural-language-processing features of their first drafts. The central claim is that this adaptive feedback improves students' use of text evidence when they revise. In a pilot deployment across seven fifth- and sixth-grade classrooms, 143 students' revised essays received higher human rubric scores for evidence use than their first drafts, and the NLP measures of evidence quantity and specificity improved significantly. The paper argues that NLP-driven formative feedback can make a substantive writing dimension—using text evidence from a source—more learnable, without requiring extra teacher time.","feed_headline":"Automatic feedback helps 5th and 6th graders use more text evidence","feed_subtitle":"143 students wrote, got NLP-selected feedback, and revised; evidence quality rose on both human and NLP measures.","key_machinery":"The load-bearing component is the feedback selection pipeline built on the SG automatic essay scoring model. The model represents each essay with two interpretable features: $\\mathrm{NPE}$, the number of article topics a sliding window matches via word embeddings, and $\\mathrm{SPC}$, a vector counting matched examples per category. eRevise pools $\\mathrm{SPC}$ into $\\mathrm{SPC_{total}}$, computes a duplication rate $\\mathrm{DR}$ from merged unique examples, derives $\\mathrm{SPC_{AWE}}$ for four primary topics, and buckets it into low, medium, or high. A lookup table built by content experts maps each pair of $\\mathrm{NPE}$ and $\\mathrm{SPC_{lmh}}$ to two of four predefined feedback messages addressing evidence quantity, specificity, explanation, and elaboration. These messages are what students see during revision.","core_discovery":"The paper reports that eRevise's rubric-aligned scoring features, specifically the number of evidence topics ($\\mathrm{NPE}$) and a duplication-adjusted count of specific article examples ($\\mathrm{SPC_{AWE}}$), can be converted into a small set of targeted feedback messages. In the pilot, human-scored Evidence quality improved from a mean of 2.62 to 2.72 ($p \\le 0.08$), while the $\\mathrm{NPE}$ feature rose significantly ($p \\le 0.003$) and the merged specificity count rose significantly ($p \\le 0.001$). Students whose drafts received the two least-sophisticated feedback messages showed the largest score gains ($p = 0.02$). The paper concludes that eRevise helped students improve text evidence usage after receiving formative feedback and engaging in revision.","pith_inferences":["The same selection pipeline could be applied to other response-to-text prompts if the topical word lists are generated automatically; the paper reports pilot data-driven extraction methods that still need refinement.","Hand-designed lookup tables could be learned from scored corpora, potentially replacing the content-expert mapping between feature values and feedback messages.","If students respond to feedback by copying more article text verbatim, evidence scores may rise mechanically; the paper acknowledges eRevise currently does not detect plagiarism, so future versions would need to address this.","The ceiling-effect analysis suggests fine-grained NLP features should complement rubric scores when evaluating writing interventions in elementary classrooms."],"forward_implications":["If the result holds, a rubric-based automatic essay scoring system can double as an instructional tool, giving teachers a low-cost way to support evidence use in writing.","Students with the least sophisticated evidence use stand to benefit most, as the largest gains occurred in drafts receiving the first two feedback messages.","NLP feature counts such as $\\mathrm{NPE}$ and $\\mathrm{SPC_{Total\\,Merged}}$ can detect improvement even when holistic rubric scores appear to hit a ceiling.","Because no other RTA dimensions improved, the feedback appears dimension-specific: targeting evidence use improves evidence use, not writing quality in general."],"supporting_citations":[{"why":"Defines the Response to Text Assessment and the Evidence scoring rubric used as the human outcome measure.","marker":"Correnti et al. 2013"},{"why":"Develops the rubric-aligned AES features (NPE, SPC) and supplies the topic and example word lists for the RTA article.","marker":"Rahimi et al. 2017"},{"why":"Introduces the SG word-embedding model used as eRevise's AES component and provides the scored corpus used to design the feedback lookup table.","marker":"Zhang and Litman 2017"},{"why":"Supplies the conceptual framework for effective text evidence use that shapes the four feedback messages.","marker":"Wang, Matsumura, and Correnti 2018"},{"why":"Provides the meta-analytic evidence that drafting and revising after feedback is essential, motivating the two-period design.","marker":"Graham, Harris, and Santangelo 2015"},{"why":"Supports the premise that feedback accuracy is related to successful revision, justifying automatic feedback selection.","marker":"Chapelle, Cotos, and Lee 2015"}],"fun_headline_variants":["NLP feedback lifts evidence quality in 5th-6th grade essays","Targeted feedback lets students revise for better text evidence","eRevise: NLP feedback sharpens student essay evidence","Even minimal feedback improves evidence use in student writing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The improvements are attributed to the adaptive feedback itself, but the study has no control condition, so revising with any feedback—or simply revising—might have produced the same gains.","fun_headline_variants_meta":{"raw":{"variants":["NLP feedback lifts evidence quality in 5th-6th grade essays","Targeted feedback lets students revise for better text evidence","eRevise: NLP feedback sharpens student essay evidence","Even minimal feedback improves evidence use in student writing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1253,"prompt_tokens":823,"completion_tokens":430,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":362}},"tokens_in":439,"tokens_out":430,"duration_ms":64096,"temperature":1.0,"reasoning_tokens":362,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:56:41.477567+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized trial comparing three revision conditions—eRevise adaptive feedback, generic feedback, and no feedback—would settle it. If students who revise with generic feedback or no feedback improve their evidence scores as much as students who receive eRevise's adaptive messages, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Response to Text Assessment and the Evidence scoring rubric used as the human outcome measure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Develops the rubric-aligned AES features (NPE, SPC) and supplies the topic and example word lists for the RTA article."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the SG word-embedding model used as eRevise's AES component and provides the scored corpus used to design the feedback lookup table."},{"cited_title":"C.; and Correnti, R","cited_arxiv_id":null,"evidence_quote":"Supplies the conceptual framework for effective text evidence use that shapes the four feedback messages."},{"cited_title":"R.; and Santangelo, T","cited_arxiv_id":null,"evidence_quote":"Provides the meta-analytic evidence that drafting and revising after feedback is essential, motivating the two-period design."},{"cited_title":"A.; Cotos, E.; and Lee, J","cited_arxiv_id":null,"evidence_quote":"Supports the premise that feedback accuracy is related to successful revision, justifying automatic feedback selection."}],"review_version":1}