REVIEW 4 major objections 5 minor 11 references
Product vs. Process: Exploring EFL Students' Editing of AI-Generated Text for Expository Writing
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that, for 39 Hong Kong EFL students writing expository articles with AI chatbots, the volume of AI-generated words in the final essay positively predicted human-rated quality on every scoring dimension, while most…
desk verdict Real new process data on how EFL students edit AI text, but the product-side headline is likely a length effect; the paper deserves review but needs a total-length control. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a two-perspective coding and regression design: qualitative coding of every human insert or AI-text manipulation into six expository units (title, heading, topic sentence, supporting sentence, introductory paragraph, concluding paragraph) with an edit-scope code (short, medium, long relative to a unit), temporal sequence analysis using optimal matching to cluster students' editing trajectories, and multiple linear regression linking those counts, AI-word counts, human-word counts, and runtime to human-rated scores. The product-perspective count of AI-generated words is the variable that carries the paper's headline result.
What would settle it
Record the entire writing process from first prompt to final submission (e.g., with keystroke logging) for a similar sample of EFL students, or add a no-AI baseline essay scored with the same rubric to the regression. The central claim would be refuted if process variables—edit counts, patterns, or runtime—then predict human-rated scores over and above AI-word count, or if the AI-word coefficient drops to null once baseline writing ability is controlled.
Extended reading notes
Core claim
The central claim is that the quality of EFL students' AI-assisted expository essays, as judged by human raters, is best explained by how much AI-generated text the final essay contains, not by the editing behaviors the students performed. In the product-oriented regressions $(n=39)$, the number of AI words was a statistically significant positive partial predictor of content ($\beta=0.42$), language ($\beta=0.48$), organization ($\beta=0.40$), and total scores ($\beta=0.45$; all $p<.05$), while the seven edit-count variables mostly failed to reach significance. In the process-oriented regressions $(n=25)$, only the language model reached $p=.046$ and the total model $p=.081$, with heading edits positive and concluding-paragraph edits negative predictors. The paper also claims to identify two editing patterns from temporal sequence analysis—introduction-oriented and body-oriented—and interprets the overall pattern as evidence that significant editing effort does not necessarily improve product quality.
Load-bearing premise
The load-bearing premise is that the screen recordings fairly represent each student's editing of AI-generated text, even though the paper concedes the recordings were a convenience sample covering only the workshop segment, omitted 14 of 39 students, and may have missed editing done before submission up to two weeks later.
Editorial extensions
If this is right
- If the result is correct, human-rated quality of AI-assisted expository essays in this setting tracks AI-text volume more than the measured editing actions, so increasing AI output is what visibly moves scores.
- The edit-count variables that teachers might monitor as evidence of engagement—number of edits, runtime, unit-level revisions—are not reliable predictors of final essay quality in this sample.
- The two temporal patterns give a concrete description of how students distribute editing: many revise the opening repeatedly before moving on, while a smaller group moves to body paragraphs early.
- Because most edits were short chunks, students' effort concentrated on localized wording changes rather than structural expository units, which may explain the weak link to organization scores.
- Pedagogically, the paper's corollary is that genre instruction and process-focused assessment should precede AI integration.
Reading between the lines
- Editorial inference: the positive AI-word coefficient may partly reflect selection—students who let AI write more may also prompt and select better—so the result should not be read as proof that AI text itself causes higher scores.
- Editorial inference: a direct test would ask human raters to score the original AI-generated draft and the student's edited version separately; if edited versions do not score higher despite equivalent AI-word counts, the 'disconnect' is confirmed rather than an artifact of scoring.
- Editorial inference: the editing patterns may generalize poorly because recordings captured only the workshop window; full keystroke logging from planning to submission could reveal that most consequential editing happens outside the recorded segment.
- Editorial inference: a natural extension is to rate the semantic quality of each edit (e.g., whether it fixes an error, adds relevant content, or merely rephrases) rather than counting edits; edit quality may predict scores where edit counts do not.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a mixed-methods study of 39 Hong Kong secondary EFL students who composed expository articles with AI-chatbot assistance. The authors coded editing actions from 25 screen recordings and human-word chunks from 39 final compositions using an expository-unit/edit-scope scheme; applied temporal sequence analysis to identify two editing patterns; had compositions scored on HKDSE-style content, language, and organization rubrics; and ran multiple linear regressions from process and product perspectives. The central empirical claims are that students made many mostly short edits, that process-side editing variables have little predictive power for human-rated scores, and that in the product-perspective models the number of AI-generated words positively predicts all score dimensions.
Significance. The study addresses an under-researched and practically important question: how EFL students edit AI-generated text and whether that editing affects the quality of their final compositions. Its strengths include a genuine process-plus-product design, the use of an external HKDSE marking scheme with independent raters, inter-coder agreement procedures, and unusually candid reporting of limitations (partial recordings, small sample, exploratory regressions). The descriptive results on edit targets and scopes, and the two temporal editing patterns, are potentially useful for pedagogy and for future research. However, the main product-side claim that AI-word count predicts human-rated quality is confounded by essay length and by the sample truncation toward heavy AI use, and the process-side conclusions are based on incomplete screen recordings. The paper is a valuable empirical contribution if these issues are addressed with additional specification checks and appropriately qualified claims; in its current form the headline result needs substantial revision.
major comments (4)
- [Sec. 4.2.2, Appendix 4a, Table 5] The claim that 'the number of AI-generated words positively predicted all score dimensions' is not established because the product-perspective models enter raw AI-word and human-word counts without total essay length or AI proportion. Since total length is exactly AI words plus human words, the coefficient on AI words is estimated while holding human words constant, and with AI text averaging 83.7% of each composition (median 89.5%), the positive AI coefficient may simply reflect the well-known length-quality correlation in L2 writing rather than an effect of AI provenance. The paper does not test whether the AI-word coefficient is significantly larger than the human-word coefficient, and it does not report models with total word count and AI proportion as predictors. Without such a specification check, the data are equally consistent with 'longer essays score higher regardless of source,' and the abstract's wording is an overinterpretation.
- [Sec. 3.3, Sec. 5.4, Sec. 4.1.2] The process-side results—the two editing clusters and the regressions in Table 4 and Appendix 3a—rest on 25 screen recordings that the authors themselves describe as a convenience sample: mean length 24 minutes, range 1 minute 6 seconds to 35 minutes 22 seconds, captured only during the workshop, while students could finish compositions up to two weeks later. Ten of 39 participants submitted no recording and four recordings showed no AI edits. Because substantial editing may have occurred outside the recorded window, the identified clusters and the process-perspective predictors (e.g., Heading and ConcludingP edits) may be artifacts of partial observation rather than descriptions of students' actual editing behavior. The manuscript should either analyze how much of each composition's completion is covered by the recording, or explicitly limit all process conclusions to the recorded workshop segment rather than to 'students' editing behavior' generally.
- [Sec. 3.2, Sec. 4 (first paragraph)] The sample was restricted to students who self-reported using at least 50% AI-generated text, and the resulting compositions averaged 83.7% AI text (range 52.8% to 99.7%). This truncation removes exactly the variation needed to estimate an effect of AI-text volume: the study cannot compare high- versus low-AI use, and the general statement that 'AI supports but does not replace writing skills' is not supported by a design that excludes students who used AI sparingly. The claims should be re-framed as applying to students who rely heavily on AI-generated text, or supplemented with analyses across the full range of AI usage.
- [Sec. 4.2.1, Appendix 3a] For the process perspective, no enter-method model was statistically significant, and backward elimination was then applied across four score items with up to eight candidate predictors (n=25). Only the language score model (p = .046) and total score model (p = .081) emerged, with no correction for multiple testing. Given the small sample and the large number of candidate variables, these results are at high risk of being false positives, and the statement in Section 5.2 that the analyses 'identified fine-grain patterns of student editing' overstates the evidence. The authors should present these findings explicitly as hypothesis-generating, report the multiplicity issue, or use a more conservative selection or validation procedure.
minor comments (5)
- [Sec. 3.4.1, Table 2 note] The product-perspective 'edits' are operationalized as instances of human-word chunks in the final compositions, not as observed editing actions; this should be stated more prominently in the main text and abstract so that the 261 product 'edits' are not equated with the 266 process edits.
- [Sec. 5.3] There is a typo in the final paragraph: 'machine-involuted writing' should almost certainly read 'machine-in-the-loop writing.'
- [Table 3] The 'Process' and 'Product' column headers should include the sample sizes (n=25 and n=39) directly in the table rather than only in the surrounding text.
- [Sec. 3.4.2, Figures 9 and 10] For reproducibility, the optimal matching analysis should report the substitution and indel costs used, the clustering algorithm and criterion, and the cluster validation method; currently the description is too terse to determine whether the two-cluster solution is robust to reasonable parameter choices.
- [Sec. 5.2, Table 5] The negative associations for Heading chunks and Other chunks rest on data from only 2 of 39 students; the authors already note this in the text, but the abstract and conclusion should mirror that caution rather than presenting these associations as robust findings.
Circularity Check
No significant circularity: the regression findings are data-derived and the self-citations are methodological or interpretive, not load-bearing.
full rationale
The paper is an empirical study, and its central derivation chain is not circular. The outcome variables are human-rated content, language, organization, and total scores based on an external HKDSE marking scheme, applied by independent expert raters. The predictor variables for the product perspective are measured word counts and counts of human chunks coded from students' compositions; the process perspective uses edit counts and timestamps from screen recordings. None of these variables is defined in terms of the outcome scores, so the regression result that AI word count positively predicts scores is a substantive empirical finding rather than a tautology. The edit-scope coding scheme is refined from the authors' prior work (Woo et al., 2024b), but that citation affects only the categorization labels used for descriptive coding and is not the source of the predicted-score relationship. The suggested mechanism for the AI-word effect is also borrowed from that prior study, but it is offered as an interpretive analogy, not as a premise from which the regression results are derived. The paper explicitly labels the MLR models as exploratory, uses backward elimination with acknowledged small-sample limitations, and reports null and marginal results as well as positive ones. Validity threats such as the partial screen recordings and the potential length confound between AI words and total word count are substantive concerns about interpretation and robustness, but they are not circularity: the reported coefficients are not forced by construction, and the paper does not rename a known result as a novel derivation. There are no self-citation uniqueness theorems, no ansatz smuggled in via citation, and no fitted parameter being relabeled as an out-of-sample prediction. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (4)
- AI-text usage inclusion threshold =
50% self-reported AI text
- MLR coefficients for AI-word count =
standardized beta approximately 0.42 to 0.48 across the four score models
- Edit scope thresholds (short/medium/long) =
short less than one unit, medium equal to one unit, long more than one unit
- Optimal matching substitution and indel costs =
default TraMineR costs (insertion, deletion, substitution all equal to 1)
assumptions (4)
- domain assumption Human word chunks in final compositions proxy students' edits of AI-generated text
- domain assumption Screen recordings capture representative editing behavior
- domain assumption Human-rated HKDSE scores are valid measures of composition quality
- standard math Regression linearity and independence assumptions hold for the n=25 and n=39 models
Cite this review
Pith. "Pith review of Product vs. Process: Exploring EFL Students' Editing of AI-Generated Text for Expository Writing." pith.science (2026). https://pith.science/paper/RQ75DMI4
@misc{pith2026250721073,
author = {Pith},
title = {Pith review of: Product vs. Process: Exploring EFL Students' Editing of AI-Generated Text for Expository Writing},
year = {2026},
howpublished = {\url{https://pith.science/paper/RQ75DMI4}},
note = {Machine review of arXiv:2507.21073}
}
read the original abstract
Text generated by artificial intelligence (AI) chatbots is increasingly used in English as a foreign language (EFL) writing contexts, yet its impact on students' expository writing process and compositions remains understudied. This research examines how EFL secondary students edit AI-generated text. Exploring editing behaviors in their expository writing process and in expository compositions, and their effect on human-rated scores for content, organization, language, and overall quality. Participants were 39 Hong Kong secondary students who wrote an expository composition with AI chatbots in a workshop. A convergent design was employed to analyze their screen recordings and compositions to examine students' editing behaviors and writing qualities. Analytical methods included qualitative coding, descriptive statistics, temporal sequence analysis, human-rated scoring, and multiple linear regression analysis. We analyzed over 260 edits per dataset, and identified two editing patterns: one where students refined introductory units repeatedly before progressing, and another where they quickly shifted to extensive edits in body units (e.g., topic and supporting sentences). MLR analyses revealed that the number of AI-generated words positively predicted all score dimensions, while most editing variables showed minimal impact. These results suggest a disconnect between students' significant editing effort and improved composition quality, indicating AI supports but does not replace writing skills. The findings highlight the importance of genre-specific instruction and process-focused writing before AI integration. Educators should also develop assessments valuing both process and product to encourage critical engagement with AI text.
Reference graph
Works this paper leans on
-
[1]
iterature Review 2.1. Expository Writing Expository writing is a fundamental means by which language learners exhibit the knowledge they have acquired (Nippold, 2016). Slater and Graves (1989) define expository text as a form of writing that is informative, explanatory, and straightforward. They assert that this type of text actively involves readers by e...
work page 1989
-
[2]
Research has been conducted to explore the role of a machine-in-the-loop in enhancing EFL writing
A ‘machine-in-the-loop’ approach to writing. Research has been conducted to explore the role of a machine-in-the-loop in enhancing EFL writing. This research has primarily examined completed written works and associated outcomes. For example, Song and Song (2023) assessed the influence of ChatGPT on the IELTS writing skills of Chinese EFL students. Their ...
work page 2023
-
[3]
Methodology 3.1. Research Context 9 The study was conducted in seven Hong Kong secondary schools where English is taught as a foreign language and where students in EFL lessons are not taught expository essays but text types such as articles that include expository writing (Koh, 2015; The Curriculum Development Council, 2017). In Hong Kong mainstream educ...
work page 2015
-
[5]
Discussion 5.1. Characteristics of Students’ Edits We observed 266 edits in 25 screen recordings and 261 edits in written compositions, with the majority of edits being short chunks. The large number of edits in both process and product indicate students had put considerable effort into editing, and suggests students preferred minimal, localized modificat...
work page 2024
-
[6]
An article excerpt for Question 3 with AI-generated text in red. 13 3.2. Participants Although seven Hong Kong secondary schools were purposefully sampled for this study, the students from those schools composed an opportunistic convenience sample as each school’s teacher-in-charge was responsible for recruiting students for the workshop. Furthermore, we ...
work page 2018
-
[9]
Conclusion This study has explored EFL students’ expository writing with a machine-in-the-loop from process and product perspectives. It has evidenced different characteristics of students’ editing of AI-generated text in their expository writing and how these characteristics may or may not enhance product quality in terms of human-rated scores. Insights ...
arXiv 2023
-
[10]
35 Yang, Y., Yap, N. T., & Ali, A. M. (2023). Predicting EFL expository writing quality with measures of lexical richness. Assessing Writing, 57, 100762. https://doi.org/10.1016/j.asw.2023.100762 Zhang, R., Zou, D., Cheng, G., & Xie, H. (2024). Flow in ChatGPT-based logic learning and its influences on logic and self-efficacy in English argumentative writ...
-
[21]
Two experts independently scored each composition. One of the experts was the EFL teacher-in-charge from a participating school, and the other was the first author who was also an EFL teacher-in-charge. These experts’ scores were averaged to arrive at a final content, language, organization and total score for the composition. 3.4.4. Multiple Linear Regre...
work page 2004
Show all 11 references
-
[57]
School banding
https://doi.org/10.1186/s41239-023-00425-2 Fang, Z. (2014). Writing a report: A study of preadolescents’ use of informational language. Linguistics and the Human Sciences, 10(2), 103–131. https://doi.org/10.1558/lhs.v10.2.28556 Flower, L., & Hayes, J. R. (1981). A Cognitive Pr...
2014
-
[2006]
To prepare data for coding, we noted to which school, form level, and student each written composition and screen recording belonged
by implementing a hierarchical coding scheme, and presenting descriptive statistics. To prepare data for coding, we noted to which school, form level, and student each written composition and screen recording belonged. For the latter, we also noted the total runtime. To explor...
2024
-
[2010]
and refining the categorization system used by Woo et al. (2024b), we designed three codes to differentiate between minor, localized edits and more substantial, unit-level edits: (1) short edits: less than an expository writing unit (e.g., a word in a title; words in a topic s...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.