REVIEW 3 major objections 5 minor 40 references
AI in the Writing Process: How Purposeful AI Support Fosters Student Writing
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper argues that when a large language model is embedded into structured writing stages rather than delivered through a chat window, undergraduate writers report more agency and show more markers of deep knowledge transformation.
desk verdict Solid RCT with a strong agency effect, but the abstract overclaims knowledge transformation—only one of five codes holds up, and one 'significant' code is actually knowledge telling. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the design of Script&Shift: a layered interface that separates content from rhetorical organization and lets writers summon specialized AI assistants for brainstorming, detail elaboration, audience analysis, and feedback instead of a single free-form chat. The measurement machinery is an essay-coding framework that classifies each sentence as knowledge transformation (synthesis, analysis, application, evaluation, comprehension) versus knowledge telling, combined with seven-point self-report items for agency. The experimental machinery is a three-condition randomized design with 30 participants per condition and the same LLM backend in the two AI conditions.
What would settle it
Take the same source-based writing task and compare Script&Shift against a chat condition whose usability and output quality are carefully matched (same underlying LLM, pre-tested ease-of-use ratings, and document-integrated suggestions rather than a blank separate window); if the agency and knowledge-transformation gaps vanish under that match, the claim that the interface paradigm causes them is refuted.
Extended reading notes
Core claim
The study's central claim is that process-oriented AI support, where the LLM is embedded into named writing subprocesses within a layered document interface, gives undergraduate writers greater felt agency and more markers of deep knowledge transformation than a conventional chat-based LLM assistant. In the randomized comparison, Script&Shift outperformed the chat condition on Analysis markers, Evaluation markers, and the Knowledge category, and it outperformed both the chat and standard conditions on three self-reported agency items: feeling in control, feeling content with the essay, and willingness to publish under one's own name. The paper is careful to note that final essay-quality scores did not differ reliably, and it treats that as consistent with prior findings that knowledge-transformation markers and grades are only weakly related.
Load-bearing premise
The result rests on the assumption that differences came from the interface paradigm rather than from the particular quality, usability, or familiarity of the custom chat tool, and the paper's own limitations section notes that the single 1.5-hour session may have restricted engagement.
Editorial extensions
If this is right
- Designing AI writing support around explicit writing subprocesses is a way to keep students in charge of their essays while still getting LLM help.
- Chat-style assistants can be expected to show larger agency costs, and adding structured scaffolding may be more valuable than swapping the underlying model.
- Process gains such as knowledge transformation can occur even when final essay scores do not improve, so outcome measures that rely only on grades will miss them.
- Because participants felt more agency with Script&Shift than with no AI at all, AI assistance need not be framed as a trade-off against ownership.
- Longer or repeated use would be required to see whether the knowledge-transformation markers translate into better final essays.
Reading between the lines
- An untested implication is that chat's disadvantage comes from context-switching and copy-paste overhead; a follow-up that gives the chat condition inline document integration would isolate that mechanism.
- If the tool's benefit is metacognitive rather than merely structural, the effect should survive a transfer task without Script&Shift; the single-session design cannot establish that.
- The larger agency gap over the no-AI control hints that 'control' may mean having a visible, structured process, not just the absence of automation; keystroke-level authorship analysis would test that reading.
- The null essay-quality result implies grade-based evaluations will not capture the process gains claimed here; portfolios or process-trace measures would be needed to see them.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a three-arm randomized controlled trial (N=90 undergraduates) comparing a process-oriented integrated AI writing tool (Script&Shift), a custom chat-based LLM assistant, and a standard no-AI interface on a source-based argumentative writing task. It measures writer agency via post-test self-report and knowledge transformation via rubric-based coding of final essays. The paper reports that Script&Shift users experienced greater agency and "deeper knowledge transformation overall," and uses these results to argue that LLM support embedded in structured writing stages preserves ownership and deepens engagement. The core recommendation is major revision because the overall knowledge-transformation claim is not supported by the reported analyses.
Significance. The study has genuine strengths: a randomized design, a reasonably large sample for an interface experiment, external coding of essays with inter-rater reliability (kappa=0.76 initially, kappa=0.86 after discussion), and very large agency effects (e.g., Q1 F(2,85)=101.90). If the agency finding is taken at face value, it is a meaningful contribution to the HCI and learning-sciences literature on AI-assisted writing. However, the paper's central knowledge-transformation claim currently rests on one significant subcode (Analysis), and the "Knowledge" code is direction-inverted relative to the theoretical construct; the paper's contribution after revision would be the agency result plus a carefully bounded analytical-process claim, not "deeper knowledge transformation overall." The authors should also verify the fairness of the chat baseline, since the interface-paradigm interpretation depends on it.
major comments (3)
- [Section 4.1 and Figure 2] The abstract's claim of "deeper knowledge transformation overall" (and the Section 1 claim that process-oriented tools produce "more markers of knowledge transformation") is not supported by the reported statistics. Only Analysis shows a significant omnibus Kruskal-Wallis test (H=8.72, p=.013); Synthesis (H=2.75, p=.253), Application, and Comprehension are non-significant, and no omnibus test is reported for Evaluation or Knowledge. The pairwise Script&Shift advantages on Evaluation (U=287.50, p=.033) and Knowledge (U=554.00, p=.038) are only against the Chat condition and receive no multiple-comparison correction. More seriously, Figure 2 defines the Knowledge code as "Paraphrased/copied information from a source" under the "Knowledge Telling" category; a higher count of this code is evidence of knowledge telling, not knowledge transformation, so the sentence in Section 4.1 reporting "significantly better performance in Knowledge" as a positive outcome is direction-inverted. Because Synthesis, the code most aligned with the paper's own definition of transformation as integrating information from multiple sources, did not differ across conditions, the overall deeper-knowledge-transformation claim cannot be derived from the data and should be revised to name Analysis as the only code with a significant omnibus effect.
- [Section 3] The study is framed as a test of interface paradigm, but the chat condition is described only as "a custom chat-based LLM to support writing" with no feature inventory, no usability validation, and no evidence that it was a fair or non-strawman implementation. If the chat tool was less usable, unfamiliar, or missing basic affordances (e.g., document integration or copy-paste support), the Script&Shift advantage could reflect implementation quality rather than the process-oriented design. Please report the chat interface's features, any pilot testing, and post-task usability ratings or interface logs that establish the chat condition as a competent baseline before attributing the results to interaction paradigm.
- [Section 4.1] Multiple-testing protection is missing. Five Kruskal-Wallis tests are run, and the only significant one (Analysis, p=.013) no longer survives a simple Bonferroni correction for five tests (adjusted p approximately .065). The pairwise Mann-Whitney comparisons that follow are likewise uncorrected and are reported for Evaluation and Knowledge without a significant omnibus. The agency results in Section 4.2 are protected by omnibus ANOVAs and are credible; the knowledge-transformation section should be held to the same standard, or explicitly labeled exploratory.
minor comments (5)
- [Section 4.2] The text describes a p < .05 result as "marginally higher," but "marginal" is usually reserved for .05 < p < .10; also, the sentence "For question 2..." is confusing because the immediately preceding sentence discusses a different item. Please label the two AI-agency items unambiguously.
- [Section 3.1] The demographic reporting ("mode of age range was 18-39," "median age range was 26-30") is unstandardized; report age distribution with mean and standard deviation or binned counts.
- [Section 4.4 and Figure 5] The text says "Analysis revealed a moderate positive correlation" but the figure caption specifies that the correlation is for Script&Shift while the chat-based condition shows no correlation; state the subgroup and sample size for r=0.61, and clarify whether the measure is interface transactions or LLM feature access.
- [Section 4.3] No inferential test is reported for essay quality or source integration; the discussion of "high variance indicates considerable group overlap" would benefit from a formal comparison (e.g., Welch's ANOVA or Kruskal-Wallis) and effect sizes.
- [Throughout Section 4] The paper does not report effect sizes or confidence intervals for the significant tests; adding them would help readers judge the magnitude of the agency and Analysis effects.
Circularity Check
No significant circularity: the central claim is an empirical RCT comparison whose outcomes are measured by an external coding rubric and self-report, not by the tool's definition or by a fitted parameter.
full rationale
The paper's central claim, that Script&Shift increases writer agency and knowledge transformation relative to chat and standard interfaces, rests on independent measurements rather than on the tool's construction. Knowledge transformation is operationalized via qualitative coding using the external framework of Raković et al. [28], and agency is assessed through post-test Likert-scale self-reports. The tool's design is introduced through the authors' prior work [33], but no outcome measure is defined in terms of Script&Shift's features, and no parameter is fitted to the data such that the observed differences are forced. The only substantive self-citation, reference [33], is used to describe the tool and to interpret feature-usage patterns, neither of which constitutes the load-bearing derivation of the agency or knowledge-transformation results. Concerns about the 'Knowledge' code being classified under knowledge telling while also counted as a knowledge-transformation marker, and about uncorrected multiple comparisons, are validity and statistical-inference issues rather than circularity: the coding rubric and survey items are external to the present paper and do not reduce to the conclusion by construction. The derivation chain is therefore self-contained with respect to circularity, even though the strength of the evidence is weakened by measurement and analysis limitations.
Assumptions & free parameters
free parameters (1)
- Knowledge transformation tercile cutoffs =
not reported
assumptions (4)
- domain assumption The Raković et al. coding scheme validly operationalizes knowledge transformation as distinct from knowledge telling.
- domain assumption The three self-report Likert questions (Q1-Q3) measure writer agency as a construct.
- domain assumption Random assignment created comparable groups without baseline writing ability differences.
- domain assumption The custom chat-based interface is representative of chat-based LLM writing assistants.
Cite this review
Pith. "Pith review of AI in the Writing Process: How Purposeful AI Support Fosters Student Writing." pith.science (2026). https://pith.science/paper/FUFZL43J
@misc{pith2026250620595,
author = {Pith},
title = {Pith review of: AI in the Writing Process: How Purposeful AI Support Fosters Student Writing},
year = {2026},
howpublished = {\url{https://pith.science/paper/FUFZL43J}},
note = {Machine review of arXiv:2506.20595}
}
read the original abstract
The ubiquity of technologies like ChatGPT has raised concerns about their impact on student writing, particularly regarding reduced learner agency and superficial engagement with content. While standalone chat-based LLMs often produce suboptimal writing outcomes, evidence suggests that purposefully designed AI writing support tools can enhance the writing process. This paper investigates how different AI support approaches affect writers' sense of agency and depth of knowledge transformation. Through a randomized control trial with 90 undergraduate students, we compare three conditions: (1) a chat-based LLM writing assistant, (2) an integrated AI writing tool to support diverse subprocesses, and (3) a standard writing interface (control). Our findings demonstrate that, among AI-supported conditions, students using the integrated AI writing tool exhibited greater agency over their writing process and engaged in deeper knowledge transformation overall. These results suggest that thoughtfully designed AI writing support targeting specific aspects of the writing process can help students maintain ownership of their work while facilitating improved engagement with content.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
In: Proceedings of the 16th Conference on Creativity & Cognition
Anderson, B.R., Shah, J.H., Kreminski, M.: Homogenization effects of large lan- guage models on human creative ideation. In: Proceedings of the 16th Conference on Creativity & Cognition. pp. 413–425 (2024)
work page 2024
-
[2]
https://www.anthropic.com (2025), accessed: 2025-01-22
Anthropic: Claude. https://www.anthropic.com (2025), accessed: 2025-01-22
work page 2025
-
[3]
University of Texas Press (2010)
Bakhtin, M.M.: The dialogic imagination: Four essays. University of Texas Press (2010)
work page 2010
-
[4]
Available at SSRN4895486 (2024)
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, O., Mariman, R.: Generative AI can harm learning. Available at SSRN4895486 (2024)
work page 2024
-
[5]
Bereiter, C., Scardamalia, M.: The Psychology of Written Composition. Routledge (2013)
work page 2013
-
[6]
Learning Lives: Learning, Identity, and Agency in the Life Course pp
Biesta, G., Tedder, M.: How is agency possible? towards an ecological understand- ing of agency-as-achievement. Learning Lives: Learning, Identity, and Agency in the Life Course pp. 132–149 (2006)
work page 2006
-
[7]
but it’s still not an A+ student
Bowman, E.: A new AI chatbot might do your homework for you. but it’s still not an A+ student. NPR (December 2022), https://www.npr.org/2022/12/19/1143912956/chatgpt-ai-chatbot-homework- academia
work page 2022
-
[8]
In: Gregg, L.W., Steinberg, E.R
Collins, A., Gentner, D.: A framework for a cognitive theory of writing. In: Gregg, L.W., Steinberg, E.R. (eds.) Cognitive Processes in Writing, pp. 51–72. Erlbaum (1980)
work page 1980
Show all 40 references
-
[9]
In: International Conference on Artificial Intelligence in Education
Cummings, R.E., Le, T., Shrestha, S., Smith, C.V.: The unexpected effects of google smart compose on open-ended writing tasks. In: International Conference on Artificial Intelligence in Education. pp. 468–481. Springer (2024) AI in the Writing Process 13
2024
-
[10]
Education and Information Technologies 24(2), 1563–1581 (2019)
Dahlström, H.: Digital writing tools from the student perspective: Access, affor- dances, and agency. Education and Information Technologies 24(2), 1563–1581 (2019)
2019
-
[11]
Emirbayer,M.,Mische,A.:Whatisagency?AmericanJournalofSociology 103(4), 962–1023 (1998)
1998
-
[12]
College Composition and Communication 32(4), 365–387 (1981)
Flower, L., Hayes, J.R.: A cognitive process theory of writing. College Composition and Communication 32(4), 365–387 (1981)
1981
-
[13]
arXiv preprint arXiv:2409.13686 (2024)
Geng, M., Chen, C., Wu, Y., Chen, D., Wan, Y., Zhou, P.: The impact of large language models in academia: From writing to speaking. arXiv preprint arXiv:2409.13686 (2024)
2024 arXiv
-
[14]
In: Proceedings of the 2022 ACM Designing Interactive Systems Conference
Gero, K.I., Liu, V., Chilton, L.: Sparks: Inspiration for science writing using lan- guage models. In: Proceedings of the 2022 ACM Designing Interactive Systems Conference. pp. 1002–1019 (2022)
2022
-
[15]
In: International Conference on Artificial Intelligence in Educa- tion
Gubelmann, R., Burkhard, M., Ivanova, R.V., Niklaus, C., Bermeitinger, B., Hand- schuh, S.: Exploring the usefulness of open and proprietary llms in argumentative writing support. In: International Conference on Artificial Intelligence in Educa- tion. pp. 175–182. Springer (2024)
2024
-
[16]
In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems
Hoque, M.N., Mashiat, T., Ghai, B., Shelton, C.D., Chevalier, F., Kraus, K., Elmqvist, N.: The HaLLMark effect: Supporting provenance and transparent use of large language models in writing with interactive visualization. In: Proceedings of the 2024 CHI Conference on Human Fac...
2024
-
[17]
Assessment & Evaluation in Higher Education47(8), 1301–1316 (2022)
Kim, M.K., Kim, N.J., Heidari, A.: Learner experience in artificial intelligence- scaffolded argumentation. Assessment & Evaluation in Higher Education47(8), 1301–1316 (2022)
2022
-
[18]
Lee, M., Gero, K.I., Chung, J.J.Y., Shum, S.B., Raheja, V., Shen, H., Venugopalan, S., Wambsganss, T., Zhou, D., Alghamdi, E.A., et al.: A design space for intelligent andinteractivewritingassistants.In:ProceedingsoftheCHIConferenceonHuman Factors in Computing Systems. pp. 1–35 (2024)
2024
-
[19]
In: The Routledge Handbook of Discourse Processes, pp
McNamara, D.S., Allen, L.K.: Toward an integrated perspective of writing as a discourse process. In: The Routledge Handbook of Discourse Processes, pp. 362–
-
[20]
Behavior Research Methods, Instruments, & Computers 36(2), 222–233 (2004)
McNamara, D.S., Levinstein, I.B., Boonthum, C.: istart: Interactive strategy train- ing for active reading and thinking. Behavior Research Methods, Instruments, & Computers 36(2), 222–233 (2004)
2004
-
[21]
BioData Mining 16(1), 20 (2023)
Meyer, J.G., Urbanowicz, R.J., Martin, P.C., O’Connor, K., Li, R., Peng, P.C., Bright, T.J., Tatonetti, N., Won, K.J., Gonzalez-Hernandez, G., et al.: Chatgpt and large language models in academia: opportunities and challenges. BioData Mining 16(1), 20 (2023)
2023
-
[22]
Mieczkowski, H., Hancock, J.: Examining agency, expertise, and roles of ai systems in ai-mediated communication (Jul 2022), osf.io/asnv4_v1
2022
-
[23]
In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
Mirowski, P., Mathewson, K.W., Pittman, J., Evans, R.: Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals. In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. pp. 1–34 (2023)
2023
-
[24]
Science381(6654), 187–192 (2023)
Noy, S., Zhang, W.: Experimental evidence on the productivity effects of generative artificial intelligence. Science381(6654), 187–192 (2023)
2023
-
[25]
International Journal of Artificial Intelligence in Education pp
Parker, J.L., Richard, V.M., Acabá, A., Escoffier, S., Flaherty, S., Jablonka, S., Becker, K.P.: Negotiating meaning with machines: AI’s role in doctoral writing pedagogy. International Journal of Artificial Intelligence in Education pp. 1–21 (2024) 14 M. Siddiqui et al
2024
-
[26]
AI & SOCIETY (2025)
Peterson, A.J.: AI and the problem of knowledge collapse. AI & SOCIETY (2025)
2025
-
[27]
Prolific: Prolific—participant recruitment for online research (2025), https://www.prolific.com/, accessed: 2025-01-22
2025
-
[28]
In: International Learning Analytics & Knowledge Conference
Rakovic, M., Marzouk, Z., Chang, D., Winne, P.H.: Towards knowledge- transforming in writing argumentative essays from multiple sources: A method- ological approach. In: International Learning Analytics & Knowledge Conference
-
[29]
Journal of Computer Assisted Learning37(4), 903–924 (2021)
Rakovic, M., Winne, P.H., Marzouk, Z., Chang, D.: Automatic identification of knowledge-transforming content in argument essays developed from multiple sources. Journal of Computer Assisted Learning37(4), 903–924 (2021)
2021
-
[30]
Advances in Applied Psycholinguistics2, 142–175 (1987)
Scardamalia, M., Bereiter, C.: Knowledge telling and knowledge transforming in written composition. Advances in Applied Psycholinguistics2, 142–175 (1987)
1987
-
[31]
Cognitive Science8(2), 173–190 (1984)
Scardamalia, M., Bereiter, C., Steinbach, R.: Teachability of reflective processes in written composition. Cognitive Science8(2), 173–190 (1984)
1984
-
[32]
L1-Educational Studies in Language And Literature 4, 5–33 (2004)
Segev-Miller, R.: Writing from sources: The effect of explicit instruction on col- lege students’ processes and products. L1-Educational Studies in Language And Literature 4, 5–33 (2004)
2004
-
[33]
In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems
Siddiqui, M.N., Pea, R.D., Subramonyam, H.: Script&Shift: A layered interface paradigm for integrating content development and rhetorical strategy with llm writing assistants. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. pp. 1–19 (2025)
2025
-
[34]
arXiv preprint arXiv:2502.12447 (2025)
Singh, A., Taneja, K., Guan, Z., Ghosh, A.: Protecting human cognition in the age of ai. arXiv preprint arXiv:2502.12447 (2025)
2025
-
[35]
ACM Transactions on Computer-Human Interaction30(5), 1–57 (2023)
Singh, N., Bernal, G., Savchenko, D., Glassman, E.L.: Where to hide a stolen elephant: Leaps in creative writing with multimodal machine intelligence. ACM Transactions on Computer-Human Interaction30(5), 1–57 (2023)
2023
-
[36]
Computers & Education131, 33–48 (2019)
Strobl, C., Ailhaud, E., Benetos, K., Devitt, A., Kruse, O., Proske, A., Rapp, C.: Digital support for academic writing: A review of technologies and pedagogies. Computers & Education131, 33–48 (2019)
2019
-
[37]
International Journal of Artificial Intelligence in Education pp
Tarchi, C., Zappoli, A., Casado Ledesma, L., Brante, E.W.: The use of chatgpt in source-based writing tasks. International Journal of Artificial Intelligence in Education pp. 1–21 (2024)
2024
-
[38]
arXiv preprint arXiv:2112.04359 (2021)
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., et al.: Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359 (2021)
2021 arXiv
-
[39]
International Journal of Artificial Intelligence in Education 28, 106–137 (2018)
Weston-Sementelli,J.L., Allen,L.K.,McNamara,D.S.:Comprehensionandwriting strategy training improves performance on content-specific source-based writing tasks. International Journal of Artificial Intelligence in Education 28, 106–137 (2018)
2018
-
[2019]
pp. 267–272. Society for Learning Analytics Research (2019)
2019
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.