{"id":"57adb653-6cd4-4838-b888-6480c741507a","arxiv_id":"2506.20595","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In 90 undergraduates, a process-oriented AI writing tool increased self-reported agency and several coded markers of knowledge transformation compared with a chat-based LLM interface.","lead":"A randomized trial with 90 undergraduates compared a process-oriented AI writing tool, a chat-based AI assistant, and a plain editor. The integrated tool raised students' reported sense of control and some markers of deep source engagement, but essay quality did not differ across groups.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract claims 'deeper knowledge transformation overall,' but Section 4.1 reports only one of five codes (Analysis) significant against both controls; Evaluation and 'Knowledge' are only significant versus Chat, and the 'Knowledge' code is defined as knowledge-telling, not transformation.","rationale":"The central claim has two limbs: greater agency and deeper knowledge transformation. The agency limb is supported by large, consistent differences, so it survives scrutiny even under the self-report caveat. The knowledge-transformation limb is the paper's more novel contribution, but Section 4.1 does not establish it 'overall.' Only Analysis is significant against both baselines; Evaluation differs only from Chat; and the 'Knowledge' result is reported as if it favors Script&Shift even though Figure 2 codes 'Knowledge' as knowledge-telling (paraphrasing/copying). That is not a disagreement with the field—it is an internal contradiction between the coding scheme and the interpretation. Because the abstract claims 'overall' deeper knowledge transformation, this is the load-bearing point: if the reanalysis fails to show an omnibus effect, the paper's headline must be downgraded to a narrower claim. The reader identified the broader theme of overstatement, hence partial agreement; but the reader's weakest assumption focused on the chat-baseline confound, which is an external-validity issue and not as decisive here. The proposed reanalysis is cheap and definitive, so the verdict should remain CONDITIONAL rather than REJECT.","tokens_in":9992,"tokens_out":3826,"duration_ms":40920,"concrete_test":"Reanalyze the five knowledge-transformation codes from Section 4.1 as a single multivariate outcome (e.g., MANOVA on ranked scores) and apply a conservative correction (Benjamini-Hochberg) across the omnibus tests and the pairwise comparisons. Then reverse-code the 'Knowledge' category as a knowledge-telling marker per Figure 2 and re-check the comparison with ChatLLM (U = 554.00, p = .038): if the difference actually favors more knowledge telling in Script&Shift, the 'better performance in Knowledge' claim reverses. If after correction the only surviving result is Analysis (S&S vs Standard), the abstract's 'overall' knowledge-transformation claim should be replaced with a claim about a single analytical subcode.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 4.1, the authors run Kruskal-Wallis tests per code. Only Analysis yields a significant omnibus (H = 8.72, p = .013). Pairwise tests then yield Script&Shift > Standard for Analysis only (p = .007), and > Chat for Analysis (p = .020), Evaluation (p = .033), and 'Knowledge' (p = .038). No omnibus test is reported for Evaluation or Knowledge, so those pairwise p-values lack the protection that a significant omnibus would provide. With five KT codes and multiple post-hoc comparisons, no correction is applied anywhere. More importantly, Figure 2 defines the 'Knowledge' code under Knowledge Telling as 'Paraphrased/copied information from a source.' A higher Knowledge count is therefore evidence of knowledge telling, not knowledge transformation. The paper nonetheless reports Script&Shift 'significantly better performance in Knowledge' (p = .038), which inverts the direction of the construct. Synthesis—the code most directly aligned with the paper's own definition of knowledge transformation as integrating information from multiple sources—shows no significant difference (H = 2.75, p = .253). Thus the abstract's 'deeper knowledge transformation overall' is not derivable from the reported statistics; the only stable effect is on Analysis, a single subcode. This is not an external validity quibble: it is an internal consistency problem between the paper's central claim and its own Figure 2 and Section 4.1.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a three-arm randomized controlled trial (N=90 undergraduates) comparing a process-oriented integrated AI writing tool (Script&Shift), a custom chat-based LLM assistant, and a standard no-AI interface on a source-based argumentative writing task. It measures writer agency via post-test self-report and knowledge transformation via rubric-based coding of final essays. The paper reports that Script&Shift users experienced greater agency and \"deeper knowledge transformation overall,\" and uses these results to argue that LLM support embedded in structured writing stages preserves ownership and deepens engagement. The core recommendation is major revision because the overall knowledge-transformation claim is not supported by the reported analyses.","tokens_in":10296,"tokens_out":6506,"duration_ms":72261,"significance":"The study has genuine strengths: a randomized design, a reasonably large sample for an interface experiment, external coding of essays with inter-rater reliability (kappa=0.76 initially, kappa=0.86 after discussion), and very large agency effects (e.g., Q1 F(2,85)=101.90). If the agency finding is taken at face value, it is a meaningful contribution to the HCI and learning-sciences literature on AI-assisted writing. However, the paper's central knowledge-transformation claim currently rests on one significant subcode (Analysis), and the \"Knowledge\" code is direction-inverted relative to the theoretical construct; the paper's contribution after revision would be the agency result plus a carefully bounded analytical-process claim, not \"deeper knowledge transformation overall.\" The authors should also verify the fairness of the chat baseline, since the interface-paradigm interpretation depends on it.","major_comments":[{"comment":"The abstract's claim of \"deeper knowledge transformation overall\" (and the Section 1 claim that process-oriented tools produce \"more markers of knowledge transformation\") is not supported by the reported statistics. Only Analysis shows a significant omnibus Kruskal-Wallis test (H=8.72, p=.013); Synthesis (H=2.75, p=.253), Application, and Comprehension are non-significant, and no omnibus test is reported for Evaluation or Knowledge. The pairwise Script&Shift advantages on Evaluation (U=287.50, p=.033) and Knowledge (U=554.00, p=.038) are only against the Chat condition and receive no multiple-comparison correction. More seriously, Figure 2 defines the Knowledge code as \"Paraphrased/copied information from a source\" under the \"Knowledge Telling\" category; a higher count of this code is evidence of knowledge telling, not knowledge transformation, so the sentence in Section 4.1 reporting \"significantly better performance in Knowledge\" as a positive outcome is direction-inverted. Because Synthesis, the code most aligned with the paper's own definition of transformation as integrating information from multiple sources, did not differ across conditions, the overall deeper-knowledge-transformation claim cannot be derived from the data and should be revised to name Analysis as the only code with a significant omnibus effect.","section":"Section 4.1 and Figure 2"},{"comment":"The study is framed as a test of interface paradigm, but the chat condition is described only as \"a custom chat-based LLM to support writing\" with no feature inventory, no usability validation, and no evidence that it was a fair or non-strawman implementation. If the chat tool was less usable, unfamiliar, or missing basic affordances (e.g., document integration or copy-paste support), the Script&Shift advantage could reflect implementation quality rather than the process-oriented design. Please report the chat interface's features, any pilot testing, and post-task usability ratings or interface logs that establish the chat condition as a competent baseline before attributing the results to interaction paradigm.","section":"Section 3"},{"comment":"Multiple-testing protection is missing. Five Kruskal-Wallis tests are run, and the only significant one (Analysis, p=.013) no longer survives a simple Bonferroni correction for five tests (adjusted p approximately .065). The pairwise Mann-Whitney comparisons that follow are likewise uncorrected and are reported for Evaluation and Knowledge without a significant omnibus. The agency results in Section 4.2 are protected by omnibus ANOVAs and are credible; the knowledge-transformation section should be held to the same standard, or explicitly labeled exploratory.","section":"Section 4.1"}],"minor_comments":[{"comment":"The text describes a p < .05 result as \"marginally higher,\" but \"marginal\" is usually reserved for .05 < p < .10; also, the sentence \"For question 2...\" is confusing because the immediately preceding sentence discusses a different item. Please label the two AI-agency items unambiguously.","section":"Section 4.2"},{"comment":"The demographic reporting (\"mode of age range was 18-39,\" \"median age range was 26-30\") is unstandardized; report age distribution with mean and standard deviation or binned counts.","section":"Section 3.1"},{"comment":"The text says \"Analysis revealed a moderate positive correlation\" but the figure caption specifies that the correlation is for Script&Shift while the chat-based condition shows no correlation; state the subgroup and sample size for r=0.61, and clarify whether the measure is interface transactions or LLM feature access.","section":"Section 4.4 and Figure 5"},{"comment":"No inferential test is reported for essay quality or source integration; the discussion of \"high variance indicates considerable group overlap\" would benefit from a formal comparison (e.g., Welch's ANOVA or Kruskal-Wallis) and effect sizes.","section":"Section 4.3"},{"comment":"The paper does not report effect sizes or confidence intervals for the significant tests; adding them would help readers judge the magnitude of the agency and Analysis effects.","section":"Throughout Section 4"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the main issue is the mismatch between the abstract and the results in Section 4.1. The stress-test concern lands squarely on the manuscript. I would ask the authors to either (a) revise the knowledge-transformation claim to a single-code result, or (b) add the missing omnibus tests and corrections and see what survives. The agency findings are strong enough to sustain a revised paper. The chat-baseline fairness issue should also be handled with feature-level details; this is a request for evidence, not a comment on author motivation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead Siddiqui et al.'s RCT comparing a process-oriented writing tool (Script&Shift), a chat-based LLM interface, and a standard control. The headline is: the agency results are large, consistent, and probably real. The knowledge-transformation claim in the abstract is not supported by their own statistics.\n\nWhat's new: a three-arm randomized comparison (N=90, 30 per arm) of an integrated process-oriented scaffold against a chat interface and a no-AI control, with essay coding for knowledge-transformation markers and self-report agency. That's a useful design, and the agency effects are striking—Q1 F(2,85)=101.90, with Script&Shift far above the chat condition and even above the no-AI control. Inter-rater reliability on the coding is good (κ=0.86 after reconciliation). The paper also reports a moderate correlation (r=0.61) between tool interaction and KT markers, though with high variance.\n\nThe soft spots are mainly in the knowledge-transformation analysis. Only Analysis shows an omnibus effect across all three conditions (H=8.72, p=.013). The pairwise results for Evaluation and 'Knowledge' are reported without a significant omnibus, and no multiple-comparison correction is applied anywhere. Worse, Figure 2 defines the 'Knowledge' code as paraphrased/copied information—that's knowledge telling, not transformation. Reporting Script&Shift as significantly better on Knowledge (p=.038) inverts the direction of the construct. And Synthesis, the code that best matches the authors' own definition of knowledge transformation, is null (H=2.75, p=.253). So the abstract's 'deeper knowledge transformation overall' is an overstatement. The claim should be limited to 'more analysis markers' and even then only as exploratory.\n\nThere are secondary issues: the chat baseline is a custom interface with no usability validation or feature inventory, so a less usable chat tool could inflate Script&Shift's advantage; no data or code are provided; no baseline writing measure. These are addressable.\n\nBottom line: worth engaging with as an empirical contribution, but it needs major revision before the KT claims can be trusted. I'd send it to peer review, but the authors need to redo the multiple comparisons and rewrite the abstract to match what actually held up.\n\nBest.","headline":"Solid RCT with a strong agency effect, but the abstract overclaims knowledge transformation—only one of five codes holds up, and one 'significant' code is actually knowledge telling.","tokens_in":10795,"tokens_out":2564,"would_cite":true,"duration_ms":26837,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that when a large language model is embedded into structured writing stages rather than delivered through a chat window, undergraduate writers report more agency and show more markers of deep knowledge transformation.","keywords":["AI writing tools","student agency","knowledge transformation","source-based writing","chat-based LLM interfaces","process-oriented writing support","randomized controlled trial","writing education"],"falsifier":"Take the same source-based writing task and compare Script&Shift against a chat condition whose usability and output quality are carefully matched (same underlying LLM, pre-tested ease-of-use ratings, and document-integrated suggestions rather than a blank separate window); if the agency and knowledge-transformation gaps vanish under that match, the claim that the interface paradigm causes them is refuted.","tokens_in":9805,"feed_emoji":"✍️","tokens_out":8287,"duration_ms":90616,"temperature":0.7,"pith_summary":"The paper tries to establish that the interface around an AI writing assistant, not just the model powering it, determines whether students stay in charge of their work and think deeply about sources. In a randomized trial with 90 undergraduates, students using Script&Shift, a tool that embeds LLM help into named subprocesses such as brainstorming, elaboration, audience analysis, and feedback, reported significantly more control, satisfaction, and willingness to own their essay than students using a chat-based LLM, and their essays carried more markers of analysis and evaluation. The authors argue this happens because process-oriented design keeps content and rhetoric as separate spaces that writers move between, whereas chat collapses them into ready-made text. If the claim is right, educators worried about ChatGPT eroding writing skills have a more constructive target: structuring AI support around the stages of writing rather than banning the tools.","feed_headline":"AI built into writing stages beats chat for student agency","feed_subtitle":"In a randomized trial with 90 undergraduates, an integrated tool raised perceived control and analysis markers over a chatbot.","key_machinery":"The argument is carried by the design of Script&Shift: a layered interface that separates content from rhetorical organization and lets writers summon specialized AI assistants for brainstorming, detail elaboration, audience analysis, and feedback instead of a single free-form chat. The measurement machinery is an essay-coding framework that classifies each sentence as knowledge transformation (synthesis, analysis, application, evaluation, comprehension) versus knowledge telling, combined with seven-point self-report items for agency. The experimental machinery is a three-condition randomized design with 30 participants per condition and the same LLM backend in the two AI conditions.","core_discovery":"The study's central claim is that process-oriented AI support, where the LLM is embedded into named writing subprocesses within a layered document interface, gives undergraduate writers greater felt agency and more markers of deep knowledge transformation than a conventional chat-based LLM assistant. In the randomized comparison, Script&Shift outperformed the chat condition on Analysis markers, Evaluation markers, and the Knowledge category, and it outperformed both the chat and standard conditions on three self-reported agency items: feeling in control, feeling content with the essay, and willingness to publish under one's own name. The paper is careful to note that final essay-quality scores did not differ reliably, and it treats that as consistent with prior findings that knowledge-transformation markers and grades are only weakly related.","pith_inferences":["An untested implication is that chat's disadvantage comes from context-switching and copy-paste overhead; a follow-up that gives the chat condition inline document integration would isolate that mechanism.","If the tool's benefit is metacognitive rather than merely structural, the effect should survive a transfer task without Script&Shift; the single-session design cannot establish that.","The larger agency gap over the no-AI control hints that 'control' may mean having a visible, structured process, not just the absence of automation; keystroke-level authorship analysis would test that reading.","The null essay-quality result implies grade-based evaluations will not capture the process gains claimed here; portfolios or process-trace measures would be needed to see them."],"forward_implications":["Designing AI writing support around explicit writing subprocesses is a way to keep students in charge of their essays while still getting LLM help.","Chat-style assistants can be expected to show larger agency costs, and adding structured scaffolding may be more valuable than swapping the underlying model.","Process gains such as knowledge transformation can occur even when final essay scores do not improve, so outcome measures that rely only on grades will miss them.","Because participants felt more agency with Script&Shift than with no AI at all, AI assistance need not be framed as a trade-off against ownership.","Longer or repeated use would be required to see whether the knowledge-transformation markers translate into better final essays."],"supporting_citations":[{"why":"Describes Script&Shift, the process-oriented interface whose design is the experimental treatment being tested.","marker":"[33]"},{"why":"Provides the coding scheme used to classify essay content into knowledge-transformation markers versus knowledge telling.","marker":"[28]"},{"why":"Defines the knowledge-telling versus knowledge-transforming distinction that motivates the dependent variable.","marker":"[30]"},{"why":"Reports earlier evidence that chat-based AI use reduces agency in source-based writing, the baseline this study challenges.","marker":"[37]"},{"why":"Establishes source-based writing tasks as a valid setting for eliciting knowledge transformation.","marker":"[32]"},{"why":"Documents the weak correlation between knowledge-transformation markers and essay grades, which the paper uses to explain its null essay-quality results.","marker":"[29]"},{"why":"Argues that AI integration can support metacognitive reflection and writer agency, the theoretical possibility the study tests.","marker":"[25]"},{"why":"Supplies the cognitive process model of writing that grounds the process-oriented design rationale.","marker":"[12]"}],"fun_headline_variants":["Integrated AI tool boosts student agency over chat assistant","Process-embedded AI beats chatbot for writing ownership","Designed AI support fosters student agency in writing study","AI embedded in writing stages increases perceived control","Tool-integrated AI enhances agency over chat-based LLM"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result rests on the assumption that differences came from the interface paradigm rather than from the particular quality, usability, or familiarity of the custom chat tool, and the paper's own limitations section notes that the single 1.5-hour session may have restricted engagement.","fun_headline_variants_meta":{"raw":{"variants":["Integrated AI tool boosts student agency over chat assistant","Process-embedded AI beats chatbot for writing ownership","Designed AI support fosters student agency in writing study","AI embedded in writing stages increases perceived control","Tool-integrated AI enhances agency over chat-based LLM"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":1132,"prompt_tokens":856,"completion_tokens":276,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":203}},"tokens_in":472,"tokens_out":276,"duration_ms":2859,"temperature":1.0,"reasoning_tokens":203,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:44:51.301811+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same source-based writing task and compare Script&Shift against a chat condition whose usability and output quality are carefully matched (same underlying LLM, pre-tested ease-of-use ratings, and document-integrated suggestions rather than a blank separate window); if the agency and knowledge-transformation gaps vanish under that match, the claim that the interface paradigm causes them is refuted.","supporting_citations":[{"cited_title":"In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems","cited_arxiv_id":null,"evidence_quote":"Describes Script&Shift, the process-oriented interface whose design is the experimental treatment being tested."},{"cited_title":"In: International Learning Analytics & Knowledge Conference","cited_arxiv_id":null,"evidence_quote":"Provides the coding scheme used to classify essay content into knowledge-transformation markers versus knowledge telling."},{"cited_title":"Advances in Applied Psycholinguistics2, 142–175 (1987)","cited_arxiv_id":null,"evidence_quote":"Defines the knowledge-telling versus knowledge-transforming distinction that motivates the dependent variable."},{"cited_title":"International Journal of Artificial Intelligence in Education pp","cited_arxiv_id":null,"evidence_quote":"Reports earlier evidence that chat-based AI use reduces agency in source-based writing, the baseline this study challenges."},{"cited_title":"L1-Educational Studies in Language And Literature 4, 5–33 (2004)","cited_arxiv_id":null,"evidence_quote":"Establishes source-based writing tasks as a valid setting for eliciting knowledge transformation."},{"cited_title":"Journal of Computer Assisted Learning37(4), 903–924 (2021)","cited_arxiv_id":null,"evidence_quote":"Documents the weak correlation between knowledge-transformation markers and essay grades, which the paper uses to explain its null essay-quality results."},{"cited_title":"International Journal of Artificial Intelligence in Education pp","cited_arxiv_id":null,"evidence_quote":"Argues that AI integration can support metacognitive reflection and writer agency, the theoretical possibility the study tests."},{"cited_title":"College Composition and Communication 32(4), 365–387 (1981)","cited_arxiv_id":null,"evidence_quote":"Supplies the cognitive process model of writing that grounds the process-oriented design rationale."}],"review_version":1}