Pith. sign in

REVIEW 4 major objections 5 minor 65 references

Modifying AI, Enhancing Essays: How Active Engagement with Generative AI Boosts Writing Quality

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Frequent modification of AI-generated text improves lexical sophistication, syntactic complexity, and text cohesion, while verbatim acceptance degrades them.

desk verdict First causal-inference attempt on CoAuthor is worth engaging, but the headline ATE claims lack the standard errors and tests needed to support 'significant' and 'consistently improves.' read the letter →

arxiv 2412.07200 v1 pith:EFKQC6KE submitted 2024-12-10 cs.HC cs.AIcs.CL

classification cs.HCcs.AIcs.CL
keywords generativeAI-assistedwritingcausalinferenceX-learneressayqualitylexicalsophisticationsyntacticcomplexitytextcohesiongenderbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the way a student engages with AI writing suggestions changes the quality of the final essay, not just whether AI is used. On 1,445 logged GPT-3-assisted writing sessions from the CoAuthor dataset, it compares three behaviors: asking for suggestions and not accepting them (T1), accepting them unchanged (T2), and accepting then revising them (T3). Its central result is causal: frequent T3 raises lexical sophistication, syntactic complexity, and text cohesion, while frequent T2 lowers all three. It also finds that all three behaviors reduce gender bias in essays, with the largest reduction for T1. If these estimates hold, teachers can use observable writing-process logs to distinguish meaningful engagement from passive delegation.

What carries the argument

The argument is carried by a causal graph plus the X-learner meta-algorithm, a machine-learning approach that combines treated and untreated observations to estimate average and individual treatment effects. Behaviors are encoded as binary treatments (above or below the median frequency of the pattern), the four essay metrics as outcomes, and five confounders—writing genre, writing topic, language background, GPT temperature, and frequency penalty—as the back-door adjustment set. The back-door criterion identifies the causal estimands from observational data, and X-learner estimates both average and individual treatment effects. The paper uses random common cause, placebo treatment, and data subset refutations to check that the estimates are not artifacts of the model.

What would settle it

A reanalysis that adds writer-level covariates, such as independent writing ability measured before AI exposure, GAI literacy, motivation, or prior topic knowledge, to the adjustment set would falsify the causal claim if the T3 average treatment effects on lexical sophistication, syntactic complexity, or cohesion collapse toward zero. A more direct falsifier is a randomized experiment in which writers are instructed either to revise or to accept AI suggestions verbatim; observing no difference in the three text-quality metrics would refute the paper's causal interpretation.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that accepting AI-generated text and revising it (T3) is the only one of the three GAI-assisted writing behaviors that consistently improves the three text-quality measures; the estimated average treatment effects are positive for lexical sophistication, syntactic complexity, and text cohesion (0.102, 0.963, and 0.008 in the paper's reported ATE table). Accepting AI suggestions verbatim (T2) is associated with the largest negative effects on all three, and seeking suggestions without accepting them (T1) lowers lexical sophistication and syntactic complexity while slightly raising cohesion. The same causal model yields a separate finding: all three behaviors reduce gender bias in the essays, with T1 producing the largest reduction, which the paper interprets as evidence that human independent writing itself introduces measurable linguistic bias. Because the effects pass the paper's refutation checks, the authors present these as credible causal relationships rather than correlations.

Load-bearing premise

The load-bearing premise is that the five logged confounders—genre, topic, language background, GPT temperature, and frequency penalty—are enough to block all back-door paths, so unmeasured writer traits like skill, motivation, or GAI literacy do not distort the estimated effects.

Editorial extensions

If this is right

  • Instructors should treat final essay quality alone as insufficient evidence of learning in GAI-assisted writing; process logs showing whether suggestions were revised carry assessment-relevant information.
  • Pedagogical guidance should push students toward critically editing AI suggestions rather than accepting them verbatim, since verbatim acceptance is estimated to reduce all three text-quality measures.
  • Non-native English writers may gain lexical sophistication and syntactic complexity from GAI suggestions, but they also show higher gender bias when writing independently or when revising AI text, suggesting targeted support is needed.
  • The benefit of revising AI suggestions depends on context: it improves cohesion more in creative writing than in argumentative writing, and the effects shift with GPT temperature and frequency penalty settings.
  • All three GAI behaviors reduce gender bias relative to independent writing, so AI-assisted writing can serve as a partial bias-mitigation strategy even when suggestions are accepted unchanged.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: the estimated average treatment effects come from an adult crowd-sourced writing population, so a natural extension is to test whether the T3 advantage is larger for lower-skill writers, who may have more room to learn from revision, or for higher-skill writers, who may revise more effectively.
  • An editorial inference: if verbatim acceptance genuinely lowers quality, current AI autocomplete interfaces may be nudging users toward worse text by making acceptance the default action; redesigning them to require a deliberate edit could be a low-cost experiment.
  • The gender-bias result for T1, where seeking suggestions without accepting them reduces bias the most, suggests that merely reading AI suggestions may change a writer's attention; this could be tested directly by comparing essays written after exposure to de-biased brainstorming prompts against a no-prompt control.
  • The authors list missing GAI literacy as a limitation; a follow-up could also record revision quality, such as the number of semantic changes made to AI text, which would make the hypothesized learning mechanism testable rather than inferred from behavior alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates how different patterns of interaction with generative AI (GAI) during writing affect the quality of the resulting essay. Using the CoAuthor dataset (1,445 writing sessions from 63 writers), the authors define three binary treatments: T1 (seeking suggestions but not accepting them), T2 (accepting GAI suggestions without revision), and T3 (accepting and then modifying GAI suggestions). Outcomes are four text-based quality measures: lexical sophistication, syntactic complexity, text cohesion, and gender bias. The authors construct a DAG, apply the X-learner with back-door adjustment on five confounders, and report average treatment effects (ATEs). They claim that T3 consistently improves lexical sophistication, syntactic complexity, and cohesion, while T2 reduces them, and that all three behaviors reduce gender bias. The paper includes refutation checks (Random Common Cause, Placebo Treatment, Data Subset Refuter) and SHAP-based subgroup analyses.

Significance. If the causal claims were adequately supported, the paper would make a useful contribution to GAI-assisted writing research and educational assessment. It addresses a timely question, uses a publicly available dataset, and is among the first to attempt causal effect estimation for writing behaviors in this setting. The authors provide a clearly stated DAG, a reproducible code repository, and several refutation checks, which are strengths. However, the central quantitative claims are currently not backed by inferential statistics: reported ATEs come without standard errors, confidence intervals, or tests against zero, and the paper does not account for the nested structure of the data. These gaps are load-bearing because the abstract and discussion state that effects are 'significant' and 'consistent.' The manuscript therefore needs substantial revision before its conclusions can be accepted.

major comments (4)
  1. [Section 4.1, Table 2] The ATE estimates are reported without standard errors, confidence intervals, or p-values testing the null hypothesis ATE = 0. The p-values in the RCC, Placebo, and DSR columns test whether the refutation estimate differs from the original point estimate, not whether the original ATE differs from zero. For example, ATE_T3,Y2 = 0.963 may be well within sampling noise; the reported refutation p-values do not rule out a true effect of zero. Consequently, the abstract and Discussion statements that T3 'consistently and significantly improves' all three quality measures are not supported by the evidence as presented. Please report inferential statistics for each ATE, and correct for multiple comparisons across the twelve treatment-outcome pairs.
  2. [Sections 3.1 and 4.1] The 1,445 writing sessions are produced by only 63 writers, so the observations are not independent. The paper does not account for writer-level clustering in the X-learner estimation or in the refutation procedures. Any confidence intervals or significance tests added in response to the previous comment must be cluster-robust at the writer level (or use a multilevel model); otherwise, uncertainty is understated and the refutation p-values are overconfident. The present analysis gives no indication of how much of the variation in outcomes is between-writer versus within-writer, which is essential for interpreting effects of behaviors that are inherently writer-level tendencies.
  3. [Section 3.2, treatment binarization] Treatments T1, T2, and T3 are defined by dichotomizing each behavioral frequency at the sample median (Section 3.2, third paragraph). No sensitivity analysis is reported for this choice of threshold, and the median split is effectively a free parameter. It is plausible that the estimated ATEs, especially the smaller effects on Y3 (text cohesion) and Y4 (gender bias), are sensitive to the cutoff. Please report results at alternative thresholds (e.g., quartiles) or use a continuous treatment representation (e.g., dose-response or ordinal treatment) to assess the robustness of the qualitative conclusions.
  4. [Section 3.2 and Limitations] The back-door adjustment set (C1-C5) is assumed sufficient for identifiability, but the Limitations section acknowledges that writer-level variables such as GAI literacy are missing. If writing skill, motivation, or GAI literacy affects both treatment assignment (behavioral pattern) and essay quality, the ATEs are biased. Because the paper's headline claims are causal, this is a load-bearing threat, not a routine caveat. A concrete sensitivity analysis (e.g., E-values or a negative-control outcome) would help bound the potential bias from unmeasured confounding and is necessary to support the strength of the current causal language.
minor comments (5)
  1. [Section 3.2 and Table 1] The numbering of confounders is inconsistent: the bullet list in Section 3.2 calls language background C2 and writing topic C3, while Table 1 and Table 3 use the reverse labeling. Please align the numbering throughout the text, tables, and figures.
  2. [Table 1] The description of T3 contains a typo: 'Ratio between teh accepted GAI suggestions' should be 'Ratio between the accepted GAI suggestions.'
  3. [Section 4.1, fourth paragraph] The phrase 'T3 (seek suggestions -> first accept and the revise)' should read 'T3 (seek suggestions -> accept and then revise).'
  4. [Figures 2-4] The beeswarm plots are referenced extensively in Section 4.2, but the captions in the manuscript do not describe the axis labels or the meaning of dot colors, and the axes are not legible in the provided version. Please add explicit labels and a legend.
  5. [Limitations] The sentence 'the imbalance in data distribution (e.g., no-native writers vs. native writers)' contains a typo: 'no-native' should be 'non-native.'

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the reported causal effects are estimated from external writing-process logs and text-quality metrics, not from fitted inputs or self-cited results.

full rationale

The paper's central claim—that actively modifying GAI suggestions (T3) is associated with improved lexical sophistication, syntactic complexity, and cohesion, while accepting suggestions without revision (T2) is associated with declines—is an empirical ATE estimate. The treatments are behavior counts/ratios derived from CoAuthor interaction logs; the outcomes are computed with external, literature-based metrics (Advanced Guiraud, Mean Length of T-Unit, Semantic Overlap, Genbit). No treatment is defined in terms of an outcome, and no outcome is defined in terms of a treatment, so the ATEs are not self-definitional. The X-learner and back-door adjustment are standard methods applied to observed data, and the reported refutation checks, whatever their statistical limitations, do not constitute a fitted-input-called-prediction pattern. The paper does cite prior work by the same group (e.g., Yang et al. [64]) to motivate the choice of behavioral patterns and to note prior use of CoAuthor, but these citations are background motivation rather than load-bearing evidence: the causal estimates do not reduce to those cited results, and no uniqueness theorem or ansatz is imported from the authors' own prior work. The complementarity of T2 and T3 (both ratios over total accepted suggestions) means the two treatments encode the same acceptance-versus-modification axis, but the sign and size of the ATEs are still determined by the external outcome data rather than by definition. The acknowledged absence of writer-level confounders and the lack of inferential statistics are validity concerns, not circularity. The derivation chain is therefore self-contained with respect to circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central causal claim rests on the correctness of the causal DAG and the sufficiency of the adjustment set, the validity of the text metrics as quality proxies, the representativeness of the crowdsourced dataset, and standard causal inference assumptions. The median-split treatment definitions are data-dependent thresholds that are not reported.

free parameters (1)
  • Median split threshold for each treatment = not reported
    Each treatment (T1, T2, T3) is binarized as above vs. below the sample median of the behavior frequency (Section 3.2). The thresholds are data-derived and not theory-grounded, and the paper does not report the median values or test sensitivity to alternative cutoffs.
assumptions (4)
  • domain assumption Causal sufficiency: conditioning on C1-C5 (genre, topic, language background, GPT temperature, frequency penalty) blocks all back-door paths between each treatment and each outcome.
    Required for back-door identification (Section 3.2, Figure 1). The paper acknowledges in Limitations that writer-related information such as GAI literacy is unavailable, so this assumption is fragile.
  • domain assumption The four outcome metrics (Advanced Guiraud, MLT, Semantic Overlap, Genbit Score) are valid and sufficient measures of essay quality.
    Borrowed from prior literature (Section 3.2), but the paper does not validate them against human judgments on this dataset, and 'quality' is therefore defined by these specific text statistics.
  • domain assumption The CoAuthor dataset (crowdsourced writers, GPT-3 suggestions) is representative of authentic educational GAI-assisted writing.
    The paper argues resemblance to educational scenarios (Section 3.1) but acknowledges limited generalizability in Limitations.
  • standard math Positivity/overlap: for every combination of confounders, there is a nonzero probability of receiving each treatment level.
    Required for X-learner estimation; not checked in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Modifying AI, Enhancing Essays: How Active Engagement with Generative AI Boosts Writing Quality." pith.science (2026). https://pith.science/paper/EFKQC6KE

@misc{pith2026241207200,
  author       = {Pith},
  title        = {Pith review of: Modifying AI, Enhancing Essays: How Active Engagement with Generative AI Boosts Writing Quality},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EFKQC6KE}},
  note         = {Machine review of arXiv:2412.07200}
}
read the original abstract

Students are increasingly relying on Generative AI (GAI) to support their writing-a key pedagogical practice in education. In GAI-assisted writing, students can delegate core cognitive tasks (e.g., generating ideas and turning them into sentences) to GAI while still producing high-quality essays. This creates new challenges for teachers in assessing and supporting student learning, as they often lack insight into whether students are engaging in meaningful cognitive processes during writing or how much of the essay's quality can be attributed to those processes. This study aimed to help teachers better assess and support student learning in GAI-assisted writing by examining how different writing behaviors, especially those indicative of meaningful learning versus those that are not, impact essay quality. Using a dataset of 1,445 GAI-assisted writing sessions, we applied the cutting-edge method, X-Learner, to quantify the causal impact of three GAI-assisted writing behavioral patterns (i.e., seeking suggestions but not accepting them, seeking suggestions and accepting them as they are, and seeking suggestions and accepting them with modification) on four measures of essay quality (i.e., lexical sophistication, syntactic complexity, text cohesion, and linguistic bias). Our analysis showed that writers who frequently modified GAI-generated text-suggesting active engagement in higher-order cognitive processes-consistently improved the quality of their essays in terms of lexical sophistication, syntactic complexity, and text cohesion. In contrast, those who often accepted GAI-generated text without changes, primarily engaging in lower-order processes, saw a decrease in essay quality. Additionally, while human writers tend to introduce linguistic bias when writing independently, incorporating GAI-generated text-even without modification-can help mitigate this bias.

Figures

Figures reproduced from arXiv: 2412.07200 by the authors.

Figure 1
Figure 1. Causal graph of GAI-assisted writing. 3.3 Causal Effect Identification and Estimation Our objective was to evaluate the effect of a treatment on an outcome. However, in reality, we cannot observe counterfactual values, which are essential for calculating causal effects but are inherently unobservable. To address this, we rely on identification techniques and assumptions that reduce causal estimands into statistical … view at source ↗
Figure 2
Figure 2. Beeswarm plot of seek suggestions -> not accept (T1) & outcomes [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Beeswarm plot of seek suggestions -> accept without revision (T2) & outcomes. refutation tests (i.e., RCC, Placebo, and DSR), with all p-values exceeding 0.05. This indicates that our causal inference results are both robust and credible. Manuscript submitted to ACM [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Beeswarm plot of seek suggestions -> first accept and the revise (T3) & outcomes [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 60 canonical work pages

  1. [1]

    Generative AI in the Classroom: Can Students Remain Active Learners?

    R. Abdelghani, H. Sauzéon, and P.-Y. Oudeyer. 2023. Generative AI in the Classroom: Can Students Remain Active Learners? arXiv preprint arXiv:2310.03192 (2023). Manuscript submitted to ACM How Active Engagement with Generative AI Boosts Writing Quality 15

  2. [2]

    Acerbi and J

    A. Acerbi and J. M. Stubbersfield. 2023. Large language models show human-like content biases in transmission chain experiments. Proceedings of the National Academy of Sciences 120, 44 (2023), e2313790120

  3. [3]

    Alipourfard, P

    N. Alipourfard, P. Fennell, and K. Lerman. 2018. Using Simpson’s paradox to discover interesting patterns in behavioral data. In Proceedings of the International AAAI Conference on Web and Social Media , Vol. 12. 2–11

  4. [4]

    Antwarg, R

    L. Antwarg, R. M. Miller, B. Shapira, and L. Rokach. 2021. Explaining anomalies detected by autoencoders using Shapley Additive Explanations. Expert systems with applications 186 (2021), 115736

  5. [5]

    A. N. Applebee. 1984. Writing and reasoning. Review of educational research 54, 4 (1984), 577–596

  6. [6]

    V. M. Baaijen, D. Galbraith, and K. De Glopper. 2012. Keystroke analysis: Reflections on procedures and measures. Written Communication 29, 3 (2012), 246–277

  7. [7]

    K. M. Baker. 2016. Peer review as a strategy for improving students’ writing process. Active Learning in Higher Education 17, 3 (2016), 179–192

  8. [8]

    Bardovi-Harlig

    K. Bardovi-Harlig. 1992. A second look at T-unit analysis: Reconsidering the sentence. TESOL quarterly 26, 2 (1992), 390–395

Show all 65 references
  1. [9]

    Boscolo, L

    P. Boscolo, L. Del Favero, and M. Borghetto. 2006. Writing on an interesting topic: Does writing foster interest? In Writing and motivation. Brill, 71–91

  2. [10]

    D. Boud. 2007. Reframing assessment as if learning were important. In Rethinking assessment in higher education . Routledge, 24–36

  3. [11]

    S. C. Cantrell, J. F. Almasi, M. Rintamaa, and J. C. Carter. 2016. Supplemental reading strategy instruction for adolescents: A randomized trial and follow-up study. The Journal of Educational Research 109, 1 (2016), 7–26

  4. [12]

    Cartwright

    N. Cartwright. 2010. What are randomised controlled trials good for? Philosophical studies 147, 1 (2010), 59–70

  5. [13]

    Cheng, K

    Y. Cheng, K. Lyons, G. Chen, D. Gašević, and Z. Swiecki. 2024. Evidence-centered Assessment for Writing with Generative AI. In Proceedings of the 14th Learning Analytics and Knowledge Conference . 178–188

  6. [14]

    Coenen, L

    A. Coenen, L. Davis, D. Ippolito, E. Reif, and A. Yuan. 2021. Wordcraft: A human-AI collaborative editor for story writing. arXiv preprint arXiv:2107.07430 (2021)

  7. [15]

    Condon and D

    W. Condon and D. Kelly-Riley. 2004. Assessing and teaching what we value: The relationship between college-level writing and critical thinking abilities. Assessing Writing 9, 1 (2004), 56–75

  8. [16]

    Conijn, J

    R. Conijn, J. Roeser, and M. Van Zaanen. 2019. Understanding the keystroke log: the effect of writing task on keystroke features. Reading and Writing 32, 9 (2019), 2353–2374

  9. [17]

    Crossley and D

    S. Crossley and D. McNamara. 2010. Cohesion, coherence, and expert evaluations of writing proficiency. In Proceedings of the Annual Meeting of the Cognitive Science Society, Vol. 32

  10. [18]

    S. A. Crossley. 2020. Linguistic features in writing quality and development: An overview. Journal of Writing Research 11, 3 (2020), 415–443

  11. [19]

    S. A. Crossley and D. S. McNamara. 2016. Say more and be more coherent: How text elaboration and cohesion can increase writing quality. Journal of Writing Research 7, 3 (2016), 351–370

  12. [20]

    Daller, R

    H. Daller, R. Van Hout, and J. Treffers-Daller. 2003. Lexical richness in the spontaneous speech of bilinguals.Applied linguistics 24, 2 (2003), 197–222

  13. [21]

    Deane and M

    P. Deane and M. Zhang. 2015. Exploring the feasibility of using writing process features to assess text production skills. ETS Research Report Series 2015, 2 (2015), 1–16

  14. [22]

    Denton and M

    J. Denton and M. Mabry. 1981. Causal Modeling and Research on Teacher Education. Journal of Experimental Education 49 (1981), 207–213. https://doi.org/10.1080/00220973.1981.11011785

  15. [23]

    P. S. Dhillon, S. Molaei, J. Li, M. Golub, S. Zheng, and L. P. Robert. 2024. Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–18

  16. [24]

    J. Emig. 1977. Writing as a mode of learning. College Composition & Communication 28, 2 (1977), 122–128

  17. [25]

    Feder, K

    A. Feder, K. A. Keith, E. Manzoor, R. Pryzant, D. Sridhar, Z. Wood-Doughty, J. Eisenstein, J. Grimmer, R. Reichart, M. E. Roberts, et al. 2022. Causal inference in natural language processing: Estimation, prediction, interpretation and beyond. Transactions of the Association f...

  18. [26]

    L. Flower. 1981. A cognitive process theory of writing. Composition and communication (1981)

  19. [27]

    M. J. Forgeard, S. B. Kaufman, and J. C. Kaufman. 2013. The psychology of creative writing. A Companion to creative writing (2013), 320–333

  20. [28]

    Grabe and R

    W. Grabe and R. B. Kaplan. 2014. Theory and practice of writing: An applied linguistic perspective . Routledge

  21. [29]

    R. Grimmer. 2018. The evolution of genre in the writing process. cinder 1 (2018)

  22. [30]

    M. J. Ha. 2022. Syntactic complexity in EFL writing: Within-genre topic and writing quality. Computer-Assisted Language Learning Electronic Journal 23, 1 (2022), 187–205

  23. [31]

    M. A. K. Halliday and R. Hasan. 2014. Cohesion in english. Routledge

  24. [32]

    Hariton and J

    E. Hariton and J. J. Locascio. 2018. Randomised controlled trials—the gold standard for effectiveness research. BJOG: an international journal of obstetrics and gynaecology 125, 13 (2018), 1716

  25. [33]

    Isaacson

    S. Isaacson. 1988. Assessing the writing product: Qualitative and quantitative measures. Exceptional Children 54, 6 (1988), 528–534

  26. [34]

    Y. Jin, L. Yan, V. Echeverria, D. Gašević, and R. Martinez-Maldonado. 2024. Generative AI in Higher Education: A Global Perspective of Institutional Adoption Policies and Guidelines. arXiv preprint arXiv:2405.11800 (2024)

  27. [35]

    Y.-S. G. Kim. 2020. Structural relations of language and cognitive skills, and topic knowledge to written composition: A test of the direct and indirect effects model of writing. British Journal of Educational Psychology 90, 4 (2020), 910–932. Manuscript submitted to ACM 16 Ya...

  28. [36]

    A. M. Knowles. 2022. Human-AI collaborative writing: Sharing the rhetorical task load. In 2022 IEEE International Professional Communication Conference (ProComm). IEEE, 257–261

  29. [37]

    Koppenhaver and A

    D. Koppenhaver and A. Williams. 2010. A conceptual review of writing research in augmentative and alternative communication. Augmentative and Alternative Communication 26, 3 (2010), 158–176

  30. [38]

    Kotek, R

    H. Kotek, R. Dockum, and D. Sun. 2023. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference. 12–24

  31. [39]

    S. R. Künzel, J. S. Sekhon, P. J. Bickel, and B. Yu. 2019. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences 116, 10 (2019), 4156–4165

  32. [40]

    Kyle and S

    K. Kyle and S. Crossley. 2016. The relationship between lexical sophistication and independent and source-based writing.Journal of Second Language Writing 34 (2016), 12–24

  33. [41]

    J. A. Langer and A. N. Applebee. 1987. How Writing Shapes Thinking: A Study of Teaching and Learning. NCTE Research Report No. 22. ERIC

  34. [42]

    Laufer and P

    B. Laufer and P. Nation. 1995. Vocabulary size and use: Lexical richness in L2 written production. Applied linguistics 16, 3 (1995), 307–322

  35. [43]

    Lea and B

    M. Lea and B. Street. 1998. Student writing in higher education: An academic literacies approach. Studies in Higher Education 23 (1998), 157–172. https://doi.org/10.1080/03075079812331380364

  36. [44]

    M. Lee, P. Liang, and Q. Yang. 2022. Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–19

  37. [45]

    Leijten and L

    M. Leijten and L. Van Waes. 2013. Keystroke logging in writing research: Using Inputlog to analyze and visualize writing processes. Written Communication 30, 3 (2013), 358–392

  38. [46]

    X. Lu. 2011. A corpus-based evaluation of syntactic complexity measures as indices of college-level ESL writers’ language development. TESOL quarterly 45, 1 (2011), 36–62

  39. [47]

    Luther, J

    T. Luther, J. Kimmerle, and U. Cress. 2024. Teaming up with an AI: Exploring Human–AI Collaboration in a Writing Scenario with ChatGPT. AI 5, 3 (2024), 1357–1376

  40. [48]

    M. H. Maathuis and D. Colombo. 2015. A generalized back-door criterion. (2015)

  41. [49]

    D. S. McNamara, S. A. Crossley, and P. M. McCarthy. 2010. Linguistic features of writing quality. Written communication 27, 1 (2010), 57–86

  42. [50]

    Navigli, S

    R. Navigli, S. Conia, and B. Ross. 2023. Biases in large language models: origins, inventory, and discussion. ACM Journal of Data and Information Quality 15, 2 (2023), 1–21

  43. [51]

    Nguyen, Y

    A. Nguyen, Y. Hong, B. Dang, and X. Huang. 2024. Human-AI collaboration patterns in AI-assisted academic writing. Studies in Higher Education (2024), 1–18

  44. [52]

    Parkerson, R

    J. Parkerson, R. Lomax, D. Schiller, and H. Walberg. 1984. Exploring Causal Models of Educational Achievement. Journal of Educational Psychology 76 (1984), 638–646. https://doi.org/10.1037/0022-0663.76.4.638

  45. [53]

    M. L. Petersen and M. J. van der Laan. 2014. Causal models and learning from data: integrating causal modeling and statistical estimation. Epidemiology 25, 3 (2014), 418–426

  46. [54]

    Putjorn and P

    T. Putjorn and P. Putjorn. 2023. Augmented Imagination: Exploring Generative AI from the Perspectives of Young Learners. In2023 15th International Conference on Information Technology and Electrical Engineering (ICITEE) . IEEE, 353–358

  47. [55]

    Sengupta, R

    K. Sengupta, R. Maher, D. Groves, and C. Olieman. 2021. GenBiT: measure and mitigate gender bias in language datasets. Microsoft Journal of Applied Research 16 (2021), 63–71

  48. [56]

    Sharma and E

    A. Sharma and E. Kiciman. 2020. DoWhy: An End-to-End Library for Causal Inference. arXiv preprint arXiv:2011.04216 (2020)

  49. [57]

    Shibani, R

    A. Shibani, R. Rajalakshmi, F. Mattins, S. Selvaraj, and S. Knight. 2023. Visual Representation of Co-Authorship with GPT-3: Studying Human-Machine Interaction for Effective Writing. International Educational Data Mining Society (2023)

  50. [58]

    Taguchi, W

    N. Taguchi, W. Crawford, and D. Z. Wetzel. 2013. What linguistic features are indicative of writing quality? A case of argumentative essays in a college composition program. Tesol Quarterly 47, 2 (2013), 420–430

  51. [59]

    ten Peze, T

    A. ten Peze, T. Janssen, G. Rijlaarsdam, and D. van Weijen. 2024. Instruction in creative and argumentative writing: transfer and crossover effects on writing process and text quality. Instructional Science 52, 3 (2024), 341–383

  52. [60]

    Van Der Geest and L

    T. Van Der Geest and L. Van Gemert. 1997. Review as a method for improving professional texts. Journal of business and technical communication 11, 4 (1997), 433–450

  53. [61]

    Wambsganss, X

    T. Wambsganss, X. Su, V. Swamy, S. P. Neshaei, R. Rietsche, and T. Käser. 2023. Unraveling Downstream Gender Bias from Large Language Models: A Study on AI Educational Writing Assistance. arXiv preprint arXiv:2311.03311 (2023)

  54. [62]

    C. L. Wesson, S. Deno, P. K. Mirkin, G. Maruyama, R. Skiba, R. King, and B. Sevcik. 1988. A Causal Analysis of the Relationships Among Ongoing Curriculum-Based Measurement and Evaluation, the Structure of Instruction, and Student Achievement. The Journal of Special Education 2...

  55. [63]

    U. Wingate. 2012. ‘Argument!’helping students understand what essay writing is about. Journal of English for academic purposes 11, 2 (2012), 145–154

  56. [64]

    K. Yang, Y. Cheng, L. Zhao, M. Raković, Z. Swiecki, D. Gašević, and G. Chen. 2024. Ink and Algorithm: Exploring Temporal Dynamics in Human-AI Collaborative Writing. arXiv:2406.14885 [cs.HC] https://arxiv.org/abs/2406.14885

  57. [65]

    W. Yang, X. Lu, and S. C. Weigle. 2015. Different topics, different discourse: Relationships among writing topic, measures of syntactic complexity, and judgments of writing quality. Journal of Second Language Writing 28 (2015), 53–67. Manuscript submitted to ACM

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.