{"id":"94e689f1-e155-4150-a05b-4607334eacbb","arxiv_id":"2504.13038","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Student essays in a large MOOC became longer, less lexically varied, and more likely to contain LLM-associated words after ChatGPT's release, while broad essay topics stayed similar.","lead":"A study of 56,878 essays from a free AI ethics MOOC finds that after ChatGPT launched, student essays became longer, easier to read, and more uniform in vocabulary. It is one of the first longitudinal looks at how generative AI may be shaping real student coursework outside controlled lab settings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The conclusion that a meaningful share of MOOC participants rely on LLMs is not identified from a single-course before/after contrast; without participant-level panel analysis or a comparison group, cohort and course-context shifts remain a plausible alternative.","rationale":"The reader's weakest_assumption correctly identifies the most fragile premise: the absence of demographic information and the lack of a comparison group mean that the observed pre/post differences could stem from changes in the student population or course context over the four-year window. My stress-test agrees with this assessment. The paper's own limitations section (Section 5.3) and discussion (Section 5.1) acknowledge that the data cannot rule out coincidence or quantify the role of shifting public discourse. Because the central claim includes a behavioral conclusion about LLM reliance — not just a descriptive statement about essay length — the analysis needs at least a within-participant comparison or a difference-in-differences design to support that conclusion. The descriptive statistics themselves (token counts, sentence counts, word prevalence) are large and internally consistent, and the authors are transparent about several limitations. However, the inferential step from correlation to LLM use is the weak point, and the proposed fixed-effects test would directly address it. Since the reader already issued a CONDITIONAL verdict based on this same concern, no verdict change is needed; my read leaves the reader's assessment unchanged.","tokens_in":11004,"tokens_out":4155,"duration_ms":43933,"concrete_test":"Using the participant identifiers implied by the 3,582 participants and 56,878 submissions, estimate a fixed-effects regression of log(token count) on a post-ChatGPT indicator, with participant fixed effects and assignment/prompt fixed effects, restricted to participants with at least one submission both before and after the release. If the within-participant post-period coefficient is much smaller than the pooled 79.6-token mean difference or is not significantly positive, then the headline increase is attributable to compositional turnover in who takes the course, not to a change in writing behavior. If the within-participant coefficient is comparable, the cohort-change alternative is weakened and the LLM-reliance interpretation gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference in Sections 5.2 and 6 — that a meaningful proportion of participants rely on LLMs such as ChatGPT — rests on a pooled before/after comparison within one MOOC. The authors explicitly acknowledge in Section 5.3 that demographic information is unavailable and that results may not generalize. This matters because the aggregate token count rose from 150.5 to 230.1, sentence count from 6.85 to 9.76, and TTR fell from 0.617 to 0.577, but these are only population-level shifts. Without a control group, an interrupted time-series design, or within-participant comparisons, the shifts could equally be explained by changes in who enrolls in the course (e.g., more non-native speakers after ChatGPT's public visibility), changes in course materials or peer-review norms, or the fact that the course topic is AI ethics, so public discourse about LLMs could change what students write about and how much they write regardless of whether they use LLMs. Section 5.1 itself concedes that the terminology changes could derive from 'shifting focuses of interest and public discourse applied to even a static essay prompt'. The jump in March 2023 rather than immediately after November 2022 further complicates the attribution, since it suggests a gradual diffusion or a course-cycle effect. Therefore, the descriptive findings are plausible, but the load-bearing leap from 'essays changed after ChatGPT' to 'students rely on LLMs' is not secured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a longitudinal observational study of 56,878 English-language essay submissions from 3,582 participants in a free University of Helsinki MOOC on AI ethics, spanning November 2020 to October 2024. The authors compare essays submitted before ChatGPT's release with essays submitted at least one year after, and report statistically significant increases in mean token count (150.5 to 230.1), sentence count (6.85 to 9.76), Flesch Reading Ease, and in the relative prevalence of words previously identified as LLM-associated (e.g., 'delve' and 'foster' increased roughly tenfold), alongside a decrease in type-token ratio (0.617 to 0.577). They also report changes in AI-related terminology but no meaningful changes in topic-model-derived topics. The paper interprets these shifts as evidence that a meaningful proportion of MOOC participants rely on LLMs, while acknowledging that the data cannot rule out coincidence or identify specific AI-generated essays.","tokens_in":11303,"tokens_out":4295,"duration_ms":40657,"significance":"The study's descriptive contribution is valuable: to my knowledge it is one of the first longitudinal analyses of MOOC essay text across the ChatGPT launch, using a large naturalistic corpus and externally motivated indicator words. The main findings are large, internally consistent, and in the direction predicted by prior work on LLM writing style. However, the inferential leap from aggregate before/after changes to 'students rely on LLMs' is not secured by the design, and the statistical analysis overstates precision by ignoring clustering. If the authors reframe the contribution as descriptive evidence and weaken the causal conclusion, the paper would be a solid empirical contribution.","major_comments":[{"comment":"The central claim that 'a meaningful proportion of MOOC participants rely on LLMs such as ChatGPT to produce essay answers' is not identified by the pooled before/after comparison in a single course. Section 5.3 states that demographic information is unavailable, and Section 5.1 concedes that the terminology changes could derive from 'shifting focuses of interest and public discourse applied to even a static essay prompt.' Without a comparison group, an interrupted time-series design, or within-participant panel analysis, the observed length and vocabulary shifts could equally reflect changes in who enrolls, in course materials, or in peer-review norms. The March 2023 break, rather than November 2022, further complicates attribution to ChatGPT's release. Please either add a control or placebo analysis (e.g., essays from a course on an unrelated topic over the same period) or explicitly downgrade the conclusion to a descriptive claim.","section":"Section 5.2 and Section 6"},{"comment":"All Mann-Whitney U tests treat the 56,878 essays as independent observations even though they come from only 3,582 participants, with multiple essays per participant. This violates the independence assumption and makes the reported p-values (e.g., U=165056438.0, p<0.0001 for token count) artificially small. Please aggregate at the participant level or use a mixed-effects model with participant as a random effect, and report effect sizes or standardized differences in addition to p-values. This is load-bearing for the RQ1/RQ2 significance claims, though the large raw differences may survive the correction.","section":"Section 4.1 and Section 4.2"},{"comment":"The topic-modeling analysis is described as showing no meaningful changes, but the manuscript does not report the Gensim hyperparameters (number of topics, passes, chunksize) or any measure of topic-model stability or validation. Please report these details and, if possible, provide a quantitative comparison (e.g., topic coherence or divergence) rather than relying on visual inspection, so that the negative result for RQ3 is interpretable. This comment does not affect the main descriptive findings but matters for the completeness of the RQ3 analysis.","section":"Section 3.3 and Section 4.3"}],"minor_comments":[{"comment":"There is a duplicated word in the sentence 'have mainly mainly concerned experts'; please correct it.","section":"Section 2.1"},{"comment":"The typo 'prevalance' appears in the tables and the typo 'occurences' appears in Section 4.3; the column label 'pM W U' should be formatted consistently.","section":"Tables 1 and 3 and Section 4.3"},{"comment":"The manuscript does not state whether the data or analysis code will be made available; given the novelty of the dataset, an availability statement would strengthen reproducibility.","section":"Reproducibility"},{"comment":"The figure captions state that the shaded area indicates the first and third quartiles, but the text does not define the exact period windows used for the pre/post tests; please clarify how the 'at least one year post' window maps onto the monthly bins.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The descriptive findings are plausible and likely of interest to the journal's readership. The main risk is over-claiming in the conclusion; I would not oppose publication after the causal language is tempered and the clustering issue is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core of this paper is descriptive and it holds up. They have the first longitudinal look at real MOOC essay submissions spanning the ChatGPT release: 56,878 essays from 3,582 consented participants in one AI ethics course, November 2020 to October 2024. The observed shifts are large and internally consistent: mean length up from ~150 to ~230 tokens, sentence count from 6.85 to 9.76, type-token ratio down from 0.617 to 0.577, and roughly tenfold increases in LLM-favored words like 'delve' and 'foster.' I find the descriptive claim well supported. They also deserve credit for borrowing indicator words from external prior work rather than cherry-picking their own, and for honestly flagging the demographic gap and the limitation that this is one free course on AI ethics, of all topics. The topic-modeling null result is a nice counterweight and shows they are not just reporting what they hoped to find.\n\nThe soft spots are real but mostly statistical rather than fatal. The Mann-Whitney tests treat each essay as independent, which they are not — 56,878 essays come from 3,582 participants, many of whom submitted multiple essays. That inflates significance. No multiple-comparison correction, and p-values below 0.0001 across dozens of tests are not informative. The bigger issue is the leap in Section 5.2 and the conclusion: 'a meaningful proportion of MOOC participants rely on LLMs' is an inference from a pooled before/after contrast with no control group, no within-participant panel, and no adjustment for cohort or course-context changes. The authors themselves concede in Section 5.1 that public discourse shifts could explain terminology changes even with a static prompt. The March 2023 jump rather than an immediate November 2022 effect also points to gradual diffusion or course-cycle effects. So the causal interpretation is plausible but not established.\n\nWho gets value? Anyone studying LLM impact on student writing, assessment design in MOOCs, or the external validity of the academic-abstract findings (Kobak et al., Andre et al.). The paper is honest, readable, and the descriptive findings are a useful reference point. It deserves peer review — a serious referee could push them to cluster by participant, report effect sizes with proper variance, and tone down the causal language. I would not cite it for the 'rely on LLMs' claim, but I would cite it for the longitudinal before/after data.\n\nRecommendation: send to review, with the expectation of substantial statistical revision.","headline":"A solid descriptive study of MOOC essays around ChatGPT's release; the causal reading is not secured, but the paper is worth engaging seriously.","tokens_in":11816,"tokens_out":877,"would_cite":true,"duration_ms":10234,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MOOC essays grew longer and less varied after ChatGPT","keywords":["large language models","MOOC","student essays","essay length","type-token ratio","Flesch Reading Ease","AI-generated text","longitudinal analysis"],"falsifier":"If a similar MOOC with identical prompts but no exposure to ChatGPT showed the same essay-length increase during 2020-2024, or if essay lengths were already trending upward before November 2022, the central claim would be false. More directly, if forensic AI-detection on individual essays found that the post-2023 length increase is concentrated in essays written by students who demonstrably did not use LLMs, the inference from style to LLM use would be undermined.","tokens_in":10834,"feed_emoji":"🤖","tokens_out":3766,"duration_ms":31664,"temperature":0.7,"pith_summary":"This paper asks whether the arrival of ChatGPT changed how students write in a free online course on AI ethics. Comparing 56,878 essay submissions from November 2020 to October 2024, it finds that essays submitted after ChatGPT's release are on average longer, written in shorter and simpler sentences, and draw on a narrower vocabulary. The share of tell-tale LLM words such as 'delve' and 'foster' rose roughly tenfold. The authors argue these shifts strongly suggest that a meaningful proportion of MOOC participants now rely on large language models to produce their essay answers, even though individual essays cannot be identified as AI-generated.","feed_headline":"MOOC essays grew longer and less varied after ChatGPT","feed_subtitle":"Average answers rose from 150 to 230 tokens, vocabulary narrowed, and LLM-style words like 'delve' multiplied tenfold.","key_machinery":"The argument rests on a longitudinal comparison of essay statistics before and after ChatGPT's release, with the dataset split into pre-November 2022, the first year after, and at least one year after. The carrying instruments are token and sentence counts, Flesch Reading Ease, type-token ratio, and relative prevalence of LLM-associated words, each tested for significance with the Mann-Whitney U test. The time-series plots with monthly bins show the shift arriving around March 2023 rather than gradually, which is what gives the paper its temporal claim.","core_discovery":"The central discovery is a measured before-and-after shift in student writing that lines up with ChatGPT's release. Mean answer length rose from 150.5 to 230.1 tokens, mean sentence count from 6.85 to 9.76, and mean Flesch Reading Ease from 12.70 to 15.31, while type-token ratio fell from 0.617 to 0.577. Relative prevalence of 'delve' increased 10.45-fold and 'foster' 10.84-fold. Essay topics, measured by topic modeling, stayed broadly stable. The paper interprets these statistics as signs that many students are submitting at least partially LLM-generated answers, while carefully noting that cohort changes and shifting public discourse are alternative explanations it cannot fully rule out.","pith_inferences":["A testable extension: apply an AI-detector calibrated on known human and LLM essays to this corpus and check whether the post-2023 length increase concentrates in essays flagged as machine-generated; the paper's argument predicts it should.","The same method could be applied to other written assignments, such as discussion forum posts or short-answer exams, to see whether the style shift is specific to high-stakes essays or generalizes to low-stakes writing.","If the vocabulary shift is driven by non-native speakers using LLMs as translators, as the authors speculate, then the length increase should also appear in courses taught in other languages; comparing multilingual versions of the same MOOC would separate translation use from wholesale generation.","The data imply that 'LLM-indicator' words like 'delve' may soon stop being reliable signals, because students who learn from LLM output will adopt them into genuinely human writing, so future detection work will need continuously refreshed baselines."],"forward_implications":["If the shift is genuine, MOOC providers can no longer assume peer-reviewed essays reflect a student's own writing, which pressures them to adopt proctored exams or other verification for certificate value.","Essay length and vocabulary statistics can serve as cheap, aggregate signals for monitoring LLM adoption in a course over time.","The narrowing vocabulary and rising readability imply that writing produced with LLM assistance is, on average, more homogeneous and easier to read, which may change what instructors can infer from style.","Topic stability despite term shifts suggests that students now discuss the same AI-ethics themes but with a different technical vocabulary, so syllabus content may need fewer updates than stylistic expectations do.","The near-synonym standardization (e.g., 'recommendation system' replacing 'recommender system') hints that LLM use may be pushing terminology toward one canonical form."],"supporting_citations":[{"why":"Supplies the list of LLM-associated words (delve, crucial, potential, etc.) whose relative prevalence the paper tracks.","marker":"[17]"},{"why":"Provides the expectation that AI-generated texts have a lower type-token ratio, which the paper's TTR finding mirrors.","marker":"[1]"},{"why":"Supports the observation that AI tends to produce longer text, against which the essay-length increase is interpreted.","marker":"[6]"},{"why":"Offers a comparison baseline for student versus AI-generated texts and the tendency of AI to produce longer responses.","marker":"[28]"},{"why":"Documents word-frequency changes in arXiv abstracts, used as prior evidence that LLMs affect academic writing style.","marker":"[12]"},{"why":"Supports the claim of increasing LLM use in scientific papers, which motivates the MOOC analysis.","marker":"[19]"}],"fun_headline_variants":["ChatGPT era MOOCs: longer essays, narrower vocab","Post-ChatGPT MOOC essays get longer, blander","MOOC essay style shifts after ChatGPT: longer, less diverse","LLMs reshape MOOC essays: more words, less variety","ChatGPT linked to longer, more generic MOOC essays"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pre/post comparison assumes that the student population, course materials, and peer-review behavior stayed essentially the same across the four years, so that the measured changes come from ChatGPT's arrival rather than from a shifting cohort; the authors note that demographics are unavailable.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT era MOOCs: longer essays, narrower vocab","Post-ChatGPT MOOC essays get longer, blander","MOOC essay style shifts after ChatGPT: longer, less diverse","LLMs reshape MOOC essays: more words, less variety","ChatGPT linked to longer, more generic MOOC essays"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1231,"prompt_tokens":904,"completion_tokens":327,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":247}},"tokens_in":520,"tokens_out":327,"duration_ms":3494,"temperature":1.0,"reasoning_tokens":247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:15:33.375277+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a similar MOOC with identical prompts but no exposure to ChatGPT showed the same essay-length increase during 2020-2024, or if essay lengths were already trending upward before November 2022, the central claim would be false. More directly, if forensic AI-detection on individual essays found that the post-2023 length increase is concentrated in essays written by students who demonstrably did not use LLMs, the inference from style to LLM use would be undermined.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the expectation that AI-generated texts have a lower type-token ratio, which the paper's TTR finding mirrors."},{"cited_title":"Safi and A","cited_arxiv_id":null,"evidence_quote":"Offers a comparison baseline for student versus AI-generated texts and the tendency of AI to produce longer responses."}],"review_version":1}