{"id":"1a3daacb-3a08-4f4d-9897-54665020a741","arxiv_id":"2606.22735","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Corpus analysis of 1,600+ student statistics reports shows post-LLM writing became more similar to LLM output in style and verbs, most in intro/conclusion sections, while also aligning more with expert writing.","lead":"Researchers analyzed over 1,600 undergraduate statistics reports from 2021 to 2025 and found students' writing style and verb use shifted toward LLM patterns, especially in introductions and conclusions. This suggests statistics educators may need new assessment approaches that emphasize thinking over standard report formats.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Attribution of style shifts to LLM adoption rests on an uncontrolled pre/post comparison across 2021-2025.","rationale":"The reader's weakest_assumption is precisely the load-bearing identification gap. No other internal inconsistency or measurement detail rises to the same level given the available abstract and the explicit before/after framing.","tokens_in":1728,"tokens_out":289,"duration_ms":10276,"concrete_test":"Re-fit the main similarity regressions after adding year-by-topic fixed effects and any available student-level covariates (GPA, major, prior stats courses); if the post-2023 LLM-similarity coefficient drops by more than one standard error or loses significance, the attribution is not robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim equates temporal change in embedding/verb-based similarity to LLMs (and to experts) with student LLM use. The design compares reports from 2021-2022 versus 2023-2025 without reported controls or instruments for concurrent changes in course topics, instructor pool, student demographics, or prompt wording. Because the similarity metrics are computed on the same sections (quintiles 1 and 5) that are most sensitive to assignment framing, any shift in those framings would produce the observed pattern even if no student used an LLM. The additional claim of increased expert similarity does not resolve the identification problem.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript analyzes a corpus of over 1,600 undergraduate students' data analysis reports spanning 2021–2025. It claims that post-LLM emergence, students' writing style and verb usage have shifted toward greater similarity with LLMs (most pronounced in the first and fifth quintiles, corresponding to introductions and conclusions) while simultaneously becoming more similar to statistics experts' writing. The authors discuss pedagogical implications and propose alternative assessments focused on statistical thinking.","tokens_in":1871,"tokens_out":488,"duration_ms":27642,"significance":"If the shifts can be credibly attributed to LLM adoption rather than other temporal factors, the work would supply quantitative evidence on AI's section-specific effects on student statistical writing and offer practical suggestions for assessment redesign. The large corpus and quintile-based segmentation provide a replicable framework for tracking style changes in quantitative disciplines.","major_comments":[{"comment":"The central identification strategy relies on an uncontrolled pre/post comparison (2021–2022 vs. 2023–2025) without reported covariates or instruments for changes in course topics, instructor pool, student demographics, or assignment prompts. Because the largest reported shifts occur precisely in quintiles 1 and 5 (the sections most sensitive to framing), any evolution in prompt wording or curriculum emphasis could generate the observed pattern even in the absence of LLM use by students.","section":"Methods / corpus construction and analysis period"},{"comment":"The dual finding of increased similarity both to LLMs and to experts is presented without reconciliation or auxiliary evidence; the manuscript does not test whether the expert-similarity trend is independent of the same uncontrolled temporal factors or whether the two similarity measures are collinear.","section":"Results section on expert similarity"}],"minor_comments":[{"comment":"The abstract supplies no detail on the embedding or verb-based similarity metrics, the statistical tests employed, sample-selection criteria, or handling of multiple comparisons; these should be stated explicitly even in the abstract.","section":"Abstract"},{"comment":"Provide a clearer operational definition of the quintiles and the exact mapping to report sections, along with any robustness checks that vary the number of quintiles.","section":"Quintile analysis description"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments. We address each major point below, acknowledging limitations in the observational design while outlining targeted revisions to clarify the analysis and strengthen the discussion of alternative explanations.","responses":[{"response":"We agree that the pre/post design is observational and lacks explicit controls for potential confounders such as curriculum changes or prompt evolution. The timing of the observed shifts aligns with LLM availability, and the section-specific pattern (strongest in quintiles 1 and 5) is consistent with LLM-assisted framing, but we cannot rule out other temporal factors. In revision we will add a dedicated limitations subsection that explicitly discusses these alternative explanations and note the absence of detailed prompt metadata. If any assignment prompt records exist in the corpus, we will conduct a sensitivity check; otherwise the limitation will be stated plainly.","revision_made":"partial","referee_comment":"The central identification strategy relies on an uncontrolled pre/post comparison (2021–2022 vs. 2023–2025) without reported covariates or instruments for changes in course topics, instructor pool, student demographics, or assignment prompts. Because the largest reported shifts occur precisely in quintiles 1 and 5 (the sections most sensitive to framing), any evolution in prompt wording or curriculum emphasis could generate the observed pattern even in the absence of LLM use by students."},{"response":"The manuscript reports the two trends as separate empirical observations without asserting causal independence. We will add a results subsection that computes the correlation between LLM-similarity and expert-similarity scores to assess collinearity. The revised discussion will note that greater expert similarity could partly reflect LLM assistance in producing clearer prose, while acknowledging that the same uncontrolled temporal factors could influence both measures. This clarification will be incorporated without overclaiming.","revision_made":"yes","referee_comment":"The dual finding of increased similarity both to LLMs and to experts is presented without reconciliation or auxiliary evidence; the manuscript does not test whether the expert-similarity trend is independent of the same uncontrolled temporal factors or whether the two similarity measures are collinear."}],"tokens_in":1336,"tokens_out":449,"duration_ms":22967,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core finding is that student data analysis reports from 2023-2025 look more like LLM output than earlier ones, especially in the first and last fifths of the text, while also moving closer to expert writing. They built this from a corpus of 1600-plus undergrad reports spanning 2021-2025 and applied style and verb-based comparisons.\n\nThat corpus size and the section-specific look are the useful parts. It gives a concrete before-and-after snapshot in one course context and flags a practical issue for stats teaching.\n\nThe main weakness is that the design is a simple time split with no controls or measures for other changes that happened in the same window. Curriculum tweaks, different prompts, shifts in who enrolls, or instructor turnover could all move the same sections toward more formulaic language without students using LLMs at all. The expert-similarity result does not fix the attribution problem. The abstract gives little detail on the exact similarity metrics or statistical checks, so it is hard to judge how robust the numbers are.\n\nThis is the sort of paper that stats and data science instructors who care about writing assignments would want to read and discuss. It is not a methods advance, but it surfaces a timely empirical pattern worth checking.\n\nI would send it to peer review. Referees can push on the controls and robustness steps, and the question itself is worth airing even if the causal claim needs tightening.","headline":"The paper shows measurable style and verb shifts in student stats reports after 2022 that track LLM patterns, but the pre/post design leaves the cause unclear.","tokens_in":2340,"tokens_out":367,"would_cite":false,"duration_ms":13476,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Undergraduate statistics students' writing has become more similar to large language models since 2021, especially in report introductions and conclusions, while also growing closer to expert style.","keywords":["student writing","large language models","statistics education","writing style analysis","verb usage","undergraduate reports","AI in education","data analysis reports"],"falsifier":"A controlled study that assigns identical reports to students with and without LLM access and then applies the same style and verb analysis would show whether the measured shifts appear only when LLMs are available.","tokens_in":2642,"feed_emoji":"","tokens_out":613,"duration_ms":19025,"temperature":0.7,"pith_summary":"The paper examines a corpus of more than 1,600 undergraduate data analysis reports collected from 2021 to 2025 to measure changes in writing after LLMs became widely available. It reports that student style and verb choices now align more closely with LLM output, with the strongest shifts appearing in the first and fifth quintiles that correspond to introductions and conclusions. The same analysis shows student writing has also moved closer to the style used by statistics experts. These patterns raise questions for how statistics and data science courses should assess written communication of results.","feed_headline":"Student stats reports now mirror LLM output more closely","feed_subtitle":"Analysis of 1600 reports from 2021-2025 finds style shifts toward AI in introductions and conclusions, plus greater expert alignment.","key_machinery":"Quintile-based comparison of writing style similarity and verb usage between pre- and post-LLM student reports.","core_discovery":"Analysis of the 1,600-report corpus establishes that students' writing style and verb usage have shifted toward patterns typical of LLMs, most noticeably in the opening and closing sections of reports, while the same writing has simultaneously become more similar to that of statistics experts.","pith_inferences":["If the style convergence continues, distinguishing student-authored text from LLM-assisted text may require new methods beyond surface similarity checks.","Departments could test whether requiring oral defenses of written reports restores emphasis on the underlying statistical reasoning.","Longer-term tracking of the same students after graduation could reveal whether the LLM-influenced style persists in professional statistical communication."],"forward_implications":["Statistics educators should explore assessment formats that still require students to structure statistical arguments even if LLMs handle phrasing.","Targeted writing tasks focused on report introductions can isolate and reinforce the cognitive work of framing results.","The dual convergence toward both LLM and expert styles suggests LLMs may be helping students adopt professional conventions faster."],"fun_headline_variants":["Students' stats writing echoes LLM patterns in intros conclusions","1600 reports show student style shifting toward LLMs and experts","LLM rise alters undergrad stats report verb use and style","Analysis finds stats students write more like AI and experts"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The observed changes in style and verb use result from students adopting LLMs rather than from shifts in curriculum, teaching methods, student population, or assignment design during the same period.","fun_headline_variants_meta":{"raw":{"variants":["Students' stats writing echoes LLM patterns in intros conclusions","1600 reports show student style shifting toward LLMs and experts","LLM rise alters undergrad stats report verb use and style","Analysis finds stats students write more like AI and experts"]},"model":"grok-4.3","cost_usd":0.005817,"raw_usage":{"total_tokens":2656,"prompt_tokens":605,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":58165500,"prompt_tokens_details":{"text_tokens":605,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1987,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":605,"tokens_out":64,"duration_ms":15813,"temperature":1.0,"reasoning_tokens":1987,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T06:44:41.166676+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled study that assigns identical reports to students with and without LLM access and then applies the same style and verb analysis would show whether the measured shifts appear only when LLMs are available.","supporting_citations":[],"review_version":1}