Pith. sign in

REVIEW 2 major objections 2 minor 12 references

Analyzing Students' Statistics Writing Before and After the Emergence of Large Language Models

T0 review · 2 major / 2 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read Undergraduate statistics students' writing has become more similar to large language models since 2021, especially in report introductions and conclusions, while also growing closer to expert style.

desk verdict The paper shows measurable style and verb shifts in student stats reports after 2022 that track LLM patterns, but the pre/post design leaves the cause unclear. read the letter →

arxiv 2606.22735 v1 pith:S2KMVDSJ submitted 2026-06-22 stat.AP stat.OT

classification stat.APstat.OT
keywords studentwritinglargelanguagemodelsstatisticseducationstyleanalysisverbusageundergraduatereportsAIindata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines a corpus of more than 1,600 undergraduate data analysis reports collected from 2021 to 2025 to measure changes in writing after LLMs became widely available. It reports that student style and verb choices now align more closely with LLM output, with the strongest shifts appearing in the first and fifth quintiles that correspond to introductions and conclusions. The same analysis shows student writing has also moved closer to the style used by statistics experts. These patterns raise questions for how statistics and data science courses should assess written communication of results.

What carries the argument

Quintile-based comparison of writing style similarity and verb usage between pre- and post-LLM student reports.

What would settle it

A controlled study that assigns identical reports to students with and without LLM access and then applies the same style and verb analysis would show whether the measured shifts appear only when LLMs are available.

Watch

Extended reading notes

Core claim

Analysis of the 1,600-report corpus establishes that students' writing style and verb usage have shifted toward patterns typical of LLMs, most noticeably in the opening and closing sections of reports, while the same writing has simultaneously become more similar to that of statistics experts.

Load-bearing premise

The observed changes in style and verb use result from students adopting LLMs rather than from shifts in curriculum, teaching methods, student population, or assignment design during the same period.

Editorial extensions

If this is right

  • Statistics educators should explore assessment formats that still require students to structure statistical arguments even if LLMs handle phrasing.
  • Targeted writing tasks focused on report introductions can isolate and reinforce the cognitive work of framing results.
  • The dual convergence toward both LLM and expert styles suggests LLMs may be helping students adopt professional conventions faster.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the style convergence continues, distinguishing student-authored text from LLM-assisted text may require new methods beyond surface similarity checks.
  • Departments could test whether requiring oral defenses of written reports restores emphasis on the underlying statistical reasoning.
  • Longer-term tracking of the same students after graduation could reveal whether the LLM-influenced style persists in professional statistical communication.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript analyzes a corpus of over 1,600 undergraduate students' data analysis reports spanning 2021–2025. It claims that post-LLM emergence, students' writing style and verb usage have shifted toward greater similarity with LLMs (most pronounced in the first and fifth quintiles, corresponding to introductions and conclusions) while simultaneously becoming more similar to statistics experts' writing. The authors discuss pedagogical implications and propose alternative assessments focused on statistical thinking.

Significance. If the shifts can be credibly attributed to LLM adoption rather than other temporal factors, the work would supply quantitative evidence on AI's section-specific effects on student statistical writing and offer practical suggestions for assessment redesign. The large corpus and quintile-based segmentation provide a replicable framework for tracking style changes in quantitative disciplines.

major comments (2)
  1. [Methods / corpus construction and analysis period] The central identification strategy relies on an uncontrolled pre/post comparison (2021–2022 vs. 2023–2025) without reported covariates or instruments for changes in course topics, instructor pool, student demographics, or assignment prompts. Because the largest reported shifts occur precisely in quintiles 1 and 5 (the sections most sensitive to framing), any evolution in prompt wording or curriculum emphasis could generate the observed pattern even in the absence of LLM use by students.
  2. [Results section on expert similarity] The dual finding of increased similarity both to LLMs and to experts is presented without reconciliation or auxiliary evidence; the manuscript does not test whether the expert-similarity trend is independent of the same uncontrolled temporal factors or whether the two similarity measures are collinear.
minor comments (2)
  1. [Abstract] The abstract supplies no detail on the embedding or verb-based similarity metrics, the statistical tests employed, sample-selection criteria, or handling of multiple comparisons; these should be stated explicitly even in the abstract.
  2. [Quintile analysis description] Provide a clearer operational definition of the quintiles and the exact mapping to report sections, along with any robustness checks that vary the number of quintiles.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments. We address each major point below, acknowledging limitations in the observational design while outlining targeted revisions to clarify the analysis and strengthen the discussion of alternative explanations.

read point-by-point responses
  1. Referee: The central identification strategy relies on an uncontrolled pre/post comparison (2021–2022 vs. 2023–2025) without reported covariates or instruments for changes in course topics, instructor pool, student demographics, or assignment prompts. Because the largest reported shifts occur precisely in quintiles 1 and 5 (the sections most sensitive to framing), any evolution in prompt wording or curriculum emphasis could generate the observed pattern even in the absence of LLM use by students.

    Authors: We agree that the pre/post design is observational and lacks explicit controls for potential confounders such as curriculum changes or prompt evolution. The timing of the observed shifts aligns with LLM availability, and the section-specific pattern (strongest in quintiles 1 and 5) is consistent with LLM-assisted framing, but we cannot rule out other temporal factors. In revision we will add a dedicated limitations subsection that explicitly discusses these alternative explanations and note the absence of detailed prompt metadata. If any assignment prompt records exist in the corpus, we will conduct a sensitivity check; otherwise the limitation will be stated plainly. revision: partial

  2. Referee: The dual finding of increased similarity both to LLMs and to experts is presented without reconciliation or auxiliary evidence; the manuscript does not test whether the expert-similarity trend is independent of the same uncontrolled temporal factors or whether the two similarity measures are collinear.

    Authors: The manuscript reports the two trends as separate empirical observations without asserting causal independence. We will add a results subsection that computes the correlation between LLM-similarity and expert-similarity scores to assess collinearity. The revised discussion will note that greater expert similarity could partly reflect LLM assistance in producing clearer prose, while acknowledging that the same uncontrolled temporal factors could influence both measures. This clarification will be incorporated without overclaiming. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: observational corpus comparison with direct empirical measurements.

full rationale

The paper conducts a pre/post corpus analysis of student reports (2021-2025) using embedding-based similarity and verb usage counts to compare against LLM and expert corpora. No equations, fitted parameters presented as predictions, self-definitional constructs, or load-bearing self-citations appear in the described methods or claims. The central findings rest on direct measurement of temporal shifts in the data itself rather than any reduction to inputs by construction. Attribution concerns (confounding by curriculum changes) are validity issues, not circularity.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The work is an empirical corpus study; the abstract mentions no free parameters, mathematical axioms, or newly postulated entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Analyzing Students' Statistics Writing Before and After the Emergence of Large Language Models." pith.science (2026). https://pith.science/paper/S2KMVDSJ

@misc{pith2026260622735,
  author       = {Pith},
  title        = {Pith review of: Analyzing Students' Statistics Writing Before and After the Emergence of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S2KMVDSJ}},
  note         = {Machine review of arXiv:2606.22735}
}
read the original abstract

The ability to communicate statistical results to domain experts and stakeholders is an important goal of the undergraduate statistics and data science curriculum. However, as large language models (LLMs) have become more accessible, a major concern is that students are offloading important cognitive tasks to generative AI. Using a corpus of over 1,600 undergraduate students' data analysis reports from 2021 to 2025, we show how students' writing style and verb usage have become more similar to that of LLMs. This shift is most pronounced in the first and fifth quintiles of students' reports, which roughly map onto the introduction and conclusion sections, respectively. At the same time, we demonstrate that students' writing style has become more similar to that of statistics experts with the addition of LLMs. We end by discussing the implications of our findings for statistics and data science educators. In particular, we propose alternative modes of assessment that still emphasize statistical thinking, such as targeted writing assignments for structuring a report introduction.

Figures

Figures reproduced from arXiv: 2606.22735 by the authors.

Figure 1
Figure 1. Confusion matrix on the test data for the overall LDA model, which compares [PITH_FULL_IMAGE:figures/full_fig_p017_1.png] view at source ↗
Figure 2
Figure 2. Projection of pre-LLM student reports, LLM-generated text, and expert writing onto the two linear discriminants from the overall LDA model. From [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. Contour plot of the linear discriminant scores for each text source, faceted by [PITH_FULL_IMAGE:figures/full_fig_p021_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Contour plot of the first and second linear discriminant scores for each text source, [PITH_FULL_IMAGE:figures/full_fig_p023_4.png]
Figure 5
Figure 5. Figure 5: Contour plot of the first and second linear discriminant scores for each text [PITH_FULL_IMAGE:figures/full_fig_p024_5.png]
Figure 6
Figure 6. Figure 6: The rate of use in the student corpus of sixteen lemmas that appeared in the top [PITH_FULL_IMAGE:figures/full_fig_p028_6.png]
Figure 7
Figure 7. Figure 7: Fifteen ways in which underscore was used in since-LLM student reports. The lemma underscore was never used in pre-LLM student reports. 28 [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

12 extracted references

  1. [1]

    Don’t spend the whole week on the intro and the EDA parts

    Save enough time to do the modeling. Don’t spend the whole week on the intro and the EDA parts. Get into the model building/selection, and save time afterwards for the writing. The first model you try probably won’t be your final chosen model; you will need time to look at diagnostic plots and inspect terms for significance, maybe try transformations, bui...

  2. [2]

    When working with real data, there isn't such a thing as a perfect model, but there can be useful ones

    Keep a balance between complexity and simplicity/interpretability . When working with real data, there isn't such a thing as a perfect model, but there can be useful ones. These datasets aren’t ‘nice’ like those in homeworks/labs/lectures/example -project, and your diagnostics/R 2/etc might be surprisingly bad, comparatively. Transformations will likely h...

  3. [3]

    Be sure that after you knitted your fi nal version, you read the full report

    Read your knitted report before you submit . Be sure that after you knitted your fi nal version, you read the full report. Check if your conclusions match t he presented model, and that you didn’t accidentally use a disca rded model for your conclusions. Check for typos, and if the text and figures are all there, that you remembered to title your plots an...

  4. [4]

    my project

    Write as if it is a report you would submit to a client or an academic journal, not like a homework or class assignment or a snapchat to a friend: a. Your title should be something interesting and meaningful that captures the reader’s attention and has something to do with the topic and the work (not “my project” or something similar that you might title ...

  5. [5]

    we did the following analysis

    Be consistent in your writing style. Don’t switch between past tense (like, “we did the following analysis”) versus present tense (like, “we do the following analysis”)

  6. [6]

    You should be creating your own words/phrases/transitions/paragraphs/etc

    Don’t copy verbatim sentences, phrases, or title from our posted topic prompts. You should be creating your own words/phrases/transitions/paragraphs/etc . Also the motivation for the topic that we wrote in the posted prompts isn’t the only possible motivation

  7. [7]

    echo” all of the code. In the markdown template formatting, we have set the “echo

    You don’t need to “echo” all of the code. In the markdown template formatting, we have set the “echo” to TRUE, which will make all your code chunks show up in your knitted document (it “echoes” the code chunks in the knitted document) , since seeing your code can sometimes help us in grading. But i n published research you generally would not show all the...

  8. [8]

    character

    Don’t copy/paste from other text editors into R. This could cause your document to not knit, if the text editor used different character encoding. If you get a knitting error that looks like: , then you have a character the knitter doesn’t understand, and you will need to find and delete it. [What the character is will appear in the error notification aft...

Show all 12 references
  1. [9]

    Be cautious copy/pasting between text areas and code chunks. This could similarly cause your document to not knit, for instance if some text character isn’t understood in code, or if some code symbol isn’t understood by the math font editor in the text area. [See item #8, abov...

  2. [10]

    This could similarly cause your document to not knit, if your keyboard uses different character encoding

    Be cautious using different keyboard encoding. This could similarly cause your document to not knit, if your keyboard uses different character encoding. [See item #8, above, for details.]

  3. [11]

    Put extra line breaks (i.e., hit “enter” a few times ) between graphs and text as needed; or use \ newline. If, when you knit the document, you find that graphs or text seem to be pushed off the sides of the page, it’s probably because the knitter is interpreting some graphs a...

  4. [12]

    common-plotting-issues-in-R.pdf

    To make a barplot, or to control the font size of parts of common R graphs , or to re- order bars on a barplot, see the posted document “common-plotting-issues-in-R.pdf” in the project1-materials folder on Canvas ; and to control whitespace or page breaks in your knitted docum...

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.