Pith. sign in

REVIEW 3 major objections 6 minor 62 references

Training objectives, not model scale or RLHF, drive the extreme stylistic redistribution that makes LLM text sound like AI.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 08:56 UTC pith:VJSXRF4F

load-bearing objection Solid multi-model stylometry of the AI voice; the base–instruct non-effect is useful; the λ=5.0 “beats frontier” claim is oversold on a high-perplexity 410M model. the 3 major comments →

arxiv 2605.28826 v1 pith:VJSXRF4F submitted 2026-04-08 cs.CL

From Context Shift to Stylistic Collapse: Why Training Objectives Matter More Than Scale

classification cs.CL
keywords stylistic divergenceentropy regularizationcontext shiftabsorbing stylistic statesRLHF independenceamplification ratioAI text detectionmode collapse
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that modern language models systematically reallocate probability mass over linguistic features, amplifying discourse markers and structural scaffolding by thousands of percent relative to human training corpora while suppressing complex punctuation. Across seventeen models spanning 410M to frontier scale and twenty-four probes, the same selective pattern appears. Matched base and instruction-tuned pairs show statistically indistinguishable divergence, so RLHF and instruction tuning are not the primary drivers. The author attributes the effect to deployment context shift into formal expository regimes plus self-reinforcing low-entropy “absorbing” stylistic states during autoregressive generation. Weak entropy regularization makes collapse worse; only strong regularization (λ=5.0) reduces divergence substantially and can outperform far larger frontier models on distributional naturalness. The result matters because the redistribution is invisible to ordinary quality metrics yet detectable by simple probes, with consequences for AI detection, future training data, and how human writing norms may evolve.

Core claim

Instruction-tuned and frontier LLMs systematically reallocate stylistic probability mass—amplifying discourse and structural features by mean factors of roughly 1,949–16,853 percent (peaks to ~209,675 percent) while suppressing complex punctuation to 3.2–23.2 percent of corpus baselines—and this divergence is statistically indistinguishable across matched base versus instruction-tuned pairs (p>0.25). Therefore the “AI voice” is not primarily created or worsened by RLHF; only sufficiently strong entropy regularization, not weak smoothing or scale, substantially reduces it.

What carries the argument

Amplification ratio AR_M(f) = P_M(f)/P_C(f) across a 24-feature taxonomy, together with the control-strength principle that entropy regularization L_CE − λH(P_θ) only mitigates collapse when λ is large enough (λ=5.0 works; λ=1.0 worsens it).

Load-bearing premise

That amplification ratios against Pile/Dolma baselines, measured on a thousand generations from fifteen formal expository English prompts, are a valid proxy for real deployment context shift—and that a from-scratch 410M model with high perplexity is still a fair test of whether strong regularization beats scale.

What would settle it

If matched base and instruction-tuned pairs of the same architecture, evaluated on the same 24 probes and the same prompt set, produced statistically significant differences in mean amplification (p≪0.25), or if strong entropy regularization at larger scale failed to reduce divergence relative to unregularized controls, the central claim would be falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that modern LLMs systematically reallocate stylistic probability mass relative to human training corpora: across 17 models and 24 string/regex probes, discourse and structural features are amplified by mean factors of roughly 1,949–16,853% (peaks to ~209,675%) while complex punctuation is suppressed to ~3.2–23.2% of baseline. Matched base vs instruction-tuned pairs show statistically indistinguishable divergence (p > 0.25), so the authors conclude the effect is not caused or worsened by RLHF and is instead driven by deployment context shift plus low-entropy “absorbing stylistic states.” They further claim that only strong entropy regularization (λ=5.0) during from-scratch Pythia-410M pretraining reduces divergence (40.5% improvement; 96.7–98.2% closer to AR=1 than frontier APIs) while weak λ=1.0 exacerbates collapse, establishing a “control strength principle” that training objectives matter more than scale.

Significance. If the survey results hold, the paper supplies a concrete, reproducible stylometric characterization of the “AI voice” across open and commercial systems, with clear implications for AI detection, recursive training-data contamination, and long-term linguistic drift. The base–instruct comparison is a useful corrective to narratives that pin formulaic style solely on RLHF. The entropy-regularization ablations are a genuine attempt at a training-time fix and report useful diversity metrics (distinct-n, repetition, vocab diversity). Strengths include a fully specified 24-feature taxonomy, large generation samples (1,000 per model), open experimental protocol in the appendix, and explicit limitations. The work is significant as an empirical audit even if the mechanistic story and the “beats frontier despite scale” claim require tightening.

major comments (3)
  1. Table 3 and the abstract claim that alignment “does not exacerbate” stylistic divergence because all four base–instruct pairs have p > 0.25. Three of four pairs show large numerical mean-AR increases (+1,194%, +174%, +138%); non-significance is not evidence of equivalence, especially with n=4 pairs and high cross-feature variance. The load-bearing claim that the AI voice is “upstream of alignment” and “alignment-independent” needs equivalence tests (e.g., TOST), confidence intervals on the change, or a clearer statement that the study is underpowered to detect moderate exacerbation rather than that exacerbation is ruled out.
  2. Tables 4–6 and A.6 underpin the title claim that training objectives dominate scale: pythia-410m-λ=5.0 reports distance-from-1.0 of 0.22 and is said to be 96.7–98.2% better than frontier APIs. The same tables report perplexity 786.5 (vs 48.4 at λ=0). Table 8 shows non-monotonic feature restoration and many probes still at zero in both models. Without human ratings, preference win-rates, or task metrics under the same 15 prompts, the low mean AR may reflect undertraining or quality collapse rather than successful distributional control. The assertion that “perplexity is decoupled from generation quality” is currently unsupported for the model that carries the scale-vs-objectives conclusion; either add quality evidence or substantially qualify the frontier comparison.
  3. Sections 4.1–4.3 and A.2–A.4: amplification ratios for commercial APIs (and for models without open corpora) use Pile/Dolma human baselines and 1,000 generations from 15 exclusively formal expository English prompts at temperature 0.7. That prompt set itself selects the formal-expository slice the theory calls “context shift,” so measured AR may partly be an artifact of the evaluation distribution rather than a pure property of the models. For closed models the true PC(f) is unknown. The survey claim remains directionally credible for open models with matched corpora, but the universality and magnitude claims for frontier systems need either multi-register prompts or explicit sensitivity analysis to baseline choice.
minor comments (6)
  1. Abstract vs body: abstract says “17 models”; main tables and A.4 list 13 evaluated models plus four trained Pythia variants—clarify the count consistently.
  2. Mean AR aggregation treats all 24 features equally; Appendix A.7.1 shows top-10 features drive rank correlation. State whether mean AR is unweighted and whether results are robust to top-k or category-weighted aggregation.
  3. Figure 1 caption mentions “OLMo-2-Instruct” while tables use “OLMo-1B-Instruct”; align naming.
  4. Section 3.2 presents “absorbing stylistic states” as mechanistic explanation; Limitations correctly notes lack of causal verification—consider moving stronger causal language to future work throughout the Theory section.
  5. Table 2 lists “It’s worth noting” and sentence-initial “Certainly”/“Absolutely” at 0.0 AR; clarify whether these are true zeros or below detection, and how zeros enter the mean AR.
  6. δ=0.1 and λ∈{0,0.1,1,5} are free parameters; a short sensitivity note (or pointer to ablations) would help readers assess robustness of the control-strength principle.

Circularity Check

1 steps flagged

Mild partial circularity only on diversity metrics under entropy regularization; core AR divergence survey and base-vs-instruct comparisons are independent measurements.

specific steps
  1. other [Section 3.3 / Eq. L_total = L_CE - λ·H(P_θ); Tables 4-5 and abstract claims on distinct-4 / vocab / repetition]
    "Ltotal = LCE −λ·H(Pθ) ... λ=5.0 delivers 15% higher distinct-4, 27% higher vocabulary diversity, and 78% lower repetition than moderate regularization, establishing that alignment requires sufficient control strength, not merely distributional smoothing."

    The loss term directly maximizes output entropy. Distinct-n, vocabulary diversity, and (inverse) repetition are standard proxies for that same entropy; reporting their improvement after raising λ is therefore partially tautological rather than an independent empirical prediction. (The AR-distance metric itself is not forced by the same construction, as λ=1.0 raises diversity while worsening AR.)

full rationale

The paper's primary claims rest on external empirical measurements: 24 fixed string/regex probes applied to 1,000 generations from 15 prompts, ratioed against independent Pile/Dolma corpus baselines, then compared across 13+ models and four base-instruct pairs (Tables 1-3). Those quantities are not defined in terms of the later mitigation objective and do not reduce to any fitted parameter of the models under test. The entropy-regularization experiments (Section 3.3, 4.4) introduce a mild, secondary circularity: the training objective explicitly maximizes predictive entropy (L = L_CE - λ H(P_θ)), after which the authors report higher distinct-n, higher vocabulary diversity, and lower repetition. Those particular metrics are known consequences of the entropy term and are therefore partly by construction; however, the same tables also report the independent stylistic AR distance-from-1.0 (non-monotonic across λ, with λ=1.0 actively worsening AR while still raising diversity), so the central control-strength claim is not forced. No self-citation is load-bearing, no uniqueness theorem is imported, and no ansatz is smuggled via prior work by the same author. Score remains low because the survey results and the AR component of the mitigation results stand independently of the loss definition.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 3 invented entities

The descriptive claim rests on a hand-built feature set, fixed prompts, and corpus baselines treated as PC(f). The causal story adds two mechanistic postulates (context shift; absorbing stylistic states). The mitigation claim adds a free strength parameter λ and treats distance of mean AR from 1 as the success criterion. No new physical entity is proposed; the invented pieces are conceptual mechanisms and a control-strength principle. Commercial-model conclusions further assume Pile/Dolma are adequate stand-ins for closed training mixtures.

free parameters (4)
  • entropy coefficient λ = 5.0 (selected as optimal)
    Hand-chosen grid {0, 0.1, 1.0, 5.0}; λ=5.0 declared optimal after seeing divergence/diversity trade-offs. Central mitigation claim depends on this strength.
  • divergence tolerance δ = 0.1
    Threshold for “significant” feature divergence set to 0.1 from “empirical analysis of within-human variation,” not derived.
  • generation temperature and prompt set = T=0.7; 15 fixed prompts
    Temperature 0.7 and 15 formal expository prompts define the measured deployment distribution; different sampling would change AR.
  • 24-feature taxonomy weights in mean AR = uniform mean over 24 probes
    Equal averaging of highly skewed features (headers vs rare discourse markers) is an author choice that drives headline mean amplification.
axioms (5)
  • domain assumption Pile/Dolma empirical frequencies PC(f) are valid human baselines for all evaluated models, including closed APIs.
    Section 4.1.2 and A.2 use these corpora as PC(f) even when true training data are inaccessible.
  • domain assumption Models correctly learn conditionals Pθ(x|context); observed style shift is therefore not mere frequency memorization error.
    Section 3.2.1; used to motivate context shift over “wrong training stats.”
  • ad hoc to paper Amplification ratio AR_M(f)=P_M(f)/P_C(f) and |D_M|/|F|>0.5 formalize “stylistic collapse.”
    Section 3.1 definitions; the success metric for both survey and mitigation.
  • ad hoc to paper Non-significant base–instruct differences (p>0.25) imply alignment does not exacerbate divergence.
    Table 3 / abstract; treats failure to reject null as positive evidence of no effect.
  • domain assumption L = L_CE − λ H(Pθ) is an appropriate training-time fix for generation-time collapse.
    Section 3.3; extends classification entropy penalties to LM pretraining without proving optimality.
invented entities (3)
  • absorbing stylistic states no independent evidence
    purpose: Explain why low-entropy structural features self-reinforce across tokens and produce extreme AR.
    Section 3.2.3 Markov/absorbing-region story; consistent with patterns but not causally validated (Limitations).
  • context shift (training vs deployment formal-expository slice) no independent evidence
    purpose: Primary mechanism for selective amplification/suppression without blaming RLHF.
    Section 3.2.2; inferred from prompt regime and feature pattern, not measured via controlled context interventions.
  • control strength principle (weak λ harms, strong λ helps) no independent evidence
    purpose: Summarize non-monotonic regularization results as a general alignment requirement.
    Sections 4.4 and 5; based on four λ values on one 410M architecture.

pith-pipeline@v1.1.0-grok45 · 26556 in / 4110 out tokens · 68938 ms · 2026-07-13T08:56:12.600853+00:00 · methodology

0 comments
read the original abstract

In modern LLMs, linguistic features function not as stylistic artifacts but as probes of probability mass, allocated under training alignment objectives. Language models trained with contemporary pipelines exhibit severe reshaping of linguistic features, leading to extreme language re-distribution. While previous stylometric analyses explored linguistic differences between AI-generated and human texts, we focus on the reshaping plaguing the LLM training pipeline itself. We analyze 17 models (410M-100B+ parameters) across 24 linguistically-motivated probes, documenting that instruction-tuned systems systematically collapse language entropy along discourse and structural dimensions (mean amplification: 1,949-16,853%, peaks: 5,181-209,675%), while selectively suppressing complex punctuation to 3.2-23.2% of baseline frequencies. These effects do not worsen under RLHF, as divergence patterns are statistically indistinguishable (p > 0.25) across matched base and instruction-tuned model pairs. Weak intervention (lambda=1.0) exacerbates collapse by 240%, while strong control (lambda=5.0) achieves 40.5% improvement and outperforms frontier models by 96.7-98.2% despite 200-1000x scale disadvantage. Additionally, lambda=5.0 delivers 15% higher distinct-4, 27% higher vocabulary diversity, and 78% lower repetition than moderate regularization, establishing that alignment requires sufficient control strength, not merely distributional smoothing. Our findings underscore how modern LLMs reallocate stylistic probability mass, despite RLHF and scale. More broadly, our work reveals a structural limitation of current alignment pipelines: preference optimization reshapes language distributions invisible to standard quality metrics yet detectable through distributional probes, with implications for AI detection, training data contamination, and long-term linguistic evolution.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

62 extracted references · 3 canonical work pages

  1. [1]

    Hype vs reality in the integration of artificial intelligence in clinical workflows.JMIR Formative Research, 9:e70921, 2025

    Alaa Abd-Alrazaq, Barry Solaiman, Yosra Magdi Mekki, Dena Al-Thani, Faisal Farooq, Metab Alkubeyyer, Mohamed Ziyad Abubacker, Rawan AlSaad, Sarah Aziz, Ahmed Serag, Rajat Thomas, Javaid Sheikh, and Arfan Ahmed. Hype vs reality in the integration of artificial intelligence in clinical workflows.JMIR Formative Research, 9:e70921, 2025. doi: 10.2196/70921

  2. [2]

    Claude 3 model card.Anthropic Technical Report, 2024

    Anthropic. Claude 3 model card.Anthropic Technical Report, 2024

  3. [3]

    Constitutional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073, 2022

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073, 2022

  4. [4]

    The multi-dimensional approach to linguistic analyses of genre variation: An overview of methodology and findings.Computers and the Humanities, 26(5/6):331–345, 1992

    Douglas Biber. The multi-dimensional approach to linguistic analyses of genre variation: An overview of methodology and findings.Computers and the Humanities, 26(5/6):331–345, 1992

  5. [5]

    Pythia: A suite for analyzing large language models across training and scaling.arXiv preprint arXiv:2304.01373, 2023

    Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling.arXiv preprint arXiv:2304.01373, 2023

  6. [6]

    Drift no more? context equilibria in multi-turn llm interactions, 10 2025

    Vardhan Dongre, Ryan Rossi, Viet Lai, Seunghyun Yoon, Dilek Hakkani-Tur, and Trung Bui. Drift no more? context equilibria in multi-turn llm interactions, 10 2025

  7. [7]

    The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  8. [8]

    Fraser, Hillary Dawkins, and Svetlana Kiritchenko

    Kathleen C. Fraser, Hillary Dawkins, and Svetlana Kiritchenko. Detecting ai-generated text: Factors influencing detectability with current methods.J. Artif. Int. Res., 82, June 2025. ISSN 1076-9757. doi: 10.1613/jair.1.16665. URLhttps://doi.org/10.1613/jair.1.16665

  9. [9]

    Riley Galpin, Bryce Anderson, and Tom S. Juzek. Exploring the structure of ai-induced language change in scientific english.arXiv preprint arXiv:2506.21817, 2025

  10. [10]

    The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020

  11. [11]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.arXiv preprint arXiv:2403.05530, 2024

    Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.arXiv preprint arXiv:2403.05530, 2024

  12. [12]

    Generative adversarial nets

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. InAdvances in Neural Information Processing Systems, volume 27, 2014

  13. [13]

    Olmo: Accelerating the science of language models.arXiv preprint arXiv:2402.00838, 2024

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models.arXiv preprint arXiv:2402.00838, 2024

  14. [14]

    The curious case of neural text degeneration

    Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. InInternational Conference on Learning Representations, 2020

  15. [15]

    Gpt-4o system card.arXiv preprint arXiv:2410.21276, 2024

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, et al. Gpt-4o system card.arXiv preprint arXiv:2410.21276, 2024

  16. [16]

    Hasan M. S. Jaashan and Wagdi Rashad Ali Bin-Hady. Stylometric analysis of ai-generated texts: a comparative study of chatgpt and deepseek.Cogent Arts & Humanities, 12(1):2553162, 2025. doi: 10.1080/23311983.2025.2553162. URLhttps://doi.org/10.1080/23311983.2025.2553162

  17. [17]

    Mistral 7b

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023

  18. [18]

    Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

  19. [19]

    A survey of temporal drift in large language models.arXiv preprint, 2025

    Sushil Khairnar. A survey of temporal drift in large language models.arXiv preprint, 2025. 10

  20. [20]

    Understanding the effects of rlhf on llm generalisation and diversity.arXiv preprint arXiv:2309.02926, 2023

    Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefen- stette, and Roberta Raileanu. Understanding the effects of rlhf on llm generalisation and diversity.arXiv preprint arXiv:2309.02926, 2023

  21. [21]

    Wesley W. Koo. Cross-lingual effects of ai-generated content on human work.Scientific Reports, 15(1): 30949, 2025. doi: 10.1038/s41598-025-16650-w

  22. [22]

    Preserving diversity in supervised fine-tuning of large language models

    Ziniu Li, Congliang Chen, Tian Xu, Zeyu Qin, Jiancong Xiao, Zhi-Quan Luo, and Ruoyu Sun. Preserving diversity in supervised fine-tuning of large language models. InInternational Conference on Learning Representations, 2025

  23. [23]

    Linguistic differences between ai and human comments in weibo: Detect ai- generated text through stylometric features

    Ziqi Li and Qi Zhang. Linguistic differences between ai and human comments in weibo: Detect ai- generated text through stylometric features. InProceedings of the 24th China National Conference on Computational Linguistics, pages 842–851, Jinan, China, 2025

  24. [24]

    Helpful, harmless, honest? sociotechnical limits of ai alignment and safety through reinforcement learning from human feedback.arXiv preprint, 2023

    Adam Dahlgren Lindström, Leila Methnani, Lea Krause, Petter Ericson, Íñigo Martínez de Rituerto de Troya, Dimitri Coelho Mollo, and Roel Dobbe. Helpful, harmless, honest? sociotechnical limits of ai alignment and safety through reinforcement learning from human feedback.arXiv preprint, 2023

  25. [25]

    Benchmark of stylistic variation in llm-generated texts

    Jiˇrí Miliˇcka, Anna Marklová, and Václav Cvrˇcek. Benchmark of stylistic variation in llm-generated texts. arXiv preprint arXiv:2509.10179, 2025

  26. [26]

    Detectgpt: Zero-shot machine-generated text detection using probability curvature

    Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. Detectgpt: Zero-shot machine-generated text detection using probability curvature. InInternational Conference on Machine Learning, 2023

  27. [27]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback....

  28. [28]

    Regularizing neural networks by penalizing confident output distributions

    Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton. Regularizing neural networks by penalizing confident output distributions. InInternational Conference on Learning Representations, 2017

  29. [29]

    Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118, 2024

    Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, et al. Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118, 2024

  30. [30]

    Unmasking ai- generated texts using linguistic and stylistic features.International Journal of Advanced Computer Science and Applications, 16(3):213–221, 2025

    Muhammad Irfaan Hossen Rujeedawa, Sameerchand Pudaruth, and Vusumuzi Malele. Unmasking ai- generated texts using linguistic and stylistic features.International Journal of Advanced Computer Science and Applications, 16(3):213–221, 2025

  31. [31]

    A mathematical theory of communication.The Bell System Technical Journal, 27(3): 379–423, 1948

    Claude E Shannon. A mathematical theory of communication.The Bell System Technical Journal, 27(3): 379–423, 1948

  32. [32]

    Ai models collapse when trained on recursively generated data.Nature, 631:755–759, 2024

    Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. Ai models collapse when trained on recursively generated data.Nature, 631:755–759, 2024. doi: 10.1038/ s41586-024-07566-y

  33. [33]

    Dolma: An open corpus of three trillion tokens for language model pretraining research.arXiv preprint arXiv:2402.00159, 2024

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al. Dolma: An open corpus of three trillion tokens for language model pretraining research.arXiv preprint arXiv:2402.00159, 2024

  34. [34]

    The homogenizing effect of large language models on human expression and thought

    Zhivar Sourati et al. The homogenizing effect of large language models on human expression and thought. Trends in Cognitive Sciences, 2026. doi: 10.1016/j.tics.2026.01.003

  35. [35]

    Linguistic characteristics of ai-generated text: A survey.arXiv preprint, 2025

    Luka Terˇcon and Kaja Dobrovoljc. Linguistic characteristics of ai-generated text: A survey.arXiv preprint, 2025

  36. [36]

    Dsdr: Dual-scale diversity regularization for exploration in llm reasoning.arXiv preprint arXiv:2602.19895, 2025

    Zhongwei Wan, Yun Shen, Zhihao Dou, Donghao Zhou, Yu Zhang, Xin Wang, Hui Shen, Jing Xiong, Chaofan Tao, Zixuan Zhong, Peizhou Huang, and Mi Zhang. Dsdr: Dual-scale diversity regularization for exploration in llm reasoning.arXiv preprint arXiv:2602.19895, 2025

  37. [37]

    The price of format: Diversity collapse in LLMs

    Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang. The price of format: Diversity collapse in LLMs. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 15454–15468, Suzhou, China, 2025. 11

  38. [38]

    Understanding deep learning requires rethinking generalization

    Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. InInternational Conference on Learning Representations, 2017

  39. [39]

    Delve into

    Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R. Tomz, Christopher D. Manning, and Weiyan Shi. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity.arXiv preprint arXiv:2510.01171, 2025. A Technical Appendix A.1 Feature Taxonomy Our 24-feature taxonomy spans four categories: Punctuation Patterns (5 features): • Em das...

  40. [40]

    Pearson r >0.95 confirms stable extraction

    Test-retest reliability: Extract features from same documents with 1-week interval. Pearson r >0.95 confirms stable extraction

  41. [41]

    Delve into

    Baseline diversity: Measure coefficient of variation across corpus samples. High baseline variance indicates natural human variation, validating that models amplify beyond normal ranges. A.2.4 Baseline Statistics Feature extraction is highly reliable, with test–retest correlations of r= 0.997 (Pile) and r= 0.999 (Dolma). Both corpora exhibit substantial s...

  42. [42]

    Write a detailed analysis of the benefits and drawbacks of remote work in modern society

    "Write a detailed analysis of the benefits and drawbacks of remote work in modern society."

  43. [43]

    Explain the complex relationship between technology and privacy in the digital age

    "Explain the complex relationship between technology and privacy in the digital age."

  44. [44]

    Discuss the potential impacts of artificial intelligence on the job market over the next decade

    "Discuss the potential impacts of artificial intelligence on the job market over the next decade."

  45. [45]

    Analyze the key factors contributing to climate change and potential solutions

    "Analyze the key factors contributing to climate change and potential solutions."

  46. [46]

    Compare and contrast different approaches to education reform in the 21st century

    "Compare and contrast different approaches to education reform in the 21st century."

  47. [47]

    Examine the role of social media in shaping public opinion and political discourse

    "Examine the role of social media in shaping public opinion and political discourse."

  48. [48]

    Discuss the ethical implications of genetic engineering and CRISPR technology

    "Discuss the ethical implications of genetic engineering and CRISPR technology."

  49. [49]

    Analyze the economic and social effects of globalization on developing nations

    "Analyze the economic and social effects of globalization on developing nations."

  50. [50]

    Explore the relationship between mental health and modern lifestyle factors

    "Explore the relationship between mental health and modern lifestyle factors."

  51. [51]

    Discuss the challenges and opportunities of renewable energy adoption

    "Discuss the challenges and opportunities of renewable energy adoption."

  52. [52]

    Examine the impact of streaming services on traditional media industries

    "Examine the impact of streaming services on traditional media industries."

  53. [53]

    Analyze the factors that contribute to successful entrepreneurship in tech startups

    "Analyze the factors that contribute to successful entrepreneurship in tech startups." 14

  54. [54]

    Discuss the implications of automation and robotics on manufacturing industries

    "Discuss the implications of automation and robotics on manufacturing industries."

  55. [55]

    Explore the concept of work-life balance in contemporary professional culture

    "Explore the concept of work-life balance in contemporary professional culture."

  56. [56]

    Analyze the role of regulation in cryptocurrency and blockchain technology

    "Analyze the role of regulation in cryptocurrency and blockchain technology." Prompts are cycled to generate 1,000 samples, ensuring topic diversity while maintaining sufficient per-prompt sample size for robust statistics. A.4 Models We evaluate 13 models: Open-Source Base Models (10 models): • EleutherAI/pythia-410m (410M parameters) [5] • allenai/OLMo-...

  57. [57]

    In conclusion

    "In conclusion": 5,048%

  58. [58]

    Delve into

    "Delve into": 3,660%

  59. [59]

    Bullet points: 3,063%

  60. [60]

    Numbered lists: 1,949% 18

  61. [61]

    Fundamentally

    "Fundamentally": 570%

  62. [62]

    However" (sentence-initial): 332% These 10 features span structural (headers, lists), discourse (

    "However" (sentence-initial): 332% These 10 features span structural (headers, lists), discourse ("in conclusion", "however"), and tonal markers ("landscape", "navigate", "robust"), confirming that AI voice manifests across multiple linguistic dimensions. A.7.2 Normalization Strategy ObjectiveTo validate percentage-based amplification ratios over alternat...