{"id":"3514fa35-08dc-49a4-ace6-a41f26e400a4","arxiv_id":"2608.06589","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLM outputs show both a cohort-level generational shift between 2024 and 2026 models and model-specific stylistic signatures, supporting the idiolectal framing.","lead":"This chapter argues that the text produced by each AI chatbot is not one generic 'AI language' but a distinct, stable linguistic style, like a human idiolect. The authors analyze model outputs from 2024 and 2026 and find a shared generational shift in style while each model keeps its own signature.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-topic evidence cannot establish idiolectal stability: PCA (§5) and contraction analyses (§6) both use only the health-misinformation topic, yet the paper's own definition (§2.1) requires stability across topics and contexts, leaving the central analogy untested.","rationale":"The paper does substantial, careful empirical work: a controlled 2026 corpus, transparent cleaning and truncation handling, an unusually thorough treatment of apostrophe-variant contamination (§6.2), and — a genuine strength — two independent Mistral-7B replication runs that land near the original data on PC1, showing the method can recover a consistent signature across runs with different seeds. The contraction-composition profiles in Table 6.2 provide a compelling qualitative demonstration that models within one cohort differ in kind, not just in rate. My concern does not dispute these observations; it targets the step from 'differences on health-misinformation essays' to the paper's central theoretical label. The paper's own definition of idiolect requires stability across topics and contexts (§2.1), and that property is asserted, not demonstrated: every stylometric result (the PC1 generational divide at 37.6% of variance, contraction rates and profiles) is computed on the single health-misinformation topic. The four-topic corpus already exists, so the deficiency is one of analysis, not data. I agree with the reader that this precludes the observed differences from supporting the full idiolect claim: the profiles could equally be topic-register effects, especially since several PC1 loading words (e.g., misinformation, truth, real, uncertainty) are semantically tied to the health-misinformation domain. The secondary concerns — no inferential separability tests, the ~25% contraction difference between the original Mistral data and its reruns, and PCA sample counts that correlate with output length (30 vs 125 samples) — reinforce the need for conditional acceptance, but the cross-topic check is the one that would settle the applicability of the construct. For those reasons my read does not change the reader's CONDITIONAL verdict; the confirmed gap is the same one, and the requested cross-topic validation is the decisive test.","tokens_in":19371,"tokens_out":20587,"duration_ms":165936,"concrete_test":"Run the §5.1 PCA (150 MFW, culled at 50%, correlation matrix, 10,000-word non-overlapping samples) and the §6.1 contraction pipeline separately on the other three topics in the existing corpus (climate, global warming, maths anxiety), then compare each model's coordinates and contraction rates across the four topic-specific analyses. The §2.1 stability requirement is met if model positions and contraction rankings stay approximately fixed across topics — for instance, if a classifier trained on health-misinformation samples labels samples from the other three topics with accuracy well above chance and cross-topic correlations of per-model contraction rates are high. If model orderings or profiles reorder substantially across topics, the observed profiles are topic-conditional and the central claim should be weakened to domain-specific style rather than idiolect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The chapter's central claim is that LLM output is idiolectal in the sense defined in §2.1: a profile that 'stays relatively stable across topics and contexts' and is unique to the individual. The load-bearing empirical evidence for both uniqueness and stability comes exclusively from one topic. §5.1 restricts the stylometric PCA to health misinformation 'to ensure that any stylistic differences reflect model behaviour rather than topic effects' — a defensible design that removes topic variance instead of testing it — and §6's contraction analysis uses the same single-domain corpus. The model-specific word-frequency and contraction profiles in Fig. 5.1 and Tables 6.1–6.2 could therefore be topic-register effects: each model may simply have a characteristic way of writing about health misinformation, not a stable cross-topic voice. The data needed to test stability already exists: §3.1 reports four topics per model, and §4 uses all four for length, TTR, and disclaimer descriptors, but no cross-topic stylometric or contraction comparison is reported. Three further features amplify the exposure: (i) no separability or significance tests accompany the PCA, so 'unique profile' rests on visual ellipse inspection, with overlapping ellipses explicitly acknowledged; (ii) the 2024-vs-2026 generational axis (PC1, 37.6% of variance) is also measured only on health misinformation, so the claimed register shift (e.g., negation contractions, §6.3) may not generalise; (iii) the original Mistral-7B contraction rate (2,955.9 per million) differs by ~25% from the two replication runs (3,630.6 and 3,688.1), setting a noise floor for the model-specific 'signature'. The Mistral replication is a genuine strength for cross-run PCA stability (reruns land near the original, Euclidean distance ~1.7 in the PC1–PC2 plane, mostly on PC2), but it tests reproducibility under fixed settings, not stability across topics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the output of individual LLM-based tools is better analysed as a model-specific idiolect than as a single collective 'AI language'. Using the 2024 TextualLLMap corpus and a newly generated 2026 corpus with the same prompts and topics, the authors report computational descriptors (length, sentence length, TTR, vocabulary size, disclaimers), a stylometric PCA on word-frequency profiles, and a targeted contraction analysis. They find a generational shift on PC1 (37.6% of variance) separating most 2024 from most 2026 models, while also reporting model-specific profiles, illustrated most dramatically by contraction frequencies ranging from 120.2 to 30,611.9 per million words. The chapter concludes that the idiolect notion is empirically warranted and useful for variation, detection, and forensic linguistics.","tokens_in":19680,"tokens_out":2467,"duration_ms":25442,"significance":"If its central claim holds, the paper makes a useful contribution by shifting attention from an undifferentiated 'AI language' to model-specific, stable stylistic profiles, with concrete implications for LLM-generated text detection, forensic attribution, and usage-based accounts of variation and change. The study has notable strengths: it uses a newly generated 2026 corpus with prompts and settings matched as closely as possible to the 2024 corpus; it includes a Mistral-7B replication control that lands in nearly identical PCA coordinates; it documents and handles a genuine technical problem (the multiplicity of Unicode apostrophe glyphs in LLM output); and the empirical analyses are direct measurements of newly generated text rather than quantities derived to force the conclusion. The principal weakness is that the empirical core—PCA and contraction analyses—is restricted to a single topic (health misinformation), so the paper's own definition of idiolect as stable 'across topics and contexts' is not tested.","major_comments":[{"comment":"The paper's definition of idiolect requires a profile that 'stays relatively stable across topics and contexts' (Section 2.1), but the load-bearing evidence for distinct model profiles comes exclusively from one topic. The PCA in Section 5.1 deliberately restricts the corpus to health misinformation, and the contraction analysis in Section 6.1 uses the same single-domain corpus. The result could therefore be a topic-register effect—each model having a characteristic way of writing about health misinformation—rather than a stable cross-topic voice. The data needed to test stability already exist: Section 3.1 states that four topics were generated per model, and Section 4 uses all four for length, TTR, and disclaimer analyses, but no cross-topic stylometric or contraction comparison is reported. Without such a comparison, the central idiolect claim overreaches the evidence.","section":"§2.1, §5.1, §6.1"},{"comment":"The claim that 'each individual model maintains a unique linguistic profile' is not supported by any statistical separability or significance test. The interpretation rests on visual inspection of 95% confidence ellipses in Figure 5.1, and the figure's own alt-text acknowledges that some ellipses are 'larger, more elongated, or overlapping.' For overlapping models, 'unique profile' is not demonstrated. I would ask for a quantitative separation analysis—for example, a permutation test on pairwise Mahalanobis distances, a MANOVA on the component scores, or a cross-validated classifier accuracy—so that readers can see which model pairs are actually distinguishable and at what confidence level.","section":"§5.2, Fig. 5.1"},{"comment":"The contraction results, while vividly illustrating variation, are also based only on the health-misinformation corpus and are reported as aggregate corpus-level counts per model. The strong generational claim for negation contractions ('t, a ratio of roughly 32:1 between 2026 and 2024 means) is thus subject to the same single-topic limitation as the PCA. It is also heavily driven by one model, Claude-Haiku-4-5; the authors do note that excluding it leaves a 9:1 ratio, but no confidence intervals or per-text variability measures are provided. Treating each model's full output as a single pooled corpus does not allow assessment of within-model consistency across texts or topics, which is precisely what the idiolect claim requires.","section":"§6.3, Table 6.1"}],"minor_comments":[{"comment":"In the answer to RQ3, the text says 'the dominant finding from the PCA (Section 7)' but the PCA results are in Section 5, not Section 7; this cross-reference should be corrected.","section":"§7"},{"comment":"Model names are inconsistent across the manuscript: the 2024 cohort is described as including 'Llama-3-8B' in Section 3.1 but appears as 'Llama-3.8B' in Table 6.1, and the 2026 GPT model is rendered both as 'GPT-5.4 Mini' (Section 3.1) and 'GPT-5-4-Mini' (Sections 5 and 6). Please standardise model names.","section":"§3.1, Table 6.1, Table 6.2"},{"comment":"The enumeration of apostrophe glyphs is clear and useful, but it would help to state explicitly whether the custom regular-expression pattern builder was validated against a manually annotated sample beyond the AntConc checks, and how many texts were manually inspected.","section":"§6.2"},{"comment":"The sentence reporting that 'within-speaker variation in a given domain is greater than within-model variation, while between-speaker variation is greater than between-model variation' is presented without supporting numbers or a reference to a table or figure; this is a substantial claim and should either be quantified in the text or moved to the repository with a pointer.","section":"§4.2(iv)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serviceable contribution to the emerging literature on LLM stylometry, and the single-topic weakness is fixable within the manuscript's scope because the four-topic corpus already exists. I would encourage the editor to request the cross-topic analysis and the separability tests before publication, as these are directly load-bearing for the idiolect framing. I saw no evidence of circularity: the analyses are direct measurements of newly generated text against an external 2024 corpus."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a useful empirical contribution wrapped in a conceptual claim that outruns its evidence. The new 2026 corpus, generated with the same prompts as TextualLLMap, is real value: six contemporary models, four topics, careful cleaning, and a human-comparison side analysis. The diachronic angle—comparing 2024 vs 2026 cohorts on the same prompts—is the genuinely new piece, and the contraction study is methodologically careful, especially the handling of eight distinct apostrophe glyphs. The Mistral rerun control is a good idea and shows the stylometric method is reproducible under fixed settings.\n\nThe soft spots are real and central. The paper argues that LLM output is idiolectal, where idiolect is defined as relatively stable across topics and contexts. But the PCA and contraction analyses—the evidence for 'each individual model maintains a unique linguistic profile'—come exclusively from the health-misinformation topic. The authors deliberately restrict to one topic 'to ensure that stylistic differences reflect model behaviour rather than topic effects,' but that removes topic variance rather than testing it. Section 4 does use all four topics for length and TTR, so it is not that no cross-topic data exist; it is that the load-bearing stylometric evidence does not. Without a cross-topic comparison, the observed differences could be topic-register effects.\n\nThere are also no separability or significance tests. The PCA ellipses overlap for several models, and the 'unique profile' claim rests on visual inspection. Contraction rates are point estimates; the Mistral replication points to a noise floor (original 2,955.9 per million vs reruns 3,630.6 and 3,688.1, roughly 25% apart) that the paper does not discuss. None of this sinks the generational-shift observation—PC1 separating 2024 from 2026 is descriptive and plausible—but it does mean the strong idiolect claim is conditional, not established.\n\nWho is this for? Computational linguists and stylometry folks studying AI text, and anyone building detectors or attribution tools. The corpus and replication protocol deserve referee time. I would send it to review, but with the clear requirement that the authors either add cross-topic stylometric and contraction analyses or soften the central claim accordingly.","headline":"A genuinely new 2026 corpus and a clean diachronic comparison, but the idiolect claim overreaches because stability across topics is never tested.","tokens_in":20362,"tokens_out":3475,"would_cite":true,"duration_ms":32877,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM output is not one 'AI language': each model carries a stable idiolect, and the 2026 generation shares a new stylistic voice.","keywords":["LLM idiolects","stylometry","principal component analysis","contractions","language variation and change","AI-generated text","corpus linguistics","model generations"],"falsifier":"Recompute the same PCA and contraction counts on texts the same 13 models produce for a second, unrelated topic (for instance, climate change) under the same prompt-and-sampling protocol. If the 95% confidence ellipses of different models overlap on that topic, or if PC1 no longer separates the 2024 from the 2026 cohort, the observed 'idiolects' are topic-bound registers rather than stable model voices.","tokens_in":19145,"feed_emoji":"🤖","tokens_out":9098,"duration_ms":79806,"temperature":0.7,"pith_summary":"The paper argues that the output of a single large language model is best understood not as part of one collective 'AI language' but as a model-specific variety similar to a human idiolect: a stable, distinctive way of using language. It backs this with two corpora of model-generated essays on societal topics, one from 2024 and one newly generated from 2026 using the same prompts. Stylometric principal component analysis separates the two cohorts on the first axis, while each model occupies its own region of the stylistic space. A contraction-frequency case study sharpens the point: rates range from 120 to over 30,000 per million words, and models differ in which contraction types they favor. If right, this reframes AI-text comparison, detection, and forensic attribution around individual tools and their versions rather than a single generic 'AI voice'.","feed_headline":"Each LLM has its own idiolect—and 2026 models share a new voice","feed_subtitle":"Stylometric profiles separate every model, and a generation gap splits 2024 from 2026 output.","key_machinery":"The load-bearing concept is the LLM idiolect: an individual model's distinctive, stable, and emergent way of using language, following the Richards, Platt and Platt definition. It is made measurable by two instruments. First, a stylometric PCA on the 150 most frequent words, culled at 50% and computed on a correlation matrix, using 10,000-word samples from the health-misinformation topic; PC1 separates generations and per-model ellipses measure internal consistency. Second, a contraction taxonomy of six categories ('t, 's, 'm, 're, 've, 'd) searched with a custom script covering eight apostrophe glyphs, giving normalized per-million-word rates that expose model-specific fingerprints. The replication rerun of Mistral-7B with different seeds is the control showing that the recovered signature is stable across generation runs.","core_discovery":"The central claim is that each LLM-based tool has its own idiolect—defined as an emergent, relatively stable, unique way of communicating—and that alongside these individual idiolects there is a cohort-level generational shift. On health-misinformation essays, PCA of the 150 most frequent words separates all six 2024 models on the positive side of PC1 (37.6% of variance) and five of six 2026 models on the negative side, with only Mistral-Nemo-12B placed with the older cohort. Within that space each model's samples form its own ellipse; re-running Mistral-7B with different seeds lands near the original, indicating a recoverable signature. Contraction counts reinforce the picture: overall frequencies span 120.2 to 30,611.9 per million words, negation contractions are roughly 32 times more frequent in the 2026 cohort, and individual models show distinct category profiles, such as Gemini-3-Flash skewing toward \"n't\" while other models favor \"'s\". The paper concludes the idiolect frame captures stability, uniqueness, and emergence better than style or a single AI super-variety.","pith_inferences":["A cross-topic replication using the same 13 models on a second, unrelated prompt set would test the definitional claim of stability: if per-model ellipses separate within a topic but merge across topics, the reported profiles are topic-register effects.","If the idiolects survive cross-topic testing, model versions become usable as 'speakers' for corpus-based studies of language change, letting researchers trace family-level stylistic drift across releases and generations.","The contraction taxonomy is a cheap attribution signal—Gemini's negation-heavy profile versus Claude's broad contraction use—but it needs validation across prompts, temperatures, and lengths before it can support forensic claims.","Some of the 2026-vs-2024 shift could be an artifact of deployment rather than training-data drift, since closed-API models are queried through vendor systems that may add their own preprompts; a controlled open/closed comparison with identical system-prompt strings would isolate this."],"forward_implications":["LLM-generated text detection and attribution can treat each model version as a separate author class instead of hunting for a single generic 'AI voice.'","Human-vs-AI comparative studies should name the specific model and version, since the 120-to-30,000-per-million contraction spread shows one cohort contains radically different defaults.","The clean 2024/2026 split on PC1 means model generation is a real variable: detectors and style corpora calibrated on older output will misjudge newer text unless time-stamped.","Variationist and usage-based linguistics can study model families as lineages, with family-level continuities occupying a middle level between the individual idiolect and the super-variety."],"supporting_citations":[{"why":"Supplies the 2024 corpus of 28,000 texts from six models plus the prompts and generation settings that the 2026 cohort replicates.","marker":"Improta et al. (2024)"},{"why":"Provides the stylo R package implementing the PCA on word-frequency profiles that produces the model ellipses and PC1/PC2 loadings.","marker":"Eder, Rybicki & Kestemont (2015)"},{"why":"Supplies the definition of idiolect as an individual's stable, distinguishing way of communicating that the paper applies to LLM output.","marker":"Richards, Platt & Platt (1992)"},{"why":"Earlier exploratory delta-method study of ChatGPT versus Gemini that first framed individual LLM tools as having idiolects.","marker":"Rudnicka (2025)"},{"why":"Survey of 44 studies documenting the field's ChatGPT-centred bias and the formal, impersonal register findings this chapter positions itself against.","marker":"Terčon & Dobrovoljc (2025)"},{"why":"Establishes the comparison that human-human stylistic variation exceeds human-LLM variation, motivating the search for model-specific idiolects.","marker":"Zamaraeva et al. (2025)"}],"fun_headline_variants":["LLMs have idiolects, not one AI language","Every LLM speaks its own dialect—2026 shifts together","No single AI voice: each model speaks an idiolect","Model-specific idiolects, plus a 2026 generational shift"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes a model's idiolect stays relatively stable across topics and contexts, but all the profile evidence comes from one topic (health misinformation), so the model-specific profiles could be topic-register effects rather than stable individual voices.","fun_headline_variants_meta":{"raw":{"variants":["LLMs have idiolects, not one AI language","Every LLM speaks its own dialect—2026 shifts together","No single AI voice: each model speaks an idiolect","Model-specific idiolects, plus a 2026 generational shift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000747,"raw_usage":{"total_tokens":3346,"prompt_tokens":982,"completion_tokens":2364,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":2290}},"tokens_in":598,"tokens_out":2364,"duration_ms":16408,"temperature":1.0,"reasoning_tokens":2290,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:16:48.694160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the same PCA and contraction counts on texts the same 13 models produce for a second, unrelated topic (for instance, climate change) under the same prompt-and-sampling protocol. If the 95% confidence ellipses of different models overlap on that topic, or if PC1 no longer separates the 2024 from the 2026 cohort, the observed 'idiolects' are topic-bound registers rather than stable model voices.","supporting_citations":[],"review_version":1}