{"id":"9073f149-d5e8-4e3a-8826-63b3f2311e12","arxiv_id":"2411.18382","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ChatGPT-generated French presidential addresses differ systematically from real ones in word-category use, vocabulary choices, and sentence-length variation.","lead":"Researchers compared real New Year's speeches by four French presidents with speeches generated by ChatGPT after feeding it the real speeches as models. They found ChatGPT's French is measurably different: more nouns, possessives, and numbers, fewer verbs, pronouns, and adverbs, and more uniform sentence lengths.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed ChatGPT style may be an artifact of comparing short generated texts (avg. 835 words) with much longer presidential speeches (avg. 1,547 words), absent any length-matched human control.","rationale":"The reader's weakest_assumption correctly identifies the lack of disclosure of prompt, model version, and temperature as a serious reproducibility threat. However, I see a more directly load-bearing internal-validity concern: the comparison is confounded by text length. The paper does not include any length-matched human control, so the POS distribution differences and the reduced sentence-length variance could stem from the simple fact that ChatGPT produces shorter texts rather than from a stable 'ChatGPT style.' This concern is independent of the prompt/version issue and would persist even if the generation details were fully documented. It is testable using only the existing corpus, by extracting matched-length segments from the presidential addresses. The reader's concern and mine are complementary; both would need to be addressed before the broad claims are accepted. Since the reader already issued a CONDITIONAL verdict, my recommendation is UNCHANGED: the paper remains conditionally acceptable, with the additional condition that a length-matched analysis be performed and the generation setup be disclosed.","tokens_in":14608,"tokens_out":9067,"duration_ms":89512,"concrete_test":"Re-run the core comparisons using length-matched presidential texts. For each of the 20 presidential speeches, randomly sample one or more contiguous or random excerpts whose total word count equals the corresponding Ngpt (e.g., an 835-word segment from each NT). Recompute Tables 2, 3, and 10 on the matched NT excerpts versus the original GPT outputs. If the overuse/underuse of POS categories shrinks to non-significance or the sentence-length CV% becomes similar, the claimed ChatGPT style is not robust. As a secondary check, also report per-speech POS frequencies to ensure the aggregate result is not driven by a few outlier generations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that ChatGPT overuses nouns, possessive determiners, and numbers while underusing verbs, pronouns, and adverbs, and produces more standardized sentence lengths—rests on an aggregate comparison of all 30,935 words of presidential addresses with 16,699 words of ChatGPT outputs. The average presidential speech is 1,547 words; the average generated text is 835 words. No control is made for text length, which is known to correlate with lexical density, pronoun use, and syntactic variation. In particular, the sentence-length analysis in Section 6 compares a 30,935-word corpus with a 16,699-word corpus; the lower variance and Gaussian-like distribution of sentence lengths in GPTs could indicate that shorter texts are more uniform in sentence length, not that ChatGPT has a distinctive style. The claim in Section 7 is even more directly affected: GPT corpora per president are between 2,843 and 5,235 words, while the corresponding NT corpora are between 5,631 and 9,818 words, so the intertextual-distance trees (Figures 2-3) are computed on unequal-length texts, and the paper's own caveat that distances cannot be computed on texts shorter than 1,000 words is not addressed for individual GPT outputs. If the observed POS differences are reproduced when comparing like-for-like text lengths, the conclusion would hold; otherwise the 'ChatGPT style' characterization would be a confound of output length. This is an internal-validity issue, independent of the additionally unspecified prompt wording, model version, and temperature that the reader flagged.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a corpus study of 20 ChatGPT-generated French presidential New Year's addresses, each prompted with a real presidential address as a model, compared with the 20 original addresses. Using lemmatized POS tagging, lemma frequency tables, sentence-length distributions, and intertextual distance, the authors report that ChatGPT overuses nouns, possessive determiners, and numbers; underuses verbs, pronouns, and adverbs; produces more uniform sentence lengths; and that intertextual distance fails to distinguish the generated texts from their models when a single homogeneous text is provided. The paper concludes that ChatGPT cannot easily be detected by existing distance-based methods.","tokens_in":14913,"tokens_out":4495,"duration_ms":35728,"significance":"The study addresses an important and timely question: whether LLM-generated political speech can be stylistically distinguished from human-authored text. The descriptive frequency tables are useful and the authors provide a transparent corpus of French political speech. If the findings survive rigorous controls, they would inform both stylometry and AI-text detection. However, the current analysis lacks basic experimental controls (text length, prompt variability, model version) and the significance index is under-specified, so the central claims are plausible but not yet established.","major_comments":[{"comment":"The aggregate comparison of POS frequencies between the NT corpus (30,935 words) and the GPT corpus (16,699 words) does not control for text length. Because the average GPT output is 835 words versus 1,547 for a presidential address (Table 1), and because lexical density, pronoun frequency, and syntactic complexity are known to vary with text length, the reported differences (e.g., common nouns +7%, adverbs -32%) may be an artifact of output length rather than a stable property of ChatGPT. Please include a length-matched human baseline (e.g., first 835 words of each presidential speech) or otherwise demonstrate that the effects persist when text length is held constant.","section":"Section 4, Table 2"},{"comment":"The paper does not report the exact prompts used to generate the GPT texts, the model version (e.g., GPT-3.5 vs GPT-4), the sampling temperature, or the number of independent generations per presidential speech. Without this information, the results are not reproducible, and the abstract's general characterization of 'ChatGPT style' may apply only to a particular snapshot and prompting strategy. The authors should disclose the full generation protocol and ideally run multiple independent generations to assess stability.","section":"Section 3"},{"comment":"The sentence-length analysis compares entire corpora of different sizes and does not control for text length. The lower standard deviation and coefficient of variation in GPTs may simply reflect the fact that shorter texts are more uniform in sentence length; a length-matched comparison with human-written texts is needed. Additionally, the claim that GPT sentence lengths follow a Gaussian distribution is not tested with a formal normality test, and the mode/median/mean equality is merely stated.","section":"Section 6, Table 10 and Figure 1"},{"comment":"The intertextual distance analysis merges all GPT outputs per president (2,843-5,235 words) to circumvent the 1,000-word minimum distance requirement that the authors themselves note. While this avoids individual short texts, the merged GPT corpora are still much shorter than the corresponding NT corpora (5,631-9,818 words), so the distances are computed on unequal-length texts. More importantly, the conclusion that 'intertextual distance is therefore no longer able to identify texts generated by ChatGPT' is an overgeneralization: only one condition (a single homogeneous model text) was tested, with no control condition such as generation without a model or with multiple heterogeneous examples. The conclusion should be restricted to the tested setting, and the current wording in the abstract should be qualified.","section":"Section 7"},{"comment":"The significance index S is undefined. The text states thresholds for 5% and 1% risk but does not give the formula or the underlying statistical test. Without this, the reader cannot assess whether S-values such as 0.009 or 0.977 support the claimed significant differences. Please define S explicitly (e.g., the probability under a stated null hypothesis) and justify the thresholds.","section":"Tables 2-9"}],"minor_comments":[{"comment":"The phrase 'It is well-kwon' contains a typo; it should be 'well-known'.","section":"Section 6"},{"comment":"The word 'mounth' should be 'month'.","section":"Table 7"},{"comment":"The entries 'firs' and 'eigth' should be 'first' and 'eighth'.","section":"Table 9"},{"comment":"The word 'ensembe' should be 'ensemble' (or 'together' in the English gloss).","section":"Table 6"},{"comment":"The sentence 'the characteristics of ChatGPT's vocabulary (see Section 4)' points to Section 4, which is about POS; vocabulary lemmas are actually discussed in Section 5. Please correct the cross-reference.","section":"Section 3"},{"comment":"The phrase 'ChatGPT has a little trouble with simple verbs' is informal; consider 'some difficulty with simple verbs'.","section":"Section 5.1"},{"comment":"The axes are not fully labeled: the x-axis should be 'sentence length (words)' and the y-axis 'percentage of sentences'.","section":"Figure 1"},{"comment":"The reference to 'Vaswami et al. 2017' contains a typo; it should be 'Vaswani et al. 2017'.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is topical and the corpus is a useful resource, but the experimental design needs substantial strengthening before it can support the strong claims in the abstract. I recommend major revision rather than rejection because the descriptive tables contain information that could be salvaged with appropriate controls. I would ask the authors to reveal the full generation protocol and add length-matched baselines."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful descriptive profile of ChatGPT's French output, but the specific style claims are underdetermined by the design. The POS and lemma tables are carefully assembled from cleaned, lemmatized texts, and the observation that ChatGPT normalizes sentence-length distributions toward a Gaussian shape is genuinely interesting. The finding that it overuses nouns, possessives, and numbers while avoiding pronouns, adverbs, and subordinate clauses is plausible and worth knowing.\n\nThe soft spots are proportionate to the claims. The paper never says which ChatGPT version was used, what prompt was sent, what temperature, or how many runs. Each prompt ran once, so there's no estimate of generation variability. And there is no length-matched control: the generated corpus is 16,699 words versus 30,935 for the presidents, with generated texts averaging about half the length. The stress-test about length is not a kill shot, but it is a valid concern. Shorter texts can show different POS densities and less sentence-length spread; without a like-for-like human control, the 'ChatGPT style' characterization may be partly a confound of output length. It should be tested.\n\nSection 7 overreaches. Showing that intertextual distance clusters GPT texts near their single model text is a neat result, but the conclusion that 'intertextual distance is no longer able to identify texts generated by ChatGPT' is too strong for one configuration and one metric. The paper's own caveat about the 1,000-word minimum is brushed aside by merging per president, and the merged corpora are still unequal in length.\n\nWho gets value? Someone working on AI-text detection in French, or on stylometric profiles of LLMs. It deserves a serious referee because the descriptive core is built on real data and standard methods, and the weaknesses are fixable rather than fatal. But it needs major revision: disclose the generation setup, add length-matched human controls, report repeated sampling, and soften the Section 7 claim.\n\nI'd send it out, with the caveat that the referee should push for the controls and the disclosure.","headline":"Useful descriptive profile of ChatGPT's French output, but missing generation details and length-matched controls leave the style claims underdetermined.","tokens_in":15406,"tokens_out":3110,"would_cite":false,"duration_ms":28836,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChatGPT-written French presidential speeches are statistically distinct from real ones, yet intertextual-distance attribution no longer flags them when a single model speech is supplied.","keywords":["ChatGPT","stylometry","authorship attribution","French presidential speeches","intertextual distance","sentence length distribution","part-of-speech analysis","large language models"],"falsifier":"Re-run the same presidential speech-generation protocol on several dated ChatGPT versions with temperature and prompts recorded; if the overuse of nouns, possessive determiners, and numbers, the underuse of pronouns, adverbs, and subordinate clauses, and the compression of sentence lengths do not reproduce, the claimed signature is version- or prompt-specific.","tokens_in":14416,"feed_emoji":"🇫🇷","tokens_out":17547,"duration_ms":138922,"temperature":0.7,"pith_summary":"This paper asks whether a machine can ghostwrite a French presidential New Year's address and whether the result is detectable. By lemmatizing (reducing words to dictionary forms) and statistically comparing twenty real end-of-year speeches by Chirac, Sarkozy, Hollande, and Macron with ChatGPT outputs produced after submitting each speech as a model, the authors identify a consistent machine profile: ChatGPT overuses nouns, possessive determiners, and numbers, underuses verbs, pronouns, and adverbs, and generates sentence lengths that cluster around the average instead of showing the spread natural to human prose. The paper also shows that when given one homogeneous model text, ChatGPT imitates that author closely enough that intertextual distance, a classical authorship-attribution measure, no longer separates the generated text from the model. The stakes are practical: if the profile is stable, conventional plagiarism and authorship-detection tools need to be rethought for LLM-generated French text.","feed_headline":"ChatGPT speeches are noun-heavy, verb-shy, and too regular","feed_subtitle":"A comparison with real presidential speeches shows the pattern, and classic authorship tests now miss it.","key_machinery":"The argument is carried by three quantitative instruments applied to corpora whose words were reduced to dictionary lemmas. Part-of-speech densities per thousand words expose the noun-versus-verb imbalance. Rank-frequency tables of the most frequent lemmas in the verb, pronoun, adverb, noun, adjective, and determiner categories reveal which words ChatGPT over- or under-uses. Sentence-length distributions, summarized by mode, median, mean, standard deviation, coefficient of variation, and the ninth-to-first-decile ratio, show that generated texts cluster around the mean. The defining diagnostic is the inequality $\\text{mode} < \\text{median} < \\text{mean}$, which the authors treat as a property of natural sentence-length distributions and which ChatGPT's near-Gaussian output violates. For the detection question, the central object is intertextual distance, a measure between zero and one computed from the absolute differences between lemma frequencies divided by total text length; below a threshold it supports single authorship, and the paper shows it no longer separates GPT from human text when one homogeneous model is supplied.","core_discovery":"On its own terms, the paper establishes that ChatGPT's French writing carries a measurable stylistic signature, not just isolated oddities. Compared with the presidents' addresses, generated texts lean toward the noun side of the language: common nouns rise by about 7%, adjectives by 13%, possessive determiners by 30%, and numbers by 36%, while verb-linked categories shrink, with pronouns down 22%, adverbs down 32%, and subordinating conjunctions down 25%. Lexically, ChatGPT overuses devoir, continuer, nous, année, défi, and valiant, and underuses être, vouloir, falloir, aller, dire, and third-person pronouns. The clearest single signature is sentence length: real speeches obey the natural inequality $\\text{mode} < \\text{median} < \\text{mean}$ and have a coefficient of variation near 78%, whereas ChatGPT's near-Gaussian distribution yields $\\text{mode} \\approx \\text{median} \\approx \\text{mean}$ and pulls the coefficient down to about 50%. Despite these divergences, ChatGPT given a single homogeneous model text produces output that intertextual distance cannot flag as machine-generated, so the paper concludes this classical detection method is no longer adequate.","pith_inferences":["Because the corpus holds one generated output per presidential speech, a natural next test is to sample many outputs from the same prompt under different model versions, temperatures, and random seeds; this would quantify the stylistic variance behind the reported averages.","The overuse of 'défi' and the feminized formula 'nos concitoyennes et nos concitoyens' suggests ChatGPT generalizes beyond the prompt from generic presidential discourse in its training data, so detectors aimed at stereotyped collocations (such as 'relever les défis qui') might work where frequency profiles fail.","If the sentence-length compression ($\\text{mode} \\approx \\text{median} \\approx \\text{mean}$) generalizes beyond French, a simple variance-based readability statistic could serve as a cheap, language-independent screening test for generated prose.","A more demanding test would be corpus-level authorship attribution, comparing the generated text against many writings of the suspected author rather than a single supplied model; the paper's negative result concerns the single-model case."],"forward_implications":["ChatGPT's French prose has a describable statistical baseline: noun-heavy, verb-shy, with fewer pronouns, adverbs, and subordinate clauses than human political writing.","Real presidential addresses and ChatGPT outputs can be told apart on part-of-speech and sentence-length grounds, even when the machine was given a real speech as a model.","Intertextual-distance authorship attribution, which previously detected machine-generated and fraudulent scientific papers, fails when the prompt supplies one homogeneous model text.","Plagiarism detectors that compare n-gram overlaps will likely miss ChatGPT output because the generator rearranges the model's vocabulary rather than copying it.","The paper itself notes these limitations may be mitigated in future generator versions and calls for further experiments on detection features and other languages."],"supporting_citations":[{"why":"This work supplies the lemmatization and orthographic standardization rules applied to both corpora before any comparison.","marker":"Muller (1977)"},{"why":"This work defines the intertextual distance and its quality index used in Section 7 to test detection.","marker":"Labbé & Labbé, 2006"},{"why":"This work provides the stylometry framework and style indicators that motivate the part-of-speech, lexicon, and sentence-length analyses.","marker":"Savoy, 2020"},{"why":"This work shows that French-language GPT detection is feasible with trained classifiers but degrades on new domains, the baseline this study extends.","marker":"Antoun et al., 2023"},{"why":"This work shows that ChatGPT-generated abstracts are hard for experts and detectors to identify, motivating the detection experiment.","marker":"Gao et al. (2023)"},{"why":"This work supplies the claim that GPT output is essentially indistinguishable from human text, which the paper tests against presidential speeches.","marker":"Bubeck et al. (2023)"},{"why":"This work documents earlier successful intertextual-distance detection of generated or fraudulent scientific papers that ChatGPT now eludes.","marker":"Byrne & Labbé, 2016"}],"fun_headline_variants":["ChatGPT's style: noun-heavy, verb-shy, too uniform","AI speechwriter beats old detection test for French","Presidential AI: more nouns, fewer verbs, no surprise","Machine vs. Macron: style gap that tricks detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the outputs collected from chat.openai.com, whose model version, temperature, and exact prompts are not reported, represent a stable 'ChatGPT style' rather than an artifact of one model snapshot or one prompting strategy.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT's style: noun-heavy, verb-shy, too uniform","AI speechwriter beats old detection test for French","Presidential AI: more nouns, fewer verbs, no surprise","Machine vs. Macron: style gap that tricks detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000228,"raw_usage":{"total_tokens":1515,"prompt_tokens":1025,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":422}},"tokens_in":641,"tokens_out":490,"duration_ms":5794,"temperature":1.0,"reasoning_tokens":422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:15:26.402898+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same presidential speech-generation protocol on several dated ChatGPT versions with temperature and prompts recorded; if the overuse of nouns, possessive determiners, and numbers, the underuse of pronouns, adverbs, and subordinate clauses, and the compression of sentence lengths do not reproduce, the claimed signature is version- or prompt-specific.","supporting_citations":[],"review_version":1}