{"id":"b91cdc14-dafd-4ea6-8a4f-e2cb15eb4f92","arxiv_id":"2411.18365","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"ChatGPT-3.5-generated State of the Union speeches are stylistically distinguishable from real presidential addresses across word, sentence, part-of-speech, and rhetorical measures.","lead":"This paper compared State of the Union speeches generated by ChatGPT-3.5 with real addresses from four US presidents and found that the AI's writing is statistically distinct: longer sentences, more nouns and commas, fewer verbs, and a positive, neutral tone. A smart generalist might read it to see how detectable AI ghostwriting is in high-stakes political speech.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'distinct GPT style' may be an artifact of prompt under-specification: real SOTU are content-rich; GPT prompts offer no policy details, so noun-heavy abstract differences likely reflect input richness, not a stable style.","rationale":"The reader's weakest_assumption concerns training-data provenance (Section 7's claim that no SOTU was in GPT-3.5's training set). I think that assumption is not load-bearing: if SOTU texts had been memorized, the generated speeches should have been more similar to the target presidents, not less; the observed large distances would make the distinctness claim stronger, not weaker. The truly load-bearing condition is that the comparison controls for the amount and type of substantive content available to each author. Real SOTU addresses are situation-bound documents: they enumerate specific programs, name foreign leaders, cite statistics, and respond to a particular year's circumstances. The GPT prompts contained a single non-SOTU example and a year, with no such content. The paper's own data show the generated texts are one-quarter to one-sixth the length of real addresses, and Section 6 notes that 'very few city, country or proper names appear' in GPT's characteristic vocabulary and that its messages appear 'out-of-time and space.' These are precisely the expected consequences of vague prompting, not of a measurable property intrinsic to the model. The paper's convergent measures (Tables 4, 5, 7; Figure 1) are internally consistent, but they all quantify the same contrast: specific, anchored, action-oriented human speech vs. generic, abstract, nominal machine prose generated under minimal instruction. Without a condition in which GPT receives realistic briefing material, the abstract's claim that 'the GPT's style exposes distinct features' overgeneralizes from one prompt setup. This is not a question of statistical correctness but of construct validity: the independent variable 'author' is confounded with 'amount of domain information provided.' The proposed test - regenerate with matched briefing content - would settle whether the stylometric fingerprint survives when content richness is controlled. If it does, the paper's claim is robust; if not, the findings describe the prompt, not the model. This is an addressable, concrete concern, so the conditional verdict stands; no reason to reject outright. Credit is due for transparent reporting of generation parameters and convergent evidence across multiple measures.","tokens_in":13149,"tokens_out":8928,"duration_ms":76615,"concrete_test":"One decisive check: re-generate the four GPT corpora using prompts that include a realistic briefing matched to the target year - for example, 5 policy achievements, 2 foreign-policy issues, names of current cabinet members, and one recent legislative bill - while keeping temperature (0.5), top_p (0.4), and penalties at 0. Then recompute Table 4 (word length, BW, MATTR, MSL), Table 5 (POS), Table 7 (LIWC categories), and Figure 1. If the GPT speeches acquire name-rich, verb-heavy, less positive, and more specific vocabularies, closing the gap toward the real presidents, the paper's stylistic fingerprint is an artifact of prompt vagueness. If the differences persist with statistical significance, the central claim is supported and the conditional verdict can be upgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim's load-bearing condition is that the comparison isolates style, not content or prompt richness. In Section 3, the authors state: 'the used prompt does not provide a long list of information about the possible content of the target speech' and the generated speeches are ~9,000 tokens vs. ~37,000-67,000 for real presidents (Table 1). Real SOTU addresses are saturated with names, dates, policy specifics, and current events; GPT was given only one Miller Center example (explicitly not a SOTU) and a year. The stylometric differences driving the conclusion - fewer names (Table 5), less negative emotion (Table 7), more adjectives/nouns, longer words, no time-space anchoring (Table 6), and the clean separation in Figure 1 - are exactly what one expects from a model asked to produce a generic 'State of the Union' without any year-specific briefing. A human speechwriter given the same bare instruction would also produce abstract, noun-heavy, positive prose. Therefore the observed 'GPT style' may be an artifact of prompt under-specification, not a stable property of the language model. The paper's own Section 3 footnote admits uncertainty about training data, and Section 7's unsupported assertion that 'the training set never includes a SOTU address' is not needed: even if SOTU had been in training, the prompt still lacks the substantive anchors that characterize real addresses. The claim that 'even when imposing a given style, the resulting speech remains distinct' is thus conditional on the specific prompt design and does not license the abstract's general statement about 'GPT's style.'","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares State of the Union (SOTU) addresses written by Reagan, Clinton, G.W. Bush, and Obama with SOTU addresses generated by GPT-3.5-turbo under prompts that ask the model to imitate each president. Using word-frequency rankings, part-of-speech distributions, lexical statistics (word length, MATTR, mean sentence length), characteristic vocabulary, Diction/LIWC-style categories, and Labbé intertextual distance, the paper reports that GPT texts overuse the lemma \"we\", nouns, and commas, use fewer verbs and personal pronouns, contain longer words and sentences, exhibit a positive and abstract tone, and cluster separately from all real presidential speeches in a tree representation. The author concludes that GPT has a distinct, didactic, neutral style that remains distinguishable even when the model is asked to write in a target president's style.","tokens_in":13453,"tokens_out":6496,"duration_ms":57962,"significance":"If the conclusions hold, the paper provides a useful empirical demonstration that current LLM-generated political speech can be distinguished from human-authored SOTU addresses using relatively simple surface stylometric features. Strengths include the use of a genre-matched corpus (SOTU addresses), the use of established external wordlists (Diction, LIWC) rather than ad-hoc categories, transparent reporting of model sampling parameters, and a global distance visualization that makes the main separation easy to inspect. However, the central claim is currently underdetermined because the comparison does not isolate style from prompt/content richness, and the statistical significance tests ignore document-level clustering and multiple comparisons. These issues are addressable, but they are load-bearing for the paper's main conclusion.","major_comments":[{"comment":"The central comparison does not separate style from prompt/content richness. The GPT outputs are roughly 9,000 tokens while real SOTU addresses are 37,000-67,000 tokens, and the prompt explicitly contains no year-specific policy brief (Section 3). The features identified as GPT's 'style' — fewer names, fewer negative terms, no time/space anchoring, noun-heavy and abstract vocabulary (Tables 5-7) — are exactly what one would expect from any writer asked to produce a generic SOTU without substantive material. To support the claim that these are stable properties of GPT's style, the paper needs a control condition: e.g., the same prompt given to human writers, or GPT prompted with the policy content of a real SOTU, or a topic-matched comparison. Without such a control, the conclusion that 'GPT's style exposes distinct features' is confounded with input richness.","section":"§3, Table 1; §4-7"},{"comment":"The significance tests treat tokens as independent observations drawn from pooled author corpora. The proportion tests and t-tests compare aggregate percentages over all words/sentences of each group, ignoring document-level clustering; with 7-8 generated speeches and 7-8 human speeches per president, the effective sample size is much smaller than the number of tokens. In addition, dozens of categories are tested at alpha=0.01 without any multiple-comparison correction, so several asterisks are expected by chance. Please report per-document means with a mixed-effects model, paired/permutation test, or cluster-robust standard errors, and apply a multiple-testing correction or restrict confirmatory claims to pre-specified hypotheses. This affects the load-bearing statements such as 'GPT overuses the lemma we' and 'GPT employs fewer verbs'.","section":"§4, Tables 2-5, 7"},{"comment":"The row 'Mean president' cannot be reproduced from the displayed data. The text says the last two rows give the average over the 'six groups of presidential addresses', but Table 4 shows only Reagan, Clinton, Bush, and Obama; for MATTR the mean of the four displayed values is 0.325, not 0.359, and similar discrepancies occur for word length (4.425 vs 4.39), BW (28.0 vs 27.7), and MSL (20.55 vs 19.71). Either include the Trump and Biden rows in the table (or in the annexe) and state that the mean covers all six presidents, or correct the text and recompute the averages. As it stands, the headline comparison 'Mean GPT 4.90 vs Mean president 4.39' is not verifiable.","section":"§4, Table 4"},{"comment":"The sentence 'the training set never includes a SOTU address but contains another presidential speech' is a strong empirical claim about OpenAI's proprietary training data, made without citation or evidence. It is also not needed for the main demonstration: Figure 1 would still show a separation between GPT and human texts even if some SOTU addresses were in the training data, but the interpretation would change from 'style generalization' to possible memorization/retrieval. Please either remove the assertion, label it as an unverifiable assumption, or support it with documentation, and discuss the consequences for the Figure 1 interpretation if it is wrong.","section":"§7, 'Intertextual Distance'"},{"comment":"The generation procedure is not described precisely enough for replication. The paper gives temperature (0.5), top_p (0.4), and penalties (0), but does not report the full prompt text, which Miller Center speech was used as the example for each president, the date of API access, or the exact model identifier beyond 'GPT-3.5-turbo'. Because the style-imitation manipulation depends entirely on the prompt, this information should be included in an appendix. Without it, a reader cannot regenerate the corpus or assess whether the provided example biased the results.","section":"§3, footnote 3; §7 prompts"}],"minor_comments":[{"comment":"There is a typo: 'GTP' should be 'GPT'.","section":"§2"},{"comment":"The introduction's preview of the section order (Section 5 'global similarity', Section 6 'characteristic vocabulary', Section 7 'rhetorical and topical analysis') does not match the actual body, where characteristic vocabulary is Section 5, rhetorical/topical analysis is Section 6, and intertextual distance is Section 7.","section":"§1 vs §5-§7"},{"comment":"Some reference names appear misspelled: 'Vaswami et al. 2017' should be 'Vaswani et al.'; 'Zhoa et al. 2023' should be 'Zhao et al.'; 'Bartélémy' should be 'Barthélemy' if referring to the standard author.","section":"References"},{"comment":"Equation (3) uses a superscript notation (tf$_0$) that is not defined; rewrite with standard notation and state which text is the reference for length normalization.","section":"§7, Eq. (3)"},{"comment":"The footnote says the t-test is used 'with the same significance level', but it is not stated whether the unit of analysis is individual speeches or pooled tokens; this matters for interpreting the asterisks.","section":"§4, Table 4 footnote"},{"comment":"Figure 1 would benefit from a statement of how the tree was fit (e.g., Neighbor-Joining or another algorithm) and a note that branch lengths are approximate; this would help readers interpret the visual separation.","section":"§7, Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a computational linguistics or stylometry venue. The main reason for major revision, rather than rejection, is that the central confound and the statistical issues are addressable with additional experiments and re-analysis. I would ask the editor to insist on a control condition that separates prompt/content richness from model style, and on a re-analysis that respects document-level clustering and multiple comparisons. The unsupported training-data claim in Section 7 should be softened or removed. The paper's contribution is modest but potentially useful if these issues are fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful thing: the paper produces a clean, multi-measure description of what GPT-3.5 writes when asked for a State of the Union address in the style of four presidents. The finding that the model's output clusters together rather than moving toward the target author, and that it lands on a noun-heavy, comma-heavy, 'we'-heavy, abstractly positive register, is credible and worth having. The use of external wordlists (Diction/LIWC) and the intertextual distance tree makes the result easy to see.\n\nThe soft spots are three. The biggest is the prompt confound. The generated speeches are roughly 9,000 tokens against 37,000-67,000 for real presidents, and the prompts ask for a SOTU with only one Miller Center example and a year. Real SOTUs are saturated with policy specifics, names, dates, and current events. The paper's own Section 3 says the prompt does not provide a long list of possible content. Most of the measured differences—fewer names, less negative emotion, less time/space anchoring, higher abstraction—are exactly what you would expect from any writer asked to produce a generic SOTU with no briefing. A human speechwriter given the same bare instruction would likely sound similar. Without a human control using the same prompt, the central claim about 'GPT's style' is really 'GPT's style under this under-specified prompt.' That does not invalidate the descriptive result, but it does undercut the abstract's generalization.\n\nSecond, there is an internal contradiction about training data. Footnote 4 says 'we don't know precisely the training sample ... one might assume that many presidential speeches have been included,' while Section 7 asserts flatly that 'the training set never includes a SOTU address.' That assertion is unsupported and load-bearing for the style-generalization reading. Remove it.\n\nThird, the significance tests pool tokens across all speeches within a group, treating each token as independent. That ignores document-level clustering and inflates significance. It is a real but fixable problem; the asterisks are stronger than the evidence.\n\nThe paper also does not release the generated corpus or the exact prompts, which is a shame for a replication-friendly subject.\n\nBottom line: a solid applied stylometry paper, clearly written and honest in places, that needs a revision addressing the prompt confound and the training-data claim. It deserves a serious referee and a conditional acceptance. I would bring it to a reading group as a good example of LLM stylometry with a design flaw worth discussing.","headline":"A credible descriptive result about GPT-3.5's default political prose, but the prompt under-specification confound and an unsupported training-data claim keep it from being a stable 'GPT style' finding.","tokens_in":13991,"tokens_out":4048,"would_cite":true,"duration_ms":32642,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-written State of the Union speeches still read as GPT","keywords":["ChatGPT","stylometry","authorship attribution","State of the Union","large language models","political speech","intertextual distance","GPT-3.5"],"falsifier":"Check ChatGPT-3.5's training data for verbatim or near-verbatim passages from any SOTU address used in this study; if one appears, re-run the distance analysis after prompting with a speech known to be absent from training. A second approach is to repeat the experiment with a model whose training set is fully documented to include SOTUs and see whether the intertextual-distance gap shrinks or disappears.","tokens_in":12928,"feed_emoji":"🤖","tokens_out":2745,"duration_ms":23756,"temperature":0.7,"pith_summary":"The paper asks whether a large language model can pass as a presidential speechwriter, and answers no. Even when ChatGPT-3.5 is explicitly asked to mimic Reagan, Clinton, Bush, or Obama, its State of the Union addresses remain separable from the real speeches using surface stylistic measurements. The generated texts use longer words and sentences, favor nouns and commas, lean heavily on the lemma \"we\", and keep a positive, abstract tone. The paper's global evidence is an intertextual-distance tree in which all GPT outputs cluster together, away from all six presidents, including the ones being imitated. If correct, this gives stylometric tools a practical target: machine-written political rhetoric carries a detectable fingerprint.","feed_headline":"GPT speeches betray a machine style even when imitating presidents","feed_subtitle":"Stylometric tests on State of the Union addresses separate ChatGPT-3.5's output from Reagan, Clinton, Bush, and Obama.","key_machinery":"The central object is the intertextual distance measure proposed by Labbé (2007), which computes a value between 0 and 1 by comparing whole-vocabulary frequencies after normalizing text lengths. This distance is used to build a tree-based visualization of the entire corpus, and the resulting figure is the paper's global proof that GPT's style is distinct. Supporting the distance evidence are several standard stylometric tools: mean word length and percentage of big words, moving-average type-token ratio (MATTR), mean sentence length, part-of-speech distributions, characteristic-vocabulary z-scores (Muller's method), and Hart's wordlists for rhetorical categories such as Symbolism, Tenacity, and Blame.","core_discovery":"The paper's central claim is that GPT-3.5 has a distinct written style that survives the instruction to imitate a specific president. In State of the Union addresses generated for Reagan, Clinton, Bush, and Obama, GPT overuses the lemma \"we\" (about 6% of tokens versus 3.8% for real presidents), uses more nouns and adjectives, more commas, fewer verbs and adverbs, longer words (mean 4.9 letters versus 4.39) and longer sentences (mean 22.62 versus 19.71 words). Its vocabulary, measured by moving-average type-token ratio, is poorer, and its characteristic words are general, neutral, and abstract rather than tied to a specific administration's issues. The intertextual-distance tree shows all GPT-generated speeches forming a cluster separate from every president, with the distance between a target president and his GPT imitation larger than the distance between that president and other real presidents.","pith_inferences":["The finding that GPT's voice is stable across four presidential masks suggests the model's style is a property of its training distribution and prompt constraints, not of the requested persona; this could be tested by fine-tuning a model on presidential corpora and seeing whether the cluster disperses.","The neutral, positive, no-blame tone may reflect reinforcement learning from human feedback rather than a constraint of next-token prediction; a comparison between models with and without RLHF would isolate that factor.","Because the study uses only GPT-3.5, newer models with different alignment and length controls could shrink the gap; measuring the distance between human and machine speeches over successive model versions is a testable prediction."],"forward_implications":["A machine-authored political speech can be flagged with standard stylometric tools, without needing a trained classifier.","GPT's overuse of \"we\" and noun-heavy style mean generated speeches read as descriptive reports rather than calls to action.","Asking an LLM to imitate a president narrows some features but not enough; the model's own voice dominates the requested persona.","The same measurements could be extended to other LLMs and to other genres of political text, such as congressional remarks or campaign stump speeches."],"supporting_citations":[{"why":"Supplies the intertextual distance measure used to build the tree that shows GPT and presidential speeches forming separate clusters.","marker":"Labbé (2007)"},{"why":"Provides the wordlists (Symbolism, Tenacity, Blame, Achieve) and the basic stylistic measurements that the paper applies to both human and GPT speeches.","marker":"Hart (1984)"},{"why":"Justifies the decision to base a stylistic study on ubiquitous and frequent words, shaping the lemma-frequency and POS analyses.","marker":"Biber & Conrad (2009)"},{"why":"Defines the characteristic-vocabulary method (binomial model and z-scores) that the paper uses to identify words overused by GPT and by each president.","marker":"Muller (1992)"},{"why":"Supplies the interpretive framework for personal-pronoun frequencies, especially the meaning of \"we\" and its rhetorical use.","marker":"Pennebaker (2011)"},{"why":"Provides the proportion test used to determine whether frequency differences between GPT and presidents are statistically significant.","marker":"Conover (1990)"},{"why":"Offers the moving-average type-token ratio (MATTR) used to measure vocabulary richness without the usual text-length bias.","marker":"Covington & McFall (2010)"}],"fun_headline_variants":["GPT's style stays robotic even when mimicking presidents","AI presidential addresses: GPT's quirks can't be hidden","ChatGPT's 'we' tic outs it in ghostwritten speeches","Machine fingerprints visible in GPT's presidential imitations","Imitation fails: GPT's speeches cluster apart from presidents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on the claim that the model's training data never included any State of the Union address, only other presidential speeches; if one of the real speeches used for comparison was in the training data, apparent stylistic differences could be retrieval artifacts rather than a stable machine style.","fun_headline_variants_meta":{"raw":{"variants":["GPT's style stays robotic even when mimicking presidents","AI presidential addresses: GPT's quirks can't be hidden","ChatGPT's 'we' tic outs it in ghostwritten speeches","Machine fingerprints visible in GPT's presidential imitations","Imitation fails: GPT's speeches cluster apart from presidents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000705,"raw_usage":{"total_tokens":3156,"prompt_tokens":902,"completion_tokens":2254,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":2173}},"tokens_in":518,"tokens_out":2254,"duration_ms":15289,"temperature":1.0,"reasoning_tokens":2173,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:15:51.699709+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check ChatGPT-3.5's training data for verbatim or near-verbatim passages from any SOTU address used in this study; if one appears, re-run the distance analysis after prompting with a speech known to be absent from training. A second approach is to repeat the experiment with a model whose training set is fully documented to include SOTUs and see whether the intertextual-distance gap shrinks or disappears.","supporting_citations":[],"review_version":1}