REVIEW 4 major objections 6 minor 36 references
The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read When asked to make French speech transcripts more persuasive, GPT-4o systematically shifts emotional lexicon and clause types, favoring surface style over stronger arguments.
desk verdict The protocol and feature set are genuinely useful for probing LLM persuasion, but the central 'not human-like / peripheral-route' claim is unsupported without a human baseline or outcome ratings, and the statistics are misapplied. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the paired-rewriting protocol and the hand-crafted interpretable feature set. Each of 90 transcripts from the 3MT_French balanced subset is rewritten by GPT-4o into an upgraded and a downgraded version; features approximate rhetorical devices through countable string patterns (repeated initial sounds for alliteration, repeated phrase starts for anaphora, reversed-phrase repetitions for antimetabole, first/last-word repetition for epanalepsis, connective patterns for polysyndeton and asyndeton), alongside discourse overlap, transitions, readability, syntactic diversity, negation counts, clause-type proportions, storytelling indicators, and LIWC-based affective and cognitive lexicons. Pairwise Mann-Whitney U tests between original and generated versions, plus averaged paired differences, separate features that shift regardless of persuasion direction (ChatGPT-style modifications) from features that shift differently under upgrade versus downgrade (persuasion-dependent modifications), and that separation grounds the peripheral-route interpretation.
What would settle it
Have trained annotators mark genuine rhetorical devices and argument strength in a sample of the original, upgraded, and downgraded transcripts, then check whether the string-pattern proxies match the human labels and whether human-rated persuasiveness favors the upgraded versions; if the proxies mismatch or the upgrades do not persuade human raters, the central interpretation fails.
Extended reading notes
Core claim
The paper's central claim is that GPT-4o, asked to adjust persuasiveness in speech transcripts, applies a consistent stylistic overlay rather than optimizing persuasiveness the way human speakers do. Across both upgrade and downgrade directions, the model raises lexical diversity, lowers syntactic and discourse complexity, removes anaphora, epanalepsis, antimetabole, storytelling markers and transitions, increases passive voice, and produces harder-to-read French. The persuasion-dependent shifts concentrate in emotional lexicon (upgraded versions carry more positive-emotion and affective words; downgraded versions are flatter) and in clause types (interrogative, exclamatory, and imperative rise in upgraded texts; declarative and conditional dominate downgraded ones). Read through the Elaboration Likelihood Model and the Heuristic-Systematic Model, the authors interpret this as evidence that ChatGPT's persuasive strategy works through heuristic cues and emotional tone rather than message-level argument elaboration, a route that may create short-term appeal but less stable attitude change.
Load-bearing premise
The results stand on the premise that the hand-crafted string-pattern proxies correctly identify rhetorical devices such as anaphora and storytelling in French; the paper acknowledges that these approximations were not validated by human annotation.
Editorial extensions
If this is right
- LLM-based speech-feedback tools that follow GPT-4o's pattern will tend to make texts more emotionally and stylistically vivid without strengthening the logical core of the argument.
- The consistent drop in anaphora, epanalepsis, antimetabole, and storytelling markers means automated rewriting may discourage exactly the mnemonic rhetorical devices that public-speaking experts recommend.
- The readability decrease (Flesch scores falling from about 64 to 47–48) suggests LLM rewriting can make spoken text harder to follow, even when it is judged more persuasive by surface criteria.
- The methodology is reusable for other models and languages, since it only requires paired original/modified transcripts and the released feature set.
Reading between the lines
- The suppression of repetition-based rhetorical devices may partly reflect GPT-4o's training objective of avoiding verbatim repetition rather than a deliberate persuasion strategy; comparing a repetition-tolerant model would separate these causes.
- If the peripheral-route reading is correct, upgraded speeches should persuade distracted or low-involvement listeners more than engaged ones, an audience experiment that would directly test the ELM/HSM interpretation.
- Because the 3MT_French dataset includes audio, the textual feature analysis could be extended with prosodic features to separate what the model changes in wording from what delivery would change.
- A useful sanity check is to ask GPT-4o for a 'less persuasive' version with a different phrasing and see whether the downgraded pattern (neutral, declarative, conditional) is stable or just one generic formal style.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies how GPT-4o modifies French speech transcripts from the 3MT_French corpus when prompted to make them more or less persuasive. The authors propose a protocol in which GPT-4o rewrites each of 90 balanced transcripts in two directions, then compute a hand-crafted feature set covering discourse, lexical, syntactic, and LIWC-based psychological/rhetorical properties. Their main empirical claim, stated in the abstract and in Section 4.3, is that GPT-4o applies systematic stylistic modifications rather than human-like persuasion strategies, and that its upgrades follow the peripheral-route/heuristic style described in ELM/HSM: more emotional vocabulary, more interrogative/exclamatory/imperative structures, and less rhetorical-device use (anaphora, antimetabole, storytelling) and less causal/explanatory language. The paper contributes an open-source implementation of the feature set and a comparison of feature shifts across three versions of each speech.
Significance. If the central claim were fully supported, the paper would make a useful contribution to the study of LLM-based text enhancement: it provides a reproducible pipeline, an open-source feature implementation, and one of the few analyses of persuasive text modification in a non-English (French) corpus. I do not see a circularity problem in the narrow sense, because no parameters are fitted and persuasiveness is not defined as a function of the features; the outcome is the LLM's independently generated text. However, the paper's headline conclusion is currently underdetermined by the reported experiment. The missing human baseline and the absence of human persuasiveness ratings of the generated outputs mean that the central 'not human-like' and 'peripheral-route' claims are not directly testable from the data. The open-source feature set and the balanced subset construction are concrete assets that the community can reuse, but they do not by themselves establish the theoretical interpretation.
major comments (4)
- [Section 4, first paragraph; Table 2] The Mann-Whitney U test is misapplied in the paired design. The comparisons are between the same 90 transcripts before and after modification, so the observations are dependent, and the U test assumes independent samples. The statement that 'significant U-statistics (i.e., p-value ≤ 0.05) for all groups of features' cannot therefore be trusted. I recommend a paired test (e.g., Wilcoxon signed-rank) on the 90 within-speech differences, together with effect sizes (e.g., rank-biserial or Cliff's delta) and confidence intervals, and a multiple-comparison correction across the roughly 25 features. This is load-bearing because the claim of systematic modifications rests on these significance results, and Table 2 currently gives only direction arrows without magnitudes or uncertainty.
- [Abstract and Section 4.3] The claim that GPT-4o modifies transcripts 'rather than optimizing persuasiveness in a human-like manner' is not supported by the experimental design. The study compares only original transcripts to GPT-upgraded and GPT-downgraded versions. There is no human comparison condition in which people were asked to enhance or diminish persuasiveness on the same texts, and no human judges rated the generated outputs for actual persuasiveness. Consequently, the observed feature shifts do not demonstrate that the model's behavior is unlike humans, nor that it failed to optimize persuasiveness; an equally plausible reading is that the model reproduces the surface-level linguistic cues humans use when explicitly instructed to make a text persuasive. This can be fixed only by adding a human enhancement baseline, human persuasiveness ratings of the generated outputs, or both, and the manuscript should then limit its claims to what that comparison actually shows.
- [Section 3.2 and Limitations (third paragraph)] The conclusions about rhetorical devices depend entirely on unvalidated string-pattern proxies: alliteration is approximated by repeated initial letters of adjacent words, anaphora by repeated phrase starts, epanalepsis by repeated first/last three words, and storytelling by counts of a small lexical list. The Limitations section correctly acknowledges that 'not all true rhetorical devices may have been captured, and some detected patterns may not correspond to genuine rhetorical usage,' but this is not merely a caveat: the detailed findings in Sections 4.1 and 4.2 about reduced anaphora, epanalepsis, antimetabole, and storytelling are load-bearing and would be unreliable if the proxies are poor. I recommend reporting precision/recall or correlation against a small human-annotated sample of rhetorical devices, or at least an explicit validation of the proxy definitions on the corpus.
- [Section 4.3] The interpretation through ELM, HSM, and the value-based decision-making framework is post-hoc. The paper argues that the observed directions 'align with' peripheral-route/heuristic persuasion, but several competing explanations are not ruled out: LLM instruction-following conventions, register shift from spoken to written style, or general preference optimization toward a 'polished' register. The brief rebuttal concerning spoken versus written language in Section 4.3 is not an empirical control. If the authors wish to retain the ELM/HSM interpretation, they should either test a differential prediction (e.g., audience effects such as argument scrutiny or attitude durability) or rephrase the claim as a hypothesis-generating observation with explicit alternative explanations listed.
minor comments (6)
- [Section 3.1] The generation procedure omits the exact model version/date and the sampling temperature; the results may not be reproducible without these details.
- [Section 3.2, Table 1] The acronym 'MTLD' is used without expansion; please define 'Measure of Textual Lexical Diversity' at first occurrence.
- [Table 2] The caption says bold and italic formatting indicate different types of feature shifts, but the plain-text rendering does not show this distinction clearly; please ensure the published table visually marks the categories.
- [Throughout] The paper switches between 'GPT-4o' and 'ChatGPT'; please use a single consistent term for the model.
- [Section 2.2] The sentence 'After verifying and removing corrupted samples' does not state the verification criteria; please specify how corrupted transcripts were identified.
- [References and in-text citations] Some citations are formatted inconsistently, for example 'Josh Misner, 2023' inside parenthetical citations; please unify the citation style.
Circularity Check
No significant circularity: the paper observes GPT-4o's output shifts with a literature-derived feature set; no prediction is fitted from or defined by those features.
full rationale
The paper's derivation chain is observational rather than self-referential. GPT-4o is prompted to upgrade or downgrade persuasiveness, and the authors measure the resulting linguistic shifts with a hand-crafted feature set (Section 3.2). The features are not fitted parameters, and persuasiveness is not defined as a function of those features; the upgraded and downgraded conditions are produced by the model, not by the measurement. The central claim that GPT-4o applies systematic stylistic modifications and favors peripheral-route persuasion is an interpretation of observed differences (Table 2, Section 4.3), not a quantity that reduces to the feature definitions. Self-citations (Barkar et al., 2023, 2024) are used for feature definitions and dataset annotation, but they are not invoked as the load-bearing evidence for the main conclusion. The acknowledged limitations (no human validation of rhetorical-device proxies, single-model scope) concern validity and generalizability, not circularity. Likewise, the absence of a human enhancement baseline or human persuasiveness ratings challenges the strength of the 'not human-like' claim empirically, but it does not make the derivation circular. No specific equation or fitted parameter equates the input to the output, so the paper does not exhibit any of the circularity patterns.
Assumptions & free parameters
assumptions (4)
- domain assumption Mann-Whitney U test is appropriate for comparing initial vs. generated transcripts
- ad hoc to paper Hand-crafted string patterns accurately approximate rhetorical devices
- domain assumption Human persuasiveness ratings in 3MT_French are valid ground truth
- domain assumption Observed feature shifts reflect GPT-4o's persuasion concept rather than a generic rewrite style
Cite this review
Pith. "Pith review of The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches." pith.science (2026). https://pith.science/paper/MEORB72L
@misc{pith2026250618621,
author = {Pith},
title = {Pith review of: The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches},
year = {2026},
howpublished = {\url{https://pith.science/paper/MEORB72L}},
note = {Machine review of arXiv:2506.18621}
}
read the original abstract
This study examines how large language models understand the concept of persuasiveness in public speaking by modifying speech transcripts from PhD candidates in the "Ma These en 180 Secondes" competition, using the 3MT French dataset. Our contributions include a novel methodology and an interpretable textual feature set integrating rhetorical devices and discourse markers. We prompt GPT-4o to enhance or diminish persuasiveness and analyze linguistic shifts between original and generated speech in terms of the new features. Results indicate that GPT-4o applies systematic stylistic modifications rather than optimizing persuasiveness in a human-like manner. Notably, it manipulates emotional lexicon and syntactic structures (such as interrogative and exclamatory clauses) to amplify rhetorical impact.
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
A. Barkar, M. Chollet, M. Labeau, B. Biancardi, and C. Clavel. 2024. Decoding persuasiveness in eloquence competitions: An investigation into the llm’s ability to assess public speaking. In Proceedings of the 17th International Conference on Agents and Artificial Intelligence (ICAART 2025), Porto, Portugal. (Accepted for Publication)
work page 2024
-
[4]
Alisa Barkar, Mathieu Chollet, Beatrice Biancardi, and Chloe Clavel. 2023. https://doi.org/10.1145/3610661.3617161 Insights into the importance of linguistic textual features on the persuasiveness of public speaking . In Companion Publication of the 25th International Conference on Multimodal Interaction, ICMI '23 Companion, page 51–55, New York, NY, USA....
-
[5]
Jonah Berger, Matthew D Rocklage, and Grant Packard. 2021. https://doi.org/10.1093/jcr/ucab076 Expression modalities: How speaking versus writing shapes word of mouth . Journal of Consumer Research, 49(3):389--408
-
[6]
Beatrice Biancardi, Mathieu Chollet, and Chlo \'e Clavel. 2024. Introducing the 3mt\_french dataset to investigate the timing of public speaking judgements. Language Resources and Evaluation, pages 1--20
work page 2024
-
[7]
Shelly Chaiken. 1980. https://api.semanticscholar.org/CorpusID:39212150 Heuristic versus systematic information processing and the use of source versus message cues in persuasion. Journal of Personality and Social Psychology, 39:752--766
work page 1980
-
[8]
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2023. https://api.semanticscholar.org/CorpusID:262826094 Art or artifice? large language models and the false promise of creativity . Proceedings of the CHI Conference on Human Factors in Computing Systems
work page 2023
Show all 36 references
-
[9]
Lei Chen, Chee Wee Leong, Gary Feng, Chong Min Lee, and Swapna Somasundaran. 2015. Utilizing multimodal cues to automatically evaluate public speaking performance. 2015 International Conference on Affective Computing and Intelligent Interaction (ACII), pages 394--400
2015
-
[10]
Zheng Chen and Huming Liu. 2023. https://api.semanticscholar.org/CorpusID:260706073 Stadee: Statistics-based deep detection of machine generated text . ArXiv, abs/2312.01672
2023 arXiv
-
[11]
Suchanek, and Chloé Clavel
Cyril Chhun, Fabian M. Suchanek, and Chloé Clavel. 2024. http://arxiv.org/abs/2405.13769 Do language models enjoy their own stories? prompting large language models for automatic story evaluation
2024 arXiv
-
[12]
Robert East, Kathy Hammond, and Malcolm Wright. 2007. https://doi.org/10.1016/j.ijresmar.2006.12.004 The relative incidence of positive and negative word of mouth: A multi-category study . International Journal of Research in Marketing, 24:175--184
2007 doi
-
[13]
Emily Falk and Christin Scholz. 2018. https://doi.org/10.1146/annurev-psych-122216-011821 Persuasion, influence, and value: Perspectives from communication and social neuroscience . Annual Review of Psychology, 69:329--356. Epub 2017 Sep 27
2018 doi
-
[14]
M. Forsyth. 2014. https://books.google.fr/books?id=Ui9BAwAAQBAJ The Elements of Eloquence: Secrets of the Perfect Turn of Phrase . Penguin Publishing Group
2014
-
[15]
Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel. 2024. http://arxiv.org/abs/2311.09807 The curious decline of linguistic diversity: Training language models on synthetic text
2024 arXiv
-
[16]
Guyer, Pablo Briñol, Thomas I
Joshua J. Guyer, Pablo Briñol, Thomas I. Vaughan-Johnston, Leandre R. Fabrigar, Lorena Moreno, and Richard E. Petty. 2021. https://doi.org/10.1007/s10919-021-00374-2 Paralinguistic features communicated through voice can affect appraisals of confidence and evaluative judgments...
2021 doi
-
[17]
Teck Ho and Eduardo Andrade. 2009. https://doi.org/10.1086/599221 Gaming emotions in social interactions . Journal of Consumer Research, 36:539--552
2009 doi
-
[18]
Josh Misner
Ph.D. Josh Misner. 2023. https://creativecommons.org/licenses/by/4.0/ Messages that Matter: Public Speaking in the Information Age – Third Edition , 3rd edition. North Idaho College, North Idaho College. This work is licensed under a Creative Commons Attribution 4.0 Internatio...
2023
-
[19]
Smiljana Komar. 2015. https://doi.org/10.4312/elope.12.2.29-52 Linguistic features of persuasive communication: The case of drtv short form spots . ELOPE: English Language Overseas Perspectives and Enquiries, 12:29
2015 doi
-
[20]
Sotiris Lamprinidis. 2023. https://api.semanticscholar.org/CorpusID:260125163 Llm cognitive judgements differ from human . ArXiv, abs/2307.11787
2023 arXiv
-
[21]
Lou Lee, Katarina Bartkova, Denis Jouvet, Mathilde Dargnat, and Yvon Keromnes. 2019. https://hal.archives-ouvertes.fr/hal-02177202 Can prosody meet pragmatics? case of discourse particles in french . In Proceedings of the 19th International Congress of Phonetic Sciences (ICPhS...
2019
-
[22]
Sebastian Lubos, Thi Ngoc Trang Tran, Alexander Felfernig, Seda Polat Erdeniz, and Viet-Man Le. 2024. https://api.semanticscholar.org/CorpusID:270796210 Llm-generated explanations for recommender systems . Adjunct Proceedings of the 32nd ACM Conference on User Modeling, Adapta...
2024
-
[23]
Rauterberg
Anh-Tuan Nguyen, Wei Chen, and G.W.M. Rauterberg. 2012. Online feedback system for public speakers. 2012 IEEE Symposium on E-Learning, E-Management and E-Services, pages 1--5
2012
-
[24]
Vishakh Padmakumar and He He. 2023. https://api.semanticscholar.org/CorpusID:261682154 Does writing with language models reduce content diversity? ArXiv, abs/2309.05196
2023 arXiv
-
[25]
James Pennebaker, Ryan Boyd, Roger Booth, Ashwini Ashokkumar, and Martha Francis. 2022. https://www.liwc.app Linguistic inquiry and word count: Liwc-22. pennebaker conglomerates
2022
-
[26]
Petty and John T
Richard E. Petty and John T. Cacioppo. 1986. https://doi.org/10.1007/978-1-4612-4964-1 Communication and Persuasion: Central and Peripheral Routes to Attitude Change . Springer-Verlag, New York
1986 doi
-
[27]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. http://arxiv.org/abs/2212.04356 Robust speech recognition via large-scale weak supervision
2022 arXiv
-
[28]
Gisela Redeker. 1984. https://doi.org/10.1080/01638538409544580 On differences between spoken and written language∗ . Discourse Processes, 7:43--55
1984 doi
-
[29]
Ernesto Ruelas Inzunza. 2020. https://doi.org/10.1525/abt.2020.82.8.563 Reconsidering the use of the passive voice in scientific writing . The American Biology Teacher, 82:563--565
2020 doi
-
[30]
Schourup
Lawrence C. Schourup. 1983. http://hdl.handle.net/1811/86004 Common Discourse Particles in English Conversation . Ph.D. thesis, The Ohio State University. Dissertation presented in partial fulfillment of the requirements for the Degree Doctor of Philosophy in the Graduate Scho...
1983
-
[31]
Schreiber, Gregory D
Lisa M. Schreiber, Gregory D. Paul, and Lisa R. Shibley. 2012. https://api.semanticscholar.org/CorpusID:145537101 The development and test of the public speaking competence rubric . Communication Education, 61:205 -- 233
2012
-
[32]
Manfred Stede and Schmitz Birte. 2000. https://doi.org/10.1023/A:1011112031877 Discourse particles and discourse functions . Machine Translation, 15
2000 doi
-
[33]
Ta, Ryan L
Vivian P. Ta, Ryan L. Boyd, Sarah Seraj, Anne Keller, Caroline Griffith, Alexia Loggarakis, and Lael Medema. 2022. https://doi.org/10.1007/s42001-021-00153-5 An inclusive, real-world investigation of persuasion in language and verbal behavior . Journal of Computational Social ...
2022 doi
-
[34]
Chenhao Tan, Vlad Niculae, Cristian Danescu-Niculescu-Mizil, and Lillian Lee. 2016. https://doi.org/10.1145/2872427.2883081 Winning arguments: Interaction dynamics and persuasion strategies in good-faith online discussions . In Proceedings of the 25th International Conference ...
2016
-
[35]
Bazen Gashaw Teferra and Jonathan Rose. 2023. https://doi.org/10.2196/44325 Predicting generalized anxiety disorder from impromptu speech transcripts using context-aware transformer-based neural networks: Model evaluation study . JMIR Mental Health, 10:e44325
2023 doi
-
[36]
Sowmya Vajjala. 2016. http://arxiv.org/abs/1612.00729 Automated assessment of non-native learner essays: Investigating the role of linguistic features . CoRR, abs/1612.00729
2016 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.