{"id":"2ace1c89-bc4f-43b5-a7f5-461c0d99ee88","arxiv_id":"2505.07874","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Right-wing and people-centric populist rhetoric in U.S. presidential speeches appears more emotionally charged, while all variants share an informal, assertive tone, according to a regression of model-assigned populism scores on LIWC text features.","lead":"This paper scores U.S. presidential inaugural and State of the Union addresses for four populism markers using a model trained on German parliamentary speeches, then regresses those scores on 94 stylistic word categories. It reports that populist rhetoric is informal and assertive, with right-wing and people-centric variants more emotional, but the result depends on an unvalidated domain transfer and uncorrected multiple comparisons.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central findings depend on RoBERTa populism scores that are never validated on U.S. presidential rhetoric; if those scores do not track human judgments in this genre, every regression coefficient in Tables 3–6 is uninterpretable.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing point, so agreement is 'agree.' I considered whether the more severe issue is the unadjusted multiple testing; however, that concern is secondary because it only matters if the dependent variable is meaningful. The construct-validity problem is prior: if the fine-tuned model does not measure populism in U.S. presidential addresses, then no amount of statistical correction can rescue the interpretation of the coefficients. The paper is transparent about the data sources and the fine-tuning setup, and the F1 numbers in Table 2 show the model performs reasonably on the German benchmark. But the benchmark is not the target domain. The authors do not report any check—human ratings, known-groups validity, or comparison to a dictionary-based populism measure—on U.S. presidential speeches. The absence of released code/data compounds this because the exact scoring pipeline cannot be audited. The proposed human-annotation check is feasible, modest in cost, and would directly settle whether the model's scores correspond to populist content in this genre. Until that check is run, the empirical core of the paper is unsupported, and the REJECT verdict stands.","tokens_in":14743,"tokens_out":3210,"duration_ms":32384,"concrete_test":"Have at least two coders, using the same codebook as Erhard et al. (2025), annotate a stratified sample of roughly 200 sentences drawn from the actual U.S. inaugural/SOTU corpus on left-wing, right-wing, anti-elitism, and people-centrism. Compare the RoBERTa scores to mean human ratings using rank correlation and examine whether speech-level averages separate known populist exemplars (e.g., Jackson and Trump) from known non-populist exemplars (e.g., Eisenhower and Biden). If the rank correlation is weak or the exemplar separation fails, the dependent variable lacks construct validity in this domain, and the regression results should be re-evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that a RoBERTa model fine-tuned on English translations of 8,795 German parliamentary sentences (Section 3.2.2, based on Erhard et al. 2025) produces valid populism-dimension scores when applied to 308 U.S. inaugural and State of the Union addresses spanning 1789–2025. The paper offers no evidence for this transfer: no human validation on the target corpus, no comparison with established populism measures for U.S. presidential rhetoric, and no analysis of how sentence-level model outputs are aggregated to speech-level scores. Because the training data are short German parliamentary sentences labeled by 'at least one annotator' agreement, the model may be detecting translationese, German party politics, or topic cues rather than the four populism constructs in American presidential speech. If so, the dependent variables in Equation 1 are mismeasured, and the significant coefficients in Tables 3–6—selected at p<0.05 among 94 LIWC regressors without multiple-comparison correction—cannot support the paper's conclusions about the 'sound' of populism. The limitation section does not address this construct-validity threat, and data/code are not yet released, so the scores cannot be independently checked.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper aims to characterize the \"sound\" of populism by combining LIWC-based linguistic features with a fine-tuned RoBERTa model that scores U.S. presidential inaugural and State of the Union addresses on four populism dimensions: left-wing, right-wing, anti-elitism, and people-centrism. The authors fit multiple linear regressions of each populism score on 94 LIWC features plus the speech year, then interpret the statistically significant coefficients as evidence that populist rhetoric is direct, assertive, informal but controlled, and emotionally differentiated across the four variants. The RoBERTa model is fine-tuned on English translations of 8,795 German parliamentary sentences annotated for the same four dimensions, and the resulting scores are used without further validation as dependent variables for the U.S. corpus.","tokens_in":14939,"tokens_out":3206,"duration_ms":34770,"significance":"If the populism scores were validated for American presidential rhetoric and the regression results were shown to be robust, the paper would offer a useful descriptive contribution to computational studies of populist language by linking interpretable LIWC features to transformer-based measures. The authors do provide a reproducible benchmark comparison with Erhard et al. and report a model comparison between Google Translate and GPT-4o translations, which are useful transparency elements. However, the central inference is currently unsupported because the model-derived dependent variable is never validated on the target corpus, and the predictor set is large relative to the sample size with no correction for multiple testing. As presented, the findings are best interpreted as properties of the fine-tuned model's scores rather than as features of populism in U.S. presidential speeches.","major_comments":[{"comment":"The dependent variable is a RoBERTa score produced by a model fine-tuned on English translations of 8,795 German parliamentary sentences and then applied to U.S. presidential inaugural and State of the Union addresses spanning 1789–2025. The manuscript provides no evidence that these scores measure the intended populism constructs in this target domain: there is no human validation on the U.S. corpus, no comparison with existing populism measures for American presidential rhetoric (e.g., Bonikowski et al. 2022), and no analysis of how translationese, genre differences, or historical register shift affect the model outputs. Consequently, the significant coefficients in Tables 3–6 could reflect artifacts of domain shift rather than properties of populist discourse. This construct-validity threat is not addressed in the limitations section and is load-bearing for every conclusion in the paper.","section":"§3.2.2 and §4, Eq. (1), Tables 3–6"},{"comment":"The regression model includes 94 LIWC predictors with only 308 observations, yet the paper reports only the coefficients that reach p < 0.05 and provides no multiple-comparison correction, no standard errors, and no model diagnostics. Under the null hypothesis, roughly 4.7 false positives would be expected across 94 tests, so the 27 reported significant coefficients are not interpretable without correction or holdout validation. LIWC categories are also highly intercorrelated, so multicollinearity and variance inflation should be assessed; otherwise the sign and magnitude of individual coefficients are unreliable.","section":"§4, Tables 3–6"},{"comment":"The aggregation procedure from sentence-level model predictions to speech-level populism scores is not described. The fine-tuned transformer produces a score per sentence according to Section 3.2.2, but Section 4 treats Y as a single value per speech. Without stating whether the scores are averaged, summed, or otherwise aggregated, the regression results are not reproducible. Table 8 labels the unit as \"Segment\" for a corpus of 308 speeches, which adds further ambiguity about the observational unit.","section":"§3.2.2 and §4"}],"minor_comments":[{"comment":"The row label \"Right-Ring\" appears to be a typo for \"Right-Wing.\"","section":"Table 7"},{"comment":"The text contains a placeholder citation as \"authority [?]\"; the reference is missing.","section":"§2.2"},{"comment":"The name \"UniPop\" appears in captions without being defined anywhere in the manuscript.","section":"Tables 2 and 7"},{"comment":"The caption says the features are extracted from \"U.S. presidential election speeches,\" but the corpus is described in Section 3.1 as inaugural addresses and State of the Union addresses, not election speeches.","section":"Table 8 caption"},{"comment":"The footnote states that the cross-validation step in Erhard et al. is not reproducible, but the authors do not explain how their own fine-tuning and checkpoint selection avoids the same issue; additional details on batch size and weight decay are said to be adopted without reporting exact values.","section":"Footnote 2"},{"comment":"Reference [1] and reference [19] are the same Mudde article and should be unified; several other references have inconsistent or incomplete bibliographic information.","section":"References"}],"recommendation":"reject","confidential_remarks":"The central problem is construct validity: the paper uses a model trained on translated German parliamentary sentences to score two centuries of U.S. presidential rhetoric, and no validation of those scores is offered. The regression findings are therefore uninterpretable in their current form. A revision would require external validation on the target corpus, a clear aggregation procedure, and correction for multiple comparisons; without these, the manuscript does not meet the evidentiary standard for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a well-written but methodologically fragile paper whose headline findings ride on an unvalidated transfer of a German-trained populism classifier to U.S. presidential speeches. The specific regression results are new, and the paper earns some credit for transparency and robustness checks, but the central measurement assumption does not hold up.\n\nWhat is actually new: the application of Erhard et al.'s four-dimension populism classifier to the full set of U.S. inaugural and SOTU addresses, and the resulting profile of LIWC associations (e.g., 'feel' and 'power' positively associated with right-wing populism, 'informal' positively with anti-elitism). I don't know of prior work reporting exactly these coefficients. The paper also does a clean job of describing the pipeline, and the appendix comparison of GPT-4o versus Google Translate translations is a nice touch.\n\nThe soft spot is large and central. The dependent variables are scores from a RoBERTa model fine-tuned on English translations of 8,795 short German parliamentary sentences, with annotations coded positive if at least one of five annotators agreed. Applied to U.S. presidential rhetoric spanning two centuries, there is no validation that the scores track human judgments of populism in that genre. The model could be detecting translationese, German party politics, or topic cues. The authors do not report any human-coded sample of U.S. speeches, no comparison with established populism measures, and no description of how sentence-level scores are aggregated to speech-level values. If the scores are mismeasured, every coefficient in Tables 3–6 is uninterpretable.\n\nThe regression analysis has a second, smaller problem. With 94 LIWC predictors and 308 observations, selecting only p<0.05 coefficients without multiple-comparison correction is a multiple-testing trap. No standard errors or model diagnostics are reported. This is fixable, but as is, the inferential claims are not reliable.\n\nThe good news is that the paper is clearly written, the methods are specified in enough detail to be reproduced, and the limitations section does acknowledge some gaps—though it misses the transfer-validity threat. The authors also compare German BERT and RoBERTa, which shows care about the modeling side.\n\nI would not publish this in its current form. The empirical foundation needs validation before the descriptive findings can be trusted. But the question is a good one, and the flaws are fixable. I would send it to peer review with a clear request that the authors validate the classifier's transfer to U.S. presidential rhetoric, apply multiple-comparison correction, and release code and data. A competent referee could turn this into a modest but legitimate contribution.\n\nFor your reading group, it would be a useful case study in what can go wrong when a classifier is moved across language, genre, and time period without validation. I wouldn't cite it in my own work until that validation is done.","headline":"A clearly written but methodologically fragile paper whose findings depend on an unvalidated cross-lingual transfer of a populism classifier to U.S. presidential rhetoric.","tokens_in":15485,"tokens_out":4112,"would_cite":false,"duration_ms":35887,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that populist rhetoric in U.S. presidential addresses has a distinctive, measurable linguistic 'sound' that varies systematically across left-wing, right-wing, anti-elitist, and people-centric variants.","keywords":["Populism","LIWC","RoBERTa","Presidential rhetoric","Political text analysis","Computational social science","Multi-label classification","Text as data"],"falsifier":"Re-score the 308 speeches with a populism model trained directly on English-language American political texts, or compare the RoBERTa scores against human annotation on a sample of these speeches; if the LIWC-regression coefficients vanish or reverse, the reported 'sound of populism' is an artifact of the translation-trained model rather than a property of presidential rhetoric.","tokens_in":14517,"feed_emoji":"🗣️","tokens_out":6132,"duration_ms":59145,"temperature":0.7,"pith_summary":"The paper tries to establish that populism in U.S. presidential rhetoric has a characteristic linguistic 'sound' — direct, assertive, informal but controlled — and that this sound differs systematically across four populist variants. It claims that right-wing populism and people-centrism are emotionally hot, drawing on identity, grievance, and crisis, while left-wing populism and anti-elitism are comparatively cool, emphasizing structural critique without vulgarity or hesitation. The authors care because if true, populism can be detected and tracked quantitatively in historical political speech, and the tone of populism is not a single register but a family of calibrated styles tied to ideology. The payoff is a measurable linguistic fingerprint of populist subtypes in a major democratic institution's core texts.","feed_headline":"Populist speech has a measurable 'sound' — and it splits by ideology","feed_subtitle":"Right-wing and people-centric populism run hot; left-wing and anti-elite populism run cool.","key_machinery":"The central machinery is a regression model in which four populism-dimension scores, generated by a fine-tuned RoBERTa-large model applied to each speech, are regressed on 94 LIWC linguistic-feature proportions plus the speech year. LIWC, a word-count lexicon that scores texts on psychological and stylistic categories, supplies the independent variables — informal, swear, nonflu, tentat, posemo, money, social, and others — that carry the interpretation of populism's 'sound.' The fine-tuned RoBERTa model, a context-aware transformer language model trained on English translations of German parliamentary sentences annotated for the four populism dimensions, supplies the dependent-variable scores for left-wing, right-wing, anti-elitism, and people-centrism. The regression's significant coefficients are the evidence for both the shared assertive tone and the ideological cleavages in emotional charge.","core_discovery":"The central discovery, as the paper states it, is that all four measured dimensions of populism share a core assertive 'sound': left-wing, right-wing, and anti-elitist discourse all show positive associations with informal language and negative associations with nonfluency, tentativeness, and assent, which the authors read as a deliberate projection of clarity, confidence, and closeness to 'the people.' At the same time, the variants diverge in emotional register: right-wing populism is marked by feeling and power vocabulary and low syntactic complexity, people-centrism by social and money references with low positive emotion and few questions, while left-wing populism and anti-elitism avoid swearing and netspeak and show negative associations with anger and anxiety. The paper concludes that populist rhetoric is strategically calibrated — informal and authentic, but not chaotic or vulgar — with the emotional charge concentrated in right-wing and people-centric variants.","pith_inferences":["Beyond the paper: the same LIWC-plus-transformer pipeline could be applied to campaign speeches, where the predicted emotional-charge gap between right-wing and left-wing populism should be larger, since campaign settings allow freer expression of grievance than formal addresses.","Beyond the paper: if the 'sound of populism' is a stable stylistic signature, speeches by presidents not usually labeled populist should show near-zero populism scores yet still show nonzero LIWC correlations, revealing the baseline tone of American presidential rhetoric against which populist variants stand out.","Beyond the paper: the finding that people-centrism correlates with money vocabulary while left-wing populism does not suggests a testable distinction between economic grievance framed as 'the people robbed' (people-centrism) and economic grievance framed as 'the system needs reform' (left-wing)."],"forward_implications":["Presidential addresses can be scored continuously for populist tone, making populism a traceable quantity across 236 years of U.S. political speech rather than a binary label.","The shared negative coefficients on nonfluency and tentativeness imply that populist discourse, whatever its ideology, avoids hedged or hesitant language; a speech high in hesitation markers should register as less populist.","The positive 'feel' and 'power' associations for right-wing populism imply that emotional and dominance-related vocabulary is a reliable stylistic marker of right-wing populist rhetoric.","The negative 'swear' and 'netspeak' coefficients for left-wing and anti-elitist populism imply that informality is calibrated: these variants sound informal but avoid vulgarity to preserve legitimacy.","The people-centrism results imply that social-reference and money vocabulary, together with avoidance of questions and positive emotion, are markers of people-centric populist speech."],"supporting_citations":[{"why":"Supplies the definition of populism as a thin ideology opposing 'the pure people' to 'the corrupt elite,' which frames the four dimensions being measured.","marker":"[1]"},{"why":"Provides the discursive account of populism as constructing collective identity, used to interpret the positive association between informal language and left-wing populism.","marker":"[3]"},{"why":"Introduces the LIWC lexicon and its psycholinguistic categories, which are the independent variables in the regression analysis.","marker":"[8]"},{"why":"Describes the RoBERTa model that the authors fine-tune to generate the four populism-dimension scores.","marker":"[9]"},{"why":"Supplies the German parliamentary dataset annotated for anti-elitism, people-centrism, left-wing, and right-wing populism, along with the baseline model and fine-tuning setup that the paper builds on.","marker":"[10]"},{"why":"Provides the ideational approach and the concept of authoritarian discourse style and charismatic leadership, used to interpret right-wing populism's feel and power associations.","marker":"[11]"},{"why":"Introduces the 'calibrated authenticity' theory used to explain why populist informality is accompanied by negative coefficients on nonfluency and netspeak.","marker":"[36]"},{"why":"Supports the assertiveness interpretation, including the claim that the populist leader 'claims to know the will of the people,' used for the negative interrogative and positive-emotion results.","marker":"[43]"}],"fun_headline_variants":["Populist 'sound' varies: right-wing runs hot, left-wing cool","Right-wing populism sounds more emotive than left-wing","Populism's tone differs: right is fiery, left is measured","The 'sound' of populism: emotional charge differs by variant","Study: right-wing populism uses more emotion in speech"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis depends on the assumption that a model fine-tuned on English translations of German parliamentary sentences gives trustworthy populism scores when applied to U.S. presidential inaugural and State of the Union addresses from 1789 to 2025, even though the training and target texts differ in language, genre, period, and register.","fun_headline_variants_meta":{"raw":{"variants":["Populist 'sound' varies: right-wing runs hot, left-wing cool","Right-wing populism sounds more emotive than left-wing","Populism's tone differs: right is fiery, left is measured","The 'sound' of populism: emotional charge differs by variant","Study: right-wing populism uses more emotion in speech"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1417,"prompt_tokens":929,"completion_tokens":488,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":397}},"tokens_in":545,"tokens_out":488,"duration_ms":4686,"temperature":1.0,"reasoning_tokens":397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:40:40.605947+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the 308 speeches with a populism model trained directly on English-language American political texts, or compare the RoBERTa scores against human annotation on a sample of these speeches; if the LIWC-regression coefficients vanish or reverse, the reported 'sound of populism' is an artifact of the translation-trained model rather than a property of presidential rhetoric.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition of populism as a thin ideology opposing 'the pure people' to 'the corrupt elite,' which frames the four dimensions being measured."},{"cited_title":"Verso Books, London (2005)","cited_arxiv_id":null,"evidence_quote":"Provides the discursive account of populism as constructing collective identity, used to interpret the positive association between informal language and left-wing populism."},{"cited_title":"Routledge, New York (2018)","cited_arxiv_id":null,"evidence_quote":"Provides the ideational approach and the concept of authoritarian discourse style and charismatic leadership, used to interpret right-wing populism's feel and power associations."},{"cited_title":"Stanford University Press, Stanford, CA (2016)","cited_arxiv_id":null,"evidence_quote":"Introduces the 'calibrated authenticity' theory used to explain why populist informality is accompanied by negative coefficients on nonfluency and netspeak."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the assertiveness interpretation, including the claim that the populist leader 'claims to know the will of the people,' used for the negative interrogative and positive-emotion results."}],"review_version":1}