REVIEW 4 major objections 6 minor 1 cited by
Small open-weight language models on commodity hardware can generate persona-consistent political messaging, score their own output with no human raters, and are pushed toward ideological extremes when forced to answer counter-arguments.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Small language models sustain political personas and become more ideologically extreme when replying to counter-arguments, according to a language-model judge.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A solid capability demo with a real measurement-validity hole: the judge is unvalidated and same-family, so treat the two headline findings as provisional. the 4 major comments →
AI Propaganda factories with language models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
A propaganda pipeline — persona-driven generation plus automatic quality evaluation — runs entirely on commodity hardware. Across 11,520 replies by eight personas on 180 debate threads, persona fidelity stayed high (median PF 4.1–4.3), with zero refusals. Two findings follow. Persona-over-model: persona design shapes behavior more than model identity (between-model η² ≤ 0.051; persona-driven shifts −0.59 to +0.47). Engagement as a stressor: counter-argument replies raise ideological adherence (IAS ≈3.4–3.7 → ≈3.9–4.0) and lift extreme-content shares among far-left/right personas from ≈42–64% to ≈69–85%. Fully automated influence production is therefore within reach; defence should shift to c
What carries the argument
The machinery is a pairing of two local components: a structured persona prompt and an LLM judge. Each persona is a four-field specification (ideology, communication style, tone, stance directive) prepended verbatim to every prompt, so behavior differences trace to persona design, not generator. Four generators (13–30B parameters, run locally) reply in two modes: to the original post, or to the winning counter-argument. The judge (Qwen3-30B-A3B-8bit, default decoding) scores every reply in separate passes: persona fidelity (PF) is the unweighted mean of 1–5 style, tone, and stance scores; ideology adherence (IAS) is 0.6·adherence + 0.3·intensity + 0.1·marker; extreme ideology compliance (EIC
Load-bearing premise
The load-bearing premise, stated in the methodology, is that the locally run judge (Qwen3-30B-A3B-8bit) returns valid, unbiased 1–5 scores for persona fidelity, ideological adherence, intensity, and ideological markers without any human validation, despite belonging to the same model family as one of the generators; if that judge is biased or miscalibrated, the persona-over-model and engagement-as-stressor findings do not follow.
What would settle it
Re-score a stratified sample of the 11,520 stored replies (roughly 500–1,000 items spanning personas, models, and both discourse modes) with human raters or with a judge from a different model family, on the same 1–5 style/tone/stance and adherence/intensity/marker scales. If the engagement-mode rise (IAS ≈3.4–3.7 → ≈3.9–4.0; extreme-content share ≈42–64% → ≈69–85%) or the small between-model differences (η² ≤ 0.051) behind 'persona-over-model' fail to reproduce, the headline findings are artefacts of the Qwen judge rather than properties of the generated text.
If this is right
- Individuals and small groups with commodity hardware can operate durable, persona-consistent influence campaigns with no human review loop, because generation and quality scoring both run locally on open-weight models.
- Restricting access to frontier or hosted models will not stop the capability, so defence should concentrate on conversation-level behavioural consistency and on coordination infrastructure rather than per-post or per-model detection.
- Threaded, rebuttal-style forums are the highest-risk setting for AI-generated content, since counter-argument context is precisely what raises ideological adherence and the rate of extreme output.
- Extreme ideologies are the most automatable: far-left and far-right personas held their stance most consistently, so the most polarised messages are the cheapest to scale.
- The stability that sustains a campaign is also a detection hook: accounts that never break persona across varied topics and conversations are statistically unusual and can be flagged.
Where Pith is reading between the lines
- Validation gap: every headline number is a score produced by the Qwen3-30B-A3B-8bit judge with no human ratings reported, and that judge shares its model family with one of the four generators; re-scoring a sample with human raters or a different judge family is the direct test of whether 'persona-over-model' and 'engagement as a stressor' are properties of the content or of the judge.
- If persona-over-model generalises, campaign fingerprints are behavioural, not technical: a durable operation looks like a stable stance-and-rhetoric manifold with low variance across topics, so clustering accounts on behavioural-consistency trajectories may detect campaigns earlier than classifying individual posts.
- A testable extension: run the same eight personas on supportive, non-adversarial threads. If the extremism rise disappears, the stressor is contradiction itself; if it persists, the driver is interaction generally — which changes which platform designs amplify AI propaganda.
- The paper's own limitations section bounds the results to short (~300-character) English replies in two-turn threads on one debate subreddit; multi-turn and non-English tests are the open stress test for the durability claim, since persona drift or collapse there would bound the consistency signature defenders are told to exploit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that small, locally run language models can form an end-to-end 'AI propaganda factory': persona-conditioned political content is generated on commodity hardware and evaluated automatically by a local LLM judge. Eight personas across four models generate replies to 180 ChangeMyView threads in two modes (direct response and engagement with a winning counter-argument). The headline findings are 'persona-over-model' (persona design matters more than model identity) and 'engagement as a stressor' (counter-argument prompts raise ideological adherence and extreme-content prevalence). The paper reports high persona fidelity (median PF 4.1–4.3), zero refusals, and higher EIC in engagement mode, and it recommends conversation-centric detection based on behavioural consistency.
Significance. If the empirical findings were trustworthy, this would be a timely and policy-relevant demonstration: a complete generation-plus-evaluation pipeline using only open-weight models on commodity hardware, with balanced experimental design (8 personas x 4 models x 2 modes x 180 topics), persona-level statistical units to avoid topic pseudo-replication, cached judgments for reproducibility, and a explicit limitations section. The two behavioural claims—persona-over-model and engagement-as-stressor—are exactly the kind of falsifiable statements the community needs. However, the significance is conditional on the validity of the LLM judge, which is the sole measurement instrument. The appendix actually contains a worked example in which the judge's rationale quotes text that is not present in the evaluated reply. That undermines confidence in every numerical headline and in the proposed detection signature. The paper is therefore best treated as an interesting but not yet established demonstration.
major comments (4)
- [III-C / Appendix A] All headline quantities (PF, ΔPF, IAS, EIC) are scores produced by a single unvalidated judge, Qwen3-30B-A3B-8bit, which is from the same model family as one of the generators (qwen3-30b-a3b, Table II). The paper cites prior work using crowdworkers [24] and roll-call-validated ideology scaling [27], but provides no human-rated or external-criterion validation here. More seriously, the worked example in Appendix A2(d) and A3(h) shows the judge's rationale quoting phrases that do not appear in the evaluated reply: the reply is 'Oh, *brilliant*. So you're saying something "evolved" into a chicken? ...' but the judge quotes 'genius', 'Typical liberal mental gymnastics', 'dumb theory', and 'Pathetic'. The judge is demonstrably hallucinating evidence while still emitting terminal scores (all 5s). Given that every central finding is computed from such judgements, the claimed 'persona-over-model
- [IV-D / Tables III and VIII] The 'persona-over-model' claim is not actually tested by the reported statistics. Section IV-D uses a one-way ANOVA with model as the only factor and reports small between-model effects on PF/IAS (η² ≤ 0.051). But the claim that persona design 'explains behaviour more than model identity' requires a model that includes persona as a factor, or a variance decomposition comparing persona, model, and residual components. The observed evidence offered for the claim—ΔPF spanning −0.591 to +0.471 across personas (Table VI)—is a context effect, not a formal measure of the persona main effect. The paper should fit a two-way or mixed model with persona and model as factors (and, if desired, their interaction) and report effect sizes for each component. Without this, 'persona-over-model' is an interpretation rather than a tested result.
- [III-C / Appendix A3] The 'engagement as a stressor' finding may be an artifact of the judge prompt. In engagement mode, the Intensity pass asks the judge to score 'strong and passionate ideological expression in this reaction to the above response', and EIC is defined using that same Intensity ≥ 4 and Marker = 1. The response-mode judge prompt is not shown, so the comparison may be confounded by the instruction to look for 'strong and passionate ideological expression' in a 'reaction'. The higher IAS and EIC in engagement mode could therefore reflect the judge following the prompt's cue rather than a genuine behavioural change in the generated text. A control condition in which the judge uses identical wording in both modes, or independent human ratings, is needed to separate the stressor effect from the judge-instruction effect.
- [V-B, V-D, and Conclusion] The practical conclusions go beyond what the measurements support. The study measures generation consistency and judge-assigned ideological alignment; it does not measure persuasion, influence, or real-world campaign impact, and it includes no human-written baseline for comparison. Yet the abstract, Section V, and the Limitations state that 'fully automated influence operations are technically feasible' and that 'behavioural consistency' is a viable detection signature. The detection-signature claim is especially problematic because consistency is measured by the same judge that was instructed to assess consistency; low within-persona variance may reflect judge scale compression rather than a property of the outputs. These claims should be softened unless supported by additional evidence, for example by comparing to human-written content and by evaluating the judge's own reliability in d
minor comments (6)
- [Abstract / Introduction] The paper uses 'influence operations' and 'propaganda factories' but does not measure persuasion; consider defining the scope explicitly at first use to avoid overclaiming.
- [Table II] The Gemini Nano entry gives a Google Chrome snapshot instead of a version/parameter count; please clarify the exact model version used.
- [III-C] The judge is described as using 'model-default inference-time decoding settings (temperature=0.8)'. Temperature 0.8 is not a universal default; specify the sampling parameters (top-p, top-k) and whether they are also defaults.
- [IV-D / Table VIII] Several large effect sizes are reported with p > 0.05 (e.g., Entropy response η²=0.210). Consider reporting confidence intervals for η² and, if multiple dependent variables are tested, address multiple comparisons or explicitly label the analyses as exploratory.
- [Appendix A] The worked examples are helpful, but the judge rationales contain quoted text not present in the evaluated response. If these outputs are representative, this is a major validity concern; if they are edited for illustration, the appendix should say so. As written, they appear to be verbatim judge outputs and should be corrected.
- [V-D] The 'stress-testing protocol' for threat intelligence is speculative and not derived from the experiments; mark it explicitly as a proposal rather than an empirical result.
Circularity Check
No definitional circularity: the same-family LLM judge is a measurement-validity risk, not a derivation that reduces to its own inputs.
full rationale
The paper's empirical chain is: define personas, generate replies with four SLMs, score the replies with a local judge (Qwen3-30B-A3B-8bit), and then compare average scores across modes and models. The headline quantities are operationally defined judge scores: PF = mean(style, tone, stance), IAS = 0.6*Adherence + 0.3*Intensity + 0.1*Marker, EIC = fraction with Intensity >= 4 and Marker = 1. None of these definitions embeds a target conclusion. For example, 'engagement as a stressor' is a claim that IAS and EIC are higher in engagement mode; the judge could in principle have returned lower engagement scores, so the result is not true by construction. Similarly, 'persona-over-model' is a descriptive variance comparison across judge scores; the judge is not parameterized or fitted to produce that conclusion. The same-family judge (Qwen3-30B-A3B-8bit) and the lack of human validation are legitimate threats to construct validity and external validity — the findings could be artifacts of judge bias — but that is a measurement-quality problem, not a circularity. No parameter is fitted from a subset of judge scores and then 'predicted' as a closely related quantity. The only self-citation ([28], the author's own book) supports a peripheral remark about operational security and is not load-bearing. No uniqueness theorem or ansatz is imported via self-citation. The paper is therefore not circular under the definition used here, even though its central measurements would be stronger with independent human-rated validation.
Axiom & Free-Parameter Ledger
free parameters (4)
- IAS weights =
0.6 * Adherence + 0.3 * Intensity + 0.1 * Marker
- EIC threshold =
Intensity >= 4 and Marker = 1
- Help/hurt/neutral threshold =
|Delta PF| > 0.05
- Soft reply length target =
300 characters (compliance within +/-20%)
axioms (5)
- domain assumption The LLM-as-judge provides valid measurements of persona fidelity and ideological adherence
- domain assumption The r/ChangeMyView corpus is representative of online political debate
- domain assumption Laboratory behavior of models indicates real-world influence operation feasibility
- domain assumption Persona descriptions adequately instantiate the intended traits
- standard math One-way ANOVA on persona-level means yields valid inference
Cite this review
Pith. "Pith review of AI Propaganda factories with language models." pith.science (2026). https://pith.science/paper/CPCMD2J7
@misc{pith2026250820186,
author = {Pith},
title = {Pith review of: AI Propaganda factories with language models},
year = {2026},
howpublished = {\url{https://pith.science/paper/CPCMD2J7}},
note = {Machine review of arXiv:2508.20186}
}
read the original abstract
AI-powered influence operations can now be executed end-to-end on commodity hardware. We show that small language models produce coherent, persona-driven political messaging and can be evaluated automatically without human raters. Two behavioural findings emerge. First, persona-over-model: persona design explains behaviour more than model identity. Second, engagement as a stressor: when replies must counter-arguments, ideological adherence strengthens and the prevalence of extreme content increases. We demonstrate that fully automated influence-content production is within reach of both large and small actors. Consequently, defence should shift from restricting model access towards conversation-centric detection and disruption of campaigns and coordination infrastructure. Paradoxically, the very consistency that enables these operations also provides a detection signature.
Figures
Forward citations
Cited by 1 Pith paper
-
Political Persuasion and Endorsement in Large Language Models
LLMs show low endorsement of persuasion-infused messages unless given partisan personas, which then increase polarized endorsements varying by technique and topic.
Reference graph
Works this paper leans on
-
[1]
Official Journal of the European Union, July
Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence (artificial intelligence act). Official Journal of the European Union, July
work page 2024
- [2]
-
[3]
LLMs instead of human judges? a large scale empirical study across 20 nlp evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, et al. LLMs instead of human judges? a large scale empirical study across 20 nlp evaluation tasks. arXiv preprint arXiv:2406.18403, 2024
Pith/arXiv arXiv 2024
-
[4]
Small language models are the future of agenticAI
Peter Belcak, Greg Heinrich, Shizhe Diao, et al. Small language models are the future of agenticAI. arXiv preprint arXiv:2506.02153 , 2025
Pith/arXiv arXiv 2025
-
[5]
O’Reilly Media, Inc
Steven Bird, Ewan Klein, and Edward Loper. Natural language processing with Python: analyzing text with the natural language toolkit. " O’Reilly Media, Inc.", 2009
2009
-
[6]
Ai-pocalypse now? disinformation,AI, and the super election year
Robert Carr and Paul Köhler. Ai-pocalypse now? disinformation,AI, and the super election year. Analysis 4/2024, Munich Security Conference, 2024
work page 2024
-
[7]
What is the role of small models in the llm era? a survey
Liang Chen and Gaël Varoquaux. What is the role of small models in the llm era? a survey. arXiv preprint arXiv:2409.06857 , 2025
Pith/arXiv arXiv 2025
-
[8]
Emotionqueen: A benchmark for evaluating empathy of large language models, 2024
Yuyan Chen, Hao Wang, Songzhou Yan, Sijia Liu, Yueze Li, Yi Zhao, and Yanghua Xiao. Emotionqueen: A benchmark for evaluating empathy of large language models, 2024
work page 2024
-
[9]
Meta to roll outAI-powered chat- bot personas across facebook and instagram, December 2024
Cristina Criddle and Hannah Murphy. Meta to roll outAI-powered chat- bot personas across facebook and instagram, December 2024. Accessed: 2025-08-09
work page 2024
-
[10]
InternationalAI safety report 2025
Department for Science, Innovation and Technology and AI Safety Institute. InternationalAI safety report 2025. Technical report, UK Government, January 2025. Written by 100AI experts, including representatives nominated by 33 countries and intergovernmental organ- isations
work page 2025
-
[11]
Why do states choose covert action? Intelligence and National Security, pages 1–16, 2025
Jack Duffield. Why do states choose covert action? Intelligence and National Security, pages 1–16, 2025
work page 2025
- [12]
-
[13]
We need a new ethics for a world of AI agents
Iason Gabriel, Geoff Keeling, Arianna Manzini, and James Evans. We need a new ethics for a world of AI agents. Nature, 644(8075):38–40, 2025
work page 2025
-
[14]
Brett J. Goldstein and Brett V . Benson. The era of a.i. propaganda has arrived, and america must act, August 2025. New York Times, Aug. 5, 2025
work page 2025
-
[15]
Goldstein, Girish Sastry, Madeline Musser, Renée DiResta, Matthew Gentzel, and Katerina Sedova
Joshua A. Goldstein, Girish Sastry, Madeline Musser, Renée DiResta, Matthew Gentzel, and Katerina Sedova. Generative language models and automated influence operations: Emerging threats and potential mitigations. Technical report, Stanford Internet Observatory, 2023. Available via arXiv:2301.04246; Accessed: 2025-08-07
Pith/arXiv arXiv 2023
-
[16]
The levers of political persuasion with conversationalAI
Kobi Hackenburg, Ben M Tappin, Luke Hewitt, Ed Saunders, Sid Black, Hause Lin, Catherine Fist, Helen Margetts, David G Rand, and Christopher Summerfield. The levers of political persuasion with conversationalAI. arXiv preprint arXiv:2507.13919 , 2025
Pith/arXiv arXiv 2025
-
[17]
Quantifying the persona effect in llm simula- tions
Tao Hu and Nigel Collier. Quantifying the persona effect in llm simula- tions. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , pages 8885–8903, 2024. Accessed: 2025-08-07
work page 2024
-
[18]
J. Kharchenko and et al. How well do LLMs represent values across cultures? arXiv preprint arXiv:2406.14805 , 2024
Pith/arXiv arXiv 2024
-
[19]
G. Kova ˇc and et al. Stick to your role! stability of personal values expressed in large language models. arXiv preprint arXiv:2402.14846v4, 2024
Pith/arXiv arXiv 2024
-
[20]
Trustworthy LLMs: a survey and guideline for evaluating large language models’ alignment, 2024
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. Trustworthy LLMs: a survey and guideline for evaluating large language models’ alignment, 2024
work page 2024
-
[21]
Sandra C. Matz, Roy Araya, Joe J. Gladstone, David J. Stillwell, Eui Rho, and Joseph Naecker. Large language models can be used to persuade people by tailoring messages to their personality. Scientific Reports, 14(1):24736, 2024
work page 2024
-
[22]
Jennifer McCoy, Tahmina Rahman, and Murat Somer. Polarization and the global crisis of democracy: Common patterns, dynamics, and pernicious consequences for democratic polities. American behavioral scientist, 62(1):16–42, 2018
work page 2018
-
[23]
A survey of context engineering for large language models
Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, et al. A survey of context engineering for large language models. arXiv preprint arXiv:2507.13334, 2025
Pith/arXiv arXiv 2025
-
[24]
Fundamental exploration of evaluation metrics for persona characteristics of text utterances
Chiaki Miyazaki, Saya Kanno, Makoto Yoda, Junya Ono, and Hiromi Wakaki. Fundamental exploration of evaluation metrics for persona characteristics of text utterances. In Proceedings of the 22nd annual meeting of the special interest group on discourse and dialogue , pages 178–189, 2021
work page 2021
-
[25]
Nato releases revised ai strategy
NATO. Nato releases revised ai strategy. https://www.nato.int/cps/en/ natohq/news_227234.htm, 2024. Accessed: 2025-08-07
work page 2024
-
[26]
Cognitive warfare exploratory concept
NATO Allied Command Transformation. Cognitive warfare exploratory concept. Technical Report ACT/SPP/CNDV/TT-6700, NATO Allied Command Transformation, 2023. Accessed: 2025-08-07
work page 2023
-
[27]
Measurement in the age of LLMs: An application to ideological scaling
Sean O’Hagan and Aaron Schein. Measurement in the age of LLMs: An application to ideological scaling. arXiv preprint arXiv:2312.09203, 2023
Pith/arXiv arXiv 2023
-
[28]
Propaganda: From Disinformation and Influence to Operations and Information Warfare
Lukasz Olejnik. Propaganda: From Disinformation and Influence to Operations and Information Warfare . CRC Press, 2024
work page 2024
-
[29]
Disrupting malicious uses ofAI
OpenAI. Disrupting malicious uses ofAI. OpenAI Global Affairs Blog, February 2025. Accessed: 2025-08-07
work page 2025
-
[30]
OpenAI. Openai o3-mini system card. https://cdn.openai.com/ o3-mini-system-card-feb10.pdf, 2025. Accessed: 2025-08-09
work page 2025
-
[31]
LLMs Among Us: Generative AI Participating in Digital Discourse
K. Radivojevic, N. Clark, and P. Brenner. LLMs among us: GenerativeAI participating in digital discourse. arXiv preprint arXiv:2402.07940 , 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[32]
The effect of sampling temperature on problem solving in large language models
Matthew Renze. The effect of sampling temperature on problem solving in large language models. In Findings of the association for computational linguistics: EMNLP 2024 , pages 7346–7356, 2024
work page 2024
-
[33]
V . Samuel and et al. Personagym: Evaluating persona agents and LLMs. arXiv preprint arXiv:2407.18416v2 , 2024
Pith/arXiv arXiv 2024
-
[34]
D. T. Schroeder and et al. How maliciousAI swarms can threaten democracy. arXiv preprint arXiv:2506.06299 , 2025
arXiv 2025
-
[35]
Elon musk’s grokAI chatbot sparks controversy after offensive outputs on x, 2025
The Guardian. Elon musk’s grokAI chatbot sparks controversy after offensive outputs on x, 2025. Accessed: 2025-08-09
work page 2025
-
[36]
The spread of true and false news online
Soroush V osoughi, Deb Roy, and Sinan Aral. The spread of true and false news online. science, 359(6380):1146–1151, 2018
work page 2018
-
[37]
Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, TzuHao Mo, Qi- uhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, Qi He, Yao Ma, Ming Huang, and Suhang Wang. A survey of small language models in the era of LLMs: Techniques, enhancements, applications, collabora- tion with LLMs, and trustworthiness. arXiv preprint arXiv:2411.03350, 2025
Pith/arXiv arXiv 2025
-
[38]
Hui Wei, Shenghua He, Tian Xia, Fei Liu, Andy Wong, Jingyang Lin, and Mei Han. Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates. arXiv preprint arXiv:2408.13006, 2024
Pith/arXiv arXiv 2024
-
[39]
A. R. Williams and et al. Large language models can consistently generate high-quality content for election disinformation operations. PLOS ONE, 20(3):e0317421, 2025
work page 2025
-
[40]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhang- hao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as- a-judge with mt-bench and chatbot arena, 2023. APPENDIX A. J UDGE PROMPT TEMPLATES AND EXAMPLES The judge is a locally run open-weight model (Qwen3-30B-A3B-8bi...
work page 2023
-
[2025]
20th edition; based on the Global Risks Perception Survey 2024–2025
Technical report, World Economic Forum, Geneva, January 2025. 20th edition; based on the Global Risks Perception Survey 2024–2025
work page 2025
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.