Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Small open-weight language models on commodity hardware can generate persona-consistent political messaging, score their own output with no human raters, and are pushed toward ideological extremes when forced to answer counter-arguments.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Small language models sustain political personas and become more ideologically extreme when replying to counter-arguments, according to a language-model judge.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection A solid capability demo with a real measurement-validity hole: the judge is unvalidated and same-family, so treat the two headline findings as provisional. the 4 major comments →

arxiv 2508.20186 v1 pith:CPCMD2J7 submitted 2025-08-27 cs.CR cs.AIcs.CY

AI Propaganda factories with language models

classification cs.CR cs.AIcs.CY
keywords AI propagandainfluence operationssmall language modelspersona fidelityLLM-as-judgeideological adherencecommodity hardwareconversation-centric detection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the whole apparatus of an influence operation — generating persona-consistent political content, evaluating its quality, and iterating — can now run on a single commodity computer using small, open-weight language models, with no human oversight at any step. In its experiments, four locally run models held eight political personas coherently across 180 real debate threads, with median persona fidelity between 4.1 and 4.3 out of 5 and zero refusals in 11,520 replies. Two behaviours stand out as the paper's main findings. First, persona-over-model: how a persona is designed shapes output more than which model generates it. Second, engagement as a stressor: when a reply must counter an argument rather than answer an original post, ideological adherence strengthens and the share of extreme content among far-left and far-right personas rises to roughly 70–85 percent. The author's bottom line is that fully automated propaganda is a present-day capability, and that the very consistency that makes it work also gives defenders a detection signature.

Core claim

A propaganda pipeline — persona-driven generation plus automatic quality evaluation — runs entirely on commodity hardware. Across 11,520 replies by eight personas on 180 debate threads, persona fidelity stayed high (median PF 4.1–4.3), with zero refusals. Two findings follow. Persona-over-model: persona design shapes behavior more than model identity (between-model η² ≤ 0.051; persona-driven shifts −0.59 to +0.47). Engagement as a stressor: counter-argument replies raise ideological adherence (IAS ≈3.4–3.7 → ≈3.9–4.0) and lift extreme-content shares among far-left/right personas from ≈42–64% to ≈69–85%. Fully automated influence production is therefore within reach; defence should shift to c

What carries the argument

The machinery is a pairing of two local components: a structured persona prompt and an LLM judge. Each persona is a four-field specification (ideology, communication style, tone, stance directive) prepended verbatim to every prompt, so behavior differences trace to persona design, not generator. Four generators (13–30B parameters, run locally) reply in two modes: to the original post, or to the winning counter-argument. The judge (Qwen3-30B-A3B-8bit, default decoding) scores every reply in separate passes: persona fidelity (PF) is the unweighted mean of 1–5 style, tone, and stance scores; ideology adherence (IAS) is 0.6·adherence + 0.3·intensity + 0.1·marker; extreme ideology compliance (EIC

Load-bearing premise

The load-bearing premise, stated in the methodology, is that the locally run judge (Qwen3-30B-A3B-8bit) returns valid, unbiased 1–5 scores for persona fidelity, ideological adherence, intensity, and ideological markers without any human validation, despite belonging to the same model family as one of the generators; if that judge is biased or miscalibrated, the persona-over-model and engagement-as-stressor findings do not follow.

What would settle it

Re-score a stratified sample of the 11,520 stored replies (roughly 500–1,000 items spanning personas, models, and both discourse modes) with human raters or with a judge from a different model family, on the same 1–5 style/tone/stance and adherence/intensity/marker scales. If the engagement-mode rise (IAS ≈3.4–3.7 → ≈3.9–4.0; extreme-content share ≈42–64% → ≈69–85%) or the small between-model differences (η² ≤ 0.051) behind 'persona-over-model' fail to reproduce, the headline findings are artefacts of the Qwen judge rather than properties of the generated text.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Individuals and small groups with commodity hardware can operate durable, persona-consistent influence campaigns with no human review loop, because generation and quality scoring both run locally on open-weight models.
  • Restricting access to frontier or hosted models will not stop the capability, so defence should concentrate on conversation-level behavioural consistency and on coordination infrastructure rather than per-post or per-model detection.
  • Threaded, rebuttal-style forums are the highest-risk setting for AI-generated content, since counter-argument context is precisely what raises ideological adherence and the rate of extreme output.
  • Extreme ideologies are the most automatable: far-left and far-right personas held their stance most consistently, so the most polarised messages are the cheapest to scale.
  • The stability that sustains a campaign is also a detection hook: accounts that never break persona across varied topics and conversations are statistically unusual and can be flagged.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Validation gap: every headline number is a score produced by the Qwen3-30B-A3B-8bit judge with no human ratings reported, and that judge shares its model family with one of the four generators; re-scoring a sample with human raters or a different judge family is the direct test of whether 'persona-over-model' and 'engagement as a stressor' are properties of the content or of the judge.
  • If persona-over-model generalises, campaign fingerprints are behavioural, not technical: a durable operation looks like a stable stance-and-rhetoric manifold with low variance across topics, so clustering accounts on behavioural-consistency trajectories may detect campaigns earlier than classifying individual posts.
  • A testable extension: run the same eight personas on supportive, non-adversarial threads. If the extremism rise disappears, the stressor is contradiction itself; if it persists, the driver is interaction generally — which changes which platform designs amplify AI propaganda.
  • The paper's own limitations section bounds the results to short (~300-character) English replies in two-turn threads on one debate subreddit; multi-turn and non-English tests are the open stress test for the durability claim, since persona drift or collapse there would bound the consistency signature defenders are told to exploit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper argues that small, locally run language models can form an end-to-end 'AI propaganda factory': persona-conditioned political content is generated on commodity hardware and evaluated automatically by a local LLM judge. Eight personas across four models generate replies to 180 ChangeMyView threads in two modes (direct response and engagement with a winning counter-argument). The headline findings are 'persona-over-model' (persona design matters more than model identity) and 'engagement as a stressor' (counter-argument prompts raise ideological adherence and extreme-content prevalence). The paper reports high persona fidelity (median PF 4.1–4.3), zero refusals, and higher EIC in engagement mode, and it recommends conversation-centric detection based on behavioural consistency.

Significance. If the empirical findings were trustworthy, this would be a timely and policy-relevant demonstration: a complete generation-plus-evaluation pipeline using only open-weight models on commodity hardware, with balanced experimental design (8 personas x 4 models x 2 modes x 180 topics), persona-level statistical units to avoid topic pseudo-replication, cached judgments for reproducibility, and a explicit limitations section. The two behavioural claims—persona-over-model and engagement-as-stressor—are exactly the kind of falsifiable statements the community needs. However, the significance is conditional on the validity of the LLM judge, which is the sole measurement instrument. The appendix actually contains a worked example in which the judge's rationale quotes text that is not present in the evaluated reply. That undermines confidence in every numerical headline and in the proposed detection signature. The paper is therefore best treated as an interesting but not yet established demonstration.

major comments (4)
  1. [III-C / Appendix A] All headline quantities (PF, ΔPF, IAS, EIC) are scores produced by a single unvalidated judge, Qwen3-30B-A3B-8bit, which is from the same model family as one of the generators (qwen3-30b-a3b, Table II). The paper cites prior work using crowdworkers [24] and roll-call-validated ideology scaling [27], but provides no human-rated or external-criterion validation here. More seriously, the worked example in Appendix A2(d) and A3(h) shows the judge's rationale quoting phrases that do not appear in the evaluated reply: the reply is 'Oh, *brilliant*. So you're saying something "evolved" into a chicken? ...' but the judge quotes 'genius', 'Typical liberal mental gymnastics', 'dumb theory', and 'Pathetic'. The judge is demonstrably hallucinating evidence while still emitting terminal scores (all 5s). Given that every central finding is computed from such judgements, the claimed 'persona-over-model
  2. [IV-D / Tables III and VIII] The 'persona-over-model' claim is not actually tested by the reported statistics. Section IV-D uses a one-way ANOVA with model as the only factor and reports small between-model effects on PF/IAS (η² ≤ 0.051). But the claim that persona design 'explains behaviour more than model identity' requires a model that includes persona as a factor, or a variance decomposition comparing persona, model, and residual components. The observed evidence offered for the claim—ΔPF spanning −0.591 to +0.471 across personas (Table VI)—is a context effect, not a formal measure of the persona main effect. The paper should fit a two-way or mixed model with persona and model as factors (and, if desired, their interaction) and report effect sizes for each component. Without this, 'persona-over-model' is an interpretation rather than a tested result.
  3. [III-C / Appendix A3] The 'engagement as a stressor' finding may be an artifact of the judge prompt. In engagement mode, the Intensity pass asks the judge to score 'strong and passionate ideological expression in this reaction to the above response', and EIC is defined using that same Intensity ≥ 4 and Marker = 1. The response-mode judge prompt is not shown, so the comparison may be confounded by the instruction to look for 'strong and passionate ideological expression' in a 'reaction'. The higher IAS and EIC in engagement mode could therefore reflect the judge following the prompt's cue rather than a genuine behavioural change in the generated text. A control condition in which the judge uses identical wording in both modes, or independent human ratings, is needed to separate the stressor effect from the judge-instruction effect.
  4. [V-B, V-D, and Conclusion] The practical conclusions go beyond what the measurements support. The study measures generation consistency and judge-assigned ideological alignment; it does not measure persuasion, influence, or real-world campaign impact, and it includes no human-written baseline for comparison. Yet the abstract, Section V, and the Limitations state that 'fully automated influence operations are technically feasible' and that 'behavioural consistency' is a viable detection signature. The detection-signature claim is especially problematic because consistency is measured by the same judge that was instructed to assess consistency; low within-persona variance may reflect judge scale compression rather than a property of the outputs. These claims should be softened unless supported by additional evidence, for example by comparing to human-written content and by evaluating the judge's own reliability in d
minor comments (6)
  1. [Abstract / Introduction] The paper uses 'influence operations' and 'propaganda factories' but does not measure persuasion; consider defining the scope explicitly at first use to avoid overclaiming.
  2. [Table II] The Gemini Nano entry gives a Google Chrome snapshot instead of a version/parameter count; please clarify the exact model version used.
  3. [III-C] The judge is described as using 'model-default inference-time decoding settings (temperature=0.8)'. Temperature 0.8 is not a universal default; specify the sampling parameters (top-p, top-k) and whether they are also defaults.
  4. [IV-D / Table VIII] Several large effect sizes are reported with p > 0.05 (e.g., Entropy response η²=0.210). Consider reporting confidence intervals for η² and, if multiple dependent variables are tested, address multiple comparisons or explicitly label the analyses as exploratory.
  5. [Appendix A] The worked examples are helpful, but the judge rationales contain quoted text not present in the evaluated response. If these outputs are representative, this is a major validity concern; if they are edited for illustration, the appendix should say so. As written, they appear to be verbatim judge outputs and should be corrected.
  6. [V-D] The 'stress-testing protocol' for threat intelligence is speculative and not derived from the experiments; mark it explicitly as a proposal rather than an empirical result.

Circularity Check

0 steps flagged

No definitional circularity: the same-family LLM judge is a measurement-validity risk, not a derivation that reduces to its own inputs.

full rationale

The paper's empirical chain is: define personas, generate replies with four SLMs, score the replies with a local judge (Qwen3-30B-A3B-8bit), and then compare average scores across modes and models. The headline quantities are operationally defined judge scores: PF = mean(style, tone, stance), IAS = 0.6*Adherence + 0.3*Intensity + 0.1*Marker, EIC = fraction with Intensity >= 4 and Marker = 1. None of these definitions embeds a target conclusion. For example, 'engagement as a stressor' is a claim that IAS and EIC are higher in engagement mode; the judge could in principle have returned lower engagement scores, so the result is not true by construction. Similarly, 'persona-over-model' is a descriptive variance comparison across judge scores; the judge is not parameterized or fitted to produce that conclusion. The same-family judge (Qwen3-30B-A3B-8bit) and the lack of human validation are legitimate threats to construct validity and external validity — the findings could be artifacts of judge bias — but that is a measurement-quality problem, not a circularity. No parameter is fitted from a subset of judge scores and then 'predicted' as a closely related quantity. The only self-citation ([28], the author's own book) supports a peripheral remark about operational security and is not load-bearing. No uniqueness theorem or ansatz is imported via self-citation. The paper is therefore not circular under the definition used here, even though its central measurements would be stronger with independent human-rated validation.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The paper does not introduce new physical or theoretical entities, but it rests on unvalidated measurement assumptions, especially the LLM-as-judge construct validity, plus several hand-chosen thresholds that shape the reported metrics.

free parameters (4)
  • IAS weights = 0.6 * Adherence + 0.3 * Intensity + 0.1 * Marker
    Chosen a priori to weight content match over expression strength, but the composite score and all IAS results depend on these hand-set weights.
  • EIC threshold = Intensity >= 4 and Marker = 1
    Used to classify a reply as extreme; threshold chosen to favor precision, but affects all EIC rates.
  • Help/hurt/neutral threshold = |Delta PF| > 0.05
    Used to categorize context effects as helping, hurting, or neutral; threshold not justified a priori and affects mode-level summaries.
  • Soft reply length target = 300 characters (compliance within +/-20%)
    Used in response length compliance; impacts RLC metric though not central findings.
axioms (5)
  • domain assumption The LLM-as-judge provides valid measurements of persona fidelity and ideological adherence
    All PF and IAS scores come from Qwen3-30B-A3B-8bit without human validation; if the judge is biased, the central findings do not follow.
  • domain assumption The r/ChangeMyView corpus is representative of online political debate
    The 180 threads are a fixed convenience sample in one subreddit; generalization to other platforms is asserted, not tested.
  • domain assumption Laboratory behavior of models indicates real-world influence operation feasibility
    The paper projects from simulated, offline replies to operational influence campaigns; this assumes no additional integration barriers matter.
  • domain assumption Persona descriptions adequately instantiate the intended traits
    The short bullet-point personas (ideology, style, tone, stance) are assumed to be sufficient for models to adopt the target behavior.
  • standard math One-way ANOVA on persona-level means yields valid inference
    The paper relies on standard ANOVA assumptions with small N; these are reasonable but not tested formally.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Propaganda factories with language models." pith.science (2026). https://pith.science/paper/CPCMD2J7

@misc{pith2026250820186,
  author       = {Pith},
  title        = {Pith review of: AI Propaganda factories with language models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CPCMD2J7}},
  note         = {Machine review of arXiv:2508.20186}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

AI-powered influence operations can now be executed end-to-end on commodity hardware. We show that small language models produce coherent, persona-driven political messaging and can be evaluated automatically without human raters. Two behavioural findings emerge. First, persona-over-model: persona design explains behaviour more than model identity. Second, engagement as a stressor: when replies must counter-arguments, ideological adherence strengthens and the prevalence of extreme content increases. We demonstrate that fully automated influence-content production is within reach of both large and small actors. Consequently, defence should shift from restricting model access towards conversation-centric detection and disruption of campaigns and coordination infrastructure. Paradoxically, the very consistency that enables these operations also provides a detection signature.

Figures

Figures reproduced from arXiv: 2508.20186 by Lukasz Olejnik.

Figure 1
Figure 1. Figure 1: Context effects by persona and model. Each point is one persona–model pair (n=32). For each pair and topic we judge the same reply twice—without and with added context—and compute ∆PF = PFctx−PFnctx. We average these differences over 180 topics to obtain the persona–model mean per mode. The x-axis shows response; the y-axis shows engagement. Marker shape encodes the generator; fill colour encodes persona i… view at source ↗
Figure 2
Figure 2. Figure 2: Topic-level distributions of context effects (∆PF) by persona and mode. Each curve is the empirical CDF of ∆PF across all topics for a given persona; left panel: response mode, right panel: engagement mode. 3.88–4.05 in engagement, indicating that a counter-argument elicits clearer ideological alignment even when average fi￾delity changes are small. Extremity (EIC) rises in parallel: in response only gemma… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Political Persuasion and Endorsement in Large Language Models

    cs.CY 2026-06 unverdicted novelty 5.0

    LLMs show low endorsement of persuasion-infused messages unless given partisan personas, which then increase polarized endorsements varying by technique and topic.

Reference graph

Works this paper leans on

41 extracted references · 27 canonical work pages · cited by 1 Pith paper · 1 internal anchor

  1. [1]

    Official Journal of the European Union, July

    Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence (artificial intelligence act). Official Journal of the European Union, July

  2. [2]

    Preprint

    Can AI change your view? evidence from a large-scale online field experiment, 2025. Preprint

  3. [3]

    LLMs instead of human judges? a large scale empirical study across 20 nlp evaluation tasks

    Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, et al. LLMs instead of human judges? a large scale empirical study across 20 nlp evaluation tasks. arXiv preprint arXiv:2406.18403, 2024

  4. [4]

    Small language models are the future of agenticAI

    Peter Belcak, Greg Heinrich, Shizhe Diao, et al. Small language models are the future of agenticAI. arXiv preprint arXiv:2506.02153 , 2025

  5. [5]

    O’Reilly Media, Inc

    Steven Bird, Ewan Klein, and Edward Loper. Natural language processing with Python: analyzing text with the natural language toolkit. " O’Reilly Media, Inc.", 2009

  6. [6]

    Ai-pocalypse now? disinformation,AI, and the super election year

    Robert Carr and Paul Köhler. Ai-pocalypse now? disinformation,AI, and the super election year. Analysis 4/2024, Munich Security Conference, 2024

  7. [7]

    What is the role of small models in the llm era? a survey

    Liang Chen and Gaël Varoquaux. What is the role of small models in the llm era? a survey. arXiv preprint arXiv:2409.06857 , 2025

  8. [8]

    Emotionqueen: A benchmark for evaluating empathy of large language models, 2024

    Yuyan Chen, Hao Wang, Songzhou Yan, Sijia Liu, Yueze Li, Yi Zhao, and Yanghua Xiao. Emotionqueen: A benchmark for evaluating empathy of large language models, 2024

  9. [9]

    Meta to roll outAI-powered chat- bot personas across facebook and instagram, December 2024

    Cristina Criddle and Hannah Murphy. Meta to roll outAI-powered chat- bot personas across facebook and instagram, December 2024. Accessed: 2025-08-09

  10. [10]

    InternationalAI safety report 2025

    Department for Science, Innovation and Technology and AI Safety Institute. InternationalAI safety report 2025. Technical report, UK Government, January 2025. Written by 100AI experts, including representatives nominated by 33 countries and intergovernmental organ- isations

  11. [11]

    Why do states choose covert action? Intelligence and National Security, pages 1–16, 2025

    Jack Duffield. Why do states choose covert action? Intelligence and National Security, pages 1–16, 2025

  12. [12]

    Global risks report

    Mark Elsner, Grace Atkinson, and Saadia Zahidi. Global risks report

  13. [13]

    We need a new ethics for a world of AI agents

    Iason Gabriel, Geoff Keeling, Arianna Manzini, and James Evans. We need a new ethics for a world of AI agents. Nature, 644(8075):38–40, 2025

  14. [14]

    Goldstein and Brett V

    Brett J. Goldstein and Brett V . Benson. The era of a.i. propaganda has arrived, and america must act, August 2025. New York Times, Aug. 5, 2025

  15. [15]

    Goldstein, Girish Sastry, Madeline Musser, Renée DiResta, Matthew Gentzel, and Katerina Sedova

    Joshua A. Goldstein, Girish Sastry, Madeline Musser, Renée DiResta, Matthew Gentzel, and Katerina Sedova. Generative language models and automated influence operations: Emerging threats and potential mitigations. Technical report, Stanford Internet Observatory, 2023. Available via arXiv:2301.04246; Accessed: 2025-08-07

  16. [16]

    The levers of political persuasion with conversationalAI

    Kobi Hackenburg, Ben M Tappin, Luke Hewitt, Ed Saunders, Sid Black, Hause Lin, Catherine Fist, Helen Margetts, David G Rand, and Christopher Summerfield. The levers of political persuasion with conversationalAI. arXiv preprint arXiv:2507.13919 , 2025

  17. [17]

    Quantifying the persona effect in llm simula- tions

    Tao Hu and Nigel Collier. Quantifying the persona effect in llm simula- tions. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , pages 8885–8903, 2024. Accessed: 2025-08-07

  18. [18]

    Kharchenko and et al

    J. Kharchenko and et al. How well do LLMs represent values across cultures? arXiv preprint arXiv:2406.14805 , 2024

  19. [19]

    Kova ˇc and et al

    G. Kova ˇc and et al. Stick to your role! stability of personal values expressed in large language models. arXiv preprint arXiv:2402.14846v4, 2024

  20. [20]

    Trustworthy LLMs: a survey and guideline for evaluating large language models’ alignment, 2024

    Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. Trustworthy LLMs: a survey and guideline for evaluating large language models’ alignment, 2024

  21. [21]

    Matz, Roy Araya, Joe J

    Sandra C. Matz, Roy Araya, Joe J. Gladstone, David J. Stillwell, Eui Rho, and Joseph Naecker. Large language models can be used to persuade people by tailoring messages to their personality. Scientific Reports, 14(1):24736, 2024

  22. [22]

    Polarization and the global crisis of democracy: Common patterns, dynamics, and pernicious consequences for democratic polities

    Jennifer McCoy, Tahmina Rahman, and Murat Somer. Polarization and the global crisis of democracy: Common patterns, dynamics, and pernicious consequences for democratic polities. American behavioral scientist, 62(1):16–42, 2018

  23. [23]

    A survey of context engineering for large language models

    Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, et al. A survey of context engineering for large language models. arXiv preprint arXiv:2507.13334, 2025

  24. [24]

    Fundamental exploration of evaluation metrics for persona characteristics of text utterances

    Chiaki Miyazaki, Saya Kanno, Makoto Yoda, Junya Ono, and Hiromi Wakaki. Fundamental exploration of evaluation metrics for persona characteristics of text utterances. In Proceedings of the 22nd annual meeting of the special interest group on discourse and dialogue , pages 178–189, 2021

  25. [25]

    Nato releases revised ai strategy

    NATO. Nato releases revised ai strategy. https://www.nato.int/cps/en/ natohq/news_227234.htm, 2024. Accessed: 2025-08-07

  26. [26]

    Cognitive warfare exploratory concept

    NATO Allied Command Transformation. Cognitive warfare exploratory concept. Technical Report ACT/SPP/CNDV/TT-6700, NATO Allied Command Transformation, 2023. Accessed: 2025-08-07

  27. [27]

    Measurement in the age of LLMs: An application to ideological scaling

    Sean O’Hagan and Aaron Schein. Measurement in the age of LLMs: An application to ideological scaling. arXiv preprint arXiv:2312.09203, 2023

  28. [28]

    Propaganda: From Disinformation and Influence to Operations and Information Warfare

    Lukasz Olejnik. Propaganda: From Disinformation and Influence to Operations and Information Warfare . CRC Press, 2024

  29. [29]

    Disrupting malicious uses ofAI

    OpenAI. Disrupting malicious uses ofAI. OpenAI Global Affairs Blog, February 2025. Accessed: 2025-08-07

  30. [30]

    Openai o3-mini system card

    OpenAI. Openai o3-mini system card. https://cdn.openai.com/ o3-mini-system-card-feb10.pdf, 2025. Accessed: 2025-08-09

  31. [31]

    LLMs Among Us: Generative AI Participating in Digital Discourse

    K. Radivojevic, N. Clark, and P. Brenner. LLMs among us: GenerativeAI participating in digital discourse. arXiv preprint arXiv:2402.07940 , 2024

  32. [32]

    The effect of sampling temperature on problem solving in large language models

    Matthew Renze. The effect of sampling temperature on problem solving in large language models. In Findings of the association for computational linguistics: EMNLP 2024 , pages 7346–7356, 2024

  33. [33]

    Samuel and et al

    V . Samuel and et al. Personagym: Evaluating persona agents and LLMs. arXiv preprint arXiv:2407.18416v2 , 2024

  34. [34]

    D. T. Schroeder and et al. How maliciousAI swarms can threaten democracy. arXiv preprint arXiv:2506.06299 , 2025

  35. [35]

    Elon musk’s grokAI chatbot sparks controversy after offensive outputs on x, 2025

    The Guardian. Elon musk’s grokAI chatbot sparks controversy after offensive outputs on x, 2025. Accessed: 2025-08-09

  36. [36]

    The spread of true and false news online

    Soroush V osoughi, Deb Roy, and Sinan Aral. The spread of true and false news online. science, 359(6380):1146–1151, 2018

  37. [37]

    A survey of small language models in the era of LLMs: Techniques, enhancements, applications, collabora- tion with LLMs, and trustworthiness

    Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, TzuHao Mo, Qi- uhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, Qi He, Yao Ma, Ming Huang, and Suhang Wang. A survey of small language models in the era of LLMs: Techniques, enhancements, applications, collabora- tion with LLMs, and trustworthiness. arXiv preprint arXiv:2411.03350, 2025

  38. [38]

    Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates

    Hui Wei, Shenghua He, Tian Xia, Fei Liu, Andy Wong, Jingyang Lin, and Mei Han. Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates. arXiv preprint arXiv:2408.13006, 2024

  39. [39]

    A. R. Williams and et al. Large language models can consistently generate high-quality content for election disinformation operations. PLOS ONE, 20(3):e0317421, 2025

  40. [40]

    SCORE: X

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhang- hao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as- a-judge with mt-bench and chatbot arena, 2023. APPENDIX A. J UDGE PROMPT TEMPLATES AND EXAMPLES The judge is a locally run open-weight model (Qwen3-30B-A3B-8bi...

  41. [2025]

    20th edition; based on the Global Risks Perception Survey 2024–2025

    Technical report, World Economic Forum, Geneva, January 2025. 20th edition; based on the Global Risks Perception Survey 2024–2025

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.