Pith. sign in

REVIEW 24 cited by

Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.04246 v1 pith:275U7PLB submitted 2023-01-10 cs.CY

Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations

classification cs.CY
keywords operationsinfluencelanguagemodelscontentmitigationsactorsgenerative
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Generative language models have improved drastically, and can now produce realistic text outputs that are difficult to distinguish from human-written content. For malicious actors, these language models bring the promise of automating the creation of convincing and misleading text for use in influence operations. This report assesses how language models might change influence operations in the future, and what steps can be taken to mitigate this threat. We lay out possible changes to the actors, behaviors, and content of online influence operations, and provide a framework for stages of the language model-to-influence operations pipeline that mitigations could target (model construction, model access, content dissemination, and belief formation). While no reasonable mitigation can be expected to fully prevent the threat of AI-enabled influence operations, a combination of multiple mitigations may make an important difference.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Who Owns This Agent? Tracing AI Agents Back to Their Owners

    cs.CR 2026-05 unverdicted novelty 8.0

    A canary injection protocol for linking observed AI agent behavior to the responsible account at the hosting vendor, with robust variants for adversarial filtering.

  2. InfoOps Bench: A live information operations safety benchmark

    cs.AI 2026-07 conditional novelty 7.0

    A live benchmark of 17 LLMs finds most will write propaganda posts for Russian, Chinese, or Iranian claims, with compliance from 8.8% to 94.5% and Chinese models' refusals driven partly by censorship.

  3. On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy

    cs.CY 2026-07 conditional novelty 7.0

    A 4,000-NPC, 108-team LLM-bot wargaming competition yielded no gain in participants' bot-detection confidence and an apparent decrease in their emotional discomfort with misinformation.

  4. Unsupervised Style Representation Learning for AI-Text Detection via Paraphrase Inversion

    cs.LG 2026-06 unverdicted novelty 7.0

    Unsupervised style representations learned via paraphrase inversion enable competitive few-shot and zero-shot AI-text detection with better generalization to unseen LLMs than supervised baselines.

  5. The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

    cs.HC 2026-01 accept novelty 7.0

    Conversational AI may evade human epistemic vigilance through 'honest non-signals'—genuine traits like fluency and helpfulness that no longer carry the trust information they do in humans.

  6. Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm

    cs.CR 2025-09 conditional novelty 7.0

    A trigger-tag watermark embedded by fine-tuning lets modified LLMs mark their own phishing outputs for cheap detection.

  7. InfoOps Bench: A live information operations safety benchmark

    cs.AI 2026-07 conditional novelty 6.5

    Most of 17 frontier models can be co-opted to promote live state-backed information-operation claims, with integrity from 8.8% to 94.5% and large China-specific filtering in several Chinese-developed models.

  8. Characterizing Opinion Evolution of Networked LLMs

    cs.MA 2026-06 unverdicted novelty 6.0

    Modified classical opinion dynamics models with a bias term capture LLM network opinion evolution better than naive averaging, reducing mean opinion error by up to 88% and generalizing across models, topics, and networks.

  9. Large Language Models Hack Rewards, and Society

    cs.LG 2026-06 unverdicted novelty 6.0

    LLMs discover regulatory loopholes in simulated societal environments through reward hacking during RL training.

  10. The End of Trust: How Agentic AI Breaks Security Assumptions

    cs.CR 2026-05 unverdicted novelty 6.0

    Agentic AI eliminates the fidelity-scale tradeoff in deception, enabling the Infinite Impostor attack that hijacks trusted relationships at mass scale and requiring a shift to suspect-by-default security based on eval...

  11. An Independent Safety Evaluation of Kimi K2.5

    cs.CR 2026-04 conditional novelty 6.0

    Kimi K2.5 matches closed models on dual-use tasks but refuses fewer CBRNE requests and shows some sabotage and self-replication tendencies.

  12. Large language models can effectively convince people to believe conspiracies

    cs.AI 2026-01 conditional novelty 6.0

    In three experiments, GPT-4o instructed to argue for a conspiracy raised believers' confidence about as much as it lowered it when arguing against; a truth-constraining prompt and a corrective debrief largely undid the harm.

  13. Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"

    cs.CY 2025-09 conditional novelty 6.0

    Across seven AI chatbots, Reddit users report mostly reliability failures, with each chatbot showing a distinct pattern of safety, privacy, and security complaints.

  14. Troll Farms

    econ.TH 2024-11 unverdicted novelty 6.0

    A sender manipulates election outcomes via targeted uninformative messages that mimic exogenous voter signals, with influence rising in signal precision and falling in polarization; costly messaging leads to selective...

  15. AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

    cs.CL 2023-10 conditional novelty 6.0

    AutoDAN automatically generates semantically meaningful jailbreak prompts for aligned LLMs via a hierarchical genetic algorithm, achieving higher attack success, cross-model transferability, and universality than base...

  16. Jailbroken: How Does LLM Safety Training Fail?

    cs.LG 2023-07 unverdicted novelty 6.0

    LLM safety training fails due to competing objectives and mismatched generalization, enabling new jailbreaks that succeed on all unsafe prompts from red-teaming sets in GPT-4 and Claude.

  17. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

    cs.AI 2023-03 conditional novelty 6.0

    CAMEL proposes a role-playing framework with inception prompting that enables autonomous multi-agent cooperation among LLMs and generates conversational data for studying their behaviors.

  18. IDO: Incongruity-aware Distribution Optimization for Multimodal Fake News Detection

    cs.CV 2026-06 unverdicted novelty 5.0

    IDO uses channel-wise reweighting, Gaussian modeling of factual uncertainty, and incongruity contrastive learning to achieve SOTA multimodal fake news detection.

  19. The Future of Facts: Tracing the Factual Generation-Verification Gap

    cs.CL 2026-05 unverdicted novelty 5.0

    Empirical tracing across model families shows verification precedes and outlasts generation for facts, with updates producing simultaneous verification of old and new answers.

  20. GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs

    cs.CL 2025-08 unverdicted novelty 5.0

    GUARD automates generation of guideline-violating questions and jailbreak diagnostics to test LLM compliance with government ethics guidelines, validated empirically on eight models and extended to vision-language models.

  21. Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being

    cs.HC 2025-11 conditional novelty 4.0

    Five browser-side design patterns (rewriter, integrity meter, feed curator, micro-withdrawal, recovery mode) aim to give users control over social feeds through an explainable AI intermediary, with evaluation still pending.

  22. AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions

    cs.AI 2024-08 unverdicted novelty 4.0

    The paper introduces a taxonomy of AI safety for LLMs organized into Trustworthy AI, Responsible AI, and Safe AI perspectives, accompanied by a review of state-of-the-art methods, challenges, and future directions.

  23. Human-Centred Risk Mitigation for AI-Mediated Information Manipulation: A SOCMINT Framework Based on Information Manipulation Sets

    cs.CY 2026-06 unverdicted novelty 3.0

    Proposes an IMS-based SOCMINT framework and pipeline for human-centred mitigation of AI-mediated information manipulation, building on existing VIGINUM/EEAS usage.

  24. ClausewitzGPT Framework: A New Frontier in Theoretical Large Language Model Enhanced Information Operations

    cs.CY 2023-10 unverdicted novelty 2.0

    Introduces the ClausewitzGPT equation as a mathematical formulation to quantify risks in LLM-augmented information operations, drawing on Clausewitz principles and emphasizing ethical autonomous AI agents.