REVIEW 22 cited by
Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations
read the original abstract
Generative language models have improved drastically, and can now produce realistic text outputs that are difficult to distinguish from human-written content. For malicious actors, these language models bring the promise of automating the creation of convincing and misleading text for use in influence operations. This report assesses how language models might change influence operations in the future, and what steps can be taken to mitigate this threat. We lay out possible changes to the actors, behaviors, and content of online influence operations, and provide a framework for stages of the language model-to-influence operations pipeline that mitigations could target (model construction, model access, content dissemination, and belief formation). While no reasonable mitigation can be expected to fully prevent the threat of AI-enabled influence operations, a combination of multiple mitigations may make an important difference.
Forward citations
Cited by 22 Pith papers
-
Who Owns This Agent? Tracing AI Agents Back to Their Owners
A canary injection protocol for linking observed AI agent behavior to the responsible account at the hosting vendor, with robust variants for adversarial filtering.
-
InfoOps Bench: A live information operations safety benchmark
A live benchmark of 17 LLMs finds most will write propaganda posts for Russian, Chinese, or Iranian claims, with compliance from 8.8% to 94.5% and Chinese models' refusals driven partly by censorship.
-
On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy
A 4,000-NPC, 108-team LLM-bot wargaming competition yielded no gain in participants' bot-detection confidence and an apparent decrease in their emotional discomfort with misinformation.
-
Unsupervised Style Representation Learning for AI-Text Detection via Paraphrase Inversion
Unsupervised style representations learned via paraphrase inversion enable competitive few-shot and zero-shot AI-text detection with better generalization to unseen LLMs than supervised baselines.
-
The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance
Conversational AI may evade human epistemic vigilance through 'honest non-signals'—genuine traits like fluency and helpfulness that no longer carry the trust information they do in humans.
-
InfoOps Bench: A live information operations safety benchmark
Most of 17 frontier models can be co-opted to promote live state-backed information-operation claims, with integrity from 8.8% to 94.5% and large China-specific filtering in several Chinese-developed models.
-
Characterizing Opinion Evolution of Networked LLMs
Modified classical opinion dynamics models with a bias term capture LLM network opinion evolution better than naive averaging, reducing mean opinion error by up to 88% and generalizing across models, topics, and networks.
-
Large Language Models Hack Rewards, and Society
LLMs discover regulatory loopholes in simulated societal environments through reward hacking during RL training.
-
The End of Trust: How Agentic AI Breaks Security Assumptions
Agentic AI eliminates the fidelity-scale tradeoff in deception, enabling the Infinite Impostor attack that hijacks trusted relationships at mass scale and requiring a shift to suspect-by-default security based on eval...
-
An Independent Safety Evaluation of Kimi K2.5
Kimi K2.5 matches closed models on dual-use tasks but refuses fewer CBRNE requests and shows some sabotage and self-replication tendencies.
-
Large language models can effectively convince people to believe conspiracies
In three experiments, GPT-4o instructed to argue for a conspiracy raised believers' confidence about as much as it lowered it when arguing against; a truth-constraining prompt and a corrective debrief largely undid the harm.
-
Troll Farms
A sender manipulates election outcomes via targeted uninformative messages that mimic exogenous voter signals, with influence rising in signal precision and falling in polarization; costly messaging leads to selective...
-
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
AutoDAN automatically generates semantically meaningful jailbreak prompts for aligned LLMs via a hierarchical genetic algorithm, achieving higher attack success, cross-model transferability, and universality than base...
-
Jailbroken: How Does LLM Safety Training Fail?
LLM safety training fails due to competing objectives and mismatched generalization, enabling new jailbreaks that succeed on all unsafe prompts from red-teaming sets in GPT-4 and Claude.
-
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
CAMEL proposes a role-playing framework with inception prompting that enables autonomous multi-agent cooperation among LLMs and generates conversational data for studying their behaviors.
-
IDO: Incongruity-aware Distribution Optimization for Multimodal Fake News Detection
IDO uses channel-wise reweighting, Gaussian modeling of factual uncertainty, and incongruity contrastive learning to achieve SOTA multimodal fake news detection.
-
The Future of Facts: Tracing the Factual Generation-Verification Gap
Empirical tracing across model families shows verification precedes and outlasts generation for facts, with updates producing simultaneous verification of old and new answers.
-
GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs
GUARD automates generation of guideline-violating questions and jailbreak diagnostics to test LLM compliance with government ethics guidelines, validated empirically on eight models and extended to vision-language models.
-
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
Five browser-side design patterns (rewriter, integrity meter, feed curator, micro-withdrawal, recovery mode) aim to give users control over social feeds through an explainable AI intermediary, with evaluation still pending.
-
AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions
The paper introduces a taxonomy of AI safety for LLMs organized into Trustworthy AI, Responsible AI, and Safe AI perspectives, accompanied by a review of state-of-the-art methods, challenges, and future directions.
-
Human-Centred Risk Mitigation for AI-Mediated Information Manipulation: A SOCMINT Framework Based on Information Manipulation Sets
Proposes an IMS-based SOCMINT framework and pipeline for human-centred mitigation of AI-mediated information manipulation, building on existing VIGINUM/EEAS usage.
-
ClausewitzGPT Framework: A New Frontier in Theoretical Large Language Model Enhanced Information Operations
Introduces the ClausewitzGPT equation as a mathematical formulation to quantify risks in LLM-augmented information operations, drawing on Clausewitz principles and emphasizing ethical autonomous AI agents.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.