Pith. sign in

REVIEW 5 cited by

CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.01900 v1 pith:M5QGPPHJ submitted 2025-05-03 cs.CL

classification cs.CL
keywords systemsadversarialagentcamouflagedetectionattackattackerattacks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automated evidence-based misinformation detection systems, which evaluate the veracity of short claims against evidence, lack comprehensive analysis of their adversarial vulnerabilities. Existing black-box text-based adversarial attacks are ill-suited for evidence-based misinformation detection systems, as these attacks primarily focus on token-level substitutions involving gradient or logit-based optimization strategies, which are incapable of fooling the multi-component nature of these detection systems. These systems incorporate both retrieval and claim-evidence comparison modules, which requires attacks to break the retrieval of evidence and/or the comparison module so that it draws incorrect inferences. We present CAMOUFLAGE, an iterative, LLM-driven approach that employs a two-agent system, a Prompt Optimization Agent and an Attacker Agent, to create adversarial claim rewritings that manipulate evidence retrieval and mislead claim-evidence comparison, effectively bypassing the system without altering the meaning of the claim. The Attacker Agent produces semantically equivalent rewrites that attempt to mislead detectors, while the Prompt Optimization Agent analyzes failed attack attempts and refines the prompt of the Attacker to guide subsequent rewrites. This enables larger structural and stylistic transformations of the text rather than token-level substitutions, adapting the magnitude of changes based on previous outcomes. Unlike existing approaches, CAMOUFLAGE optimizes its attack solely based on binary model decisions to guide its rewriting process, eliminating the need for classifier logits or extensive querying. We evaluate CAMOUFLAGE on four systems, including two recent academic systems and two real-world APIs, with an average attack success rate of 46.92\% while preserving textual coherence and semantic equivalence to the original claims.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AtomEval: Validity-Aware Atomic Evaluation of Adversarial Claim Rewriting in Fact Verification

    cs.CL 2026-04 conditional novelty 6.5 of 10

    On FEVER refuted claims, conventional ASR often counts proposition-changing rewrites as attacks; AtomEval’s SROM-based VASR separates valid verifier evasion from invalid rewrites.

  2. AtomEval: Validity-Aware Atomic Evaluation of Adversarial Claim Rewriting in Fact Verification

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    AtomEval introduces atomic claim decomposition and validity scoring to provide more reliable evaluation of adversarial rewrites than standard similarity metrics in fact verification.

  3. Large Language Models in Misinformation Ecosystems: Misuse, Defense, and Vulnerability

    cs.CR 2026-07 conditional novelty 5.0 of 10

    A role-layer survey unifies LLM misuse, LLM-based defense, and LLM-centric verification vulnerabilities across content, social, evidence, and workflow layers, then lists three open challenges.

  4. Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threats

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    Multi-agent LLM system with consensus and hierarchy outperforms individual models on disinformation detection tasks across English, Polish, Slovak, and Bulgarian datasets.

  5. Prompt Governance? On Governing Technologies Governed by Natural Language

    cs.CY 2026-04 unverdicted novelty 4.0 of 10

    Literature on system prompts for AI shows fragmented and contradictory claims that complicate policy efforts to use them as reliable governance mechanisms.

Pith tools