Pith. sign in

REVIEW 4 cited by

Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.01781 v2 pith:A2PAD6S5 submitted 2025-03-03 cs.CL

Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models

classification cs.CL
keywords modelsreasoningmodeltriggersadversarialdeepseekproblemanswer
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We investigate the robustness of reasoning models trained for step-by-step problem solving by introducing query-agnostic adversarial triggers - short, irrelevant text that, when appended to math problems, systematically mislead models to output incorrect answers without altering the problem's semantics. We propose CatAttack, an automated iterative attack pipeline for generating triggers on a weaker, less expensive proxy model (DeepSeek V3) and successfully transfer them to more advanced reasoning target models like DeepSeek R1 and DeepSeek R1-distilled-Qwen-32B, resulting in greater than 300% increase in the likelihood of the target model generating an incorrect answer. For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model getting the answer wrong. Our findings highlight critical vulnerabilities in reasoning models, revealing that even state-of-the-art models remain susceptible to subtle adversarial inputs, raising security and reliability concerns. The CatAttack triggers dataset with model responses is available at https://huggingface.co/datasets/collinear-ai/cat-attack-adversarial-triggers.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

    cs.CR 2026-07 conditional novelty 6.0

    CPInj demonstrates that federated textual prompt optimization (a TextGrad-style loop) is vulnerable to a multi-objective injection attack that persists through aggregation, degrades accuracy by up to 55 points, and ou...

  2. CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

    cs.CR 2026-07 conditional novelty 6.0

    A malicious client can poison the shared prompt of textual collaborative prompt optimization, and current prompt-injection defenses do not reliably stop it.

  3. Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models

    cs.AI 2026-05 unverdicted novelty 6.0

    An attention-guided RL reward combined with diverse persuasion strategies produces higher attack success rates against large reasoning models than prior jailbreak methods.

  4. A Neurosymbolic Approach to Natural Language Formalization and Verification

    cs.CL 2025-11 conditional novelty 6.0

    A neurosymbolic guardrail reports 99.2% soundness on a 522-item policy QA benchmark, mainly by rejecting 84% of correct answers; unvetted real-world policies score 96.8%.