REVIEW 4 cited by
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
read the original abstract
We investigate the robustness of reasoning models trained for step-by-step problem solving by introducing query-agnostic adversarial triggers - short, irrelevant text that, when appended to math problems, systematically mislead models to output incorrect answers without altering the problem's semantics. We propose CatAttack, an automated iterative attack pipeline for generating triggers on a weaker, less expensive proxy model (DeepSeek V3) and successfully transfer them to more advanced reasoning target models like DeepSeek R1 and DeepSeek R1-distilled-Qwen-32B, resulting in greater than 300% increase in the likelihood of the target model generating an incorrect answer. For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model getting the answer wrong. Our findings highlight critical vulnerabilities in reasoning models, revealing that even state-of-the-art models remain susceptible to subtle adversarial inputs, raising security and reliability concerns. The CatAttack triggers dataset with model responses is available at https://huggingface.co/datasets/collinear-ai/cat-attack-adversarial-triggers.
Forward citations
Cited by 4 Pith papers
-
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization
CPInj demonstrates that federated textual prompt optimization (a TextGrad-style loop) is vulnerable to a multi-objective injection attack that persists through aggregation, degrades accuracy by up to 55 points, and ou...
-
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization
A malicious client can poison the shared prompt of textual collaborative prompt optimization, and current prompt-injection defenses do not reliably stop it.
-
Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models
An attention-guided RL reward combined with diverse persuasion strategies produces higher attack success rates against large reasoning models than prior jailbreak methods.
-
A Neurosymbolic Approach to Natural Language Formalization and Verification
A neurosymbolic guardrail reports 99.2% soundness on a 522-item policy QA benchmark, mainly by rejecting 84% of correct answers; unvetted real-world policies score 96.8%.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.