Pith. sign in

Cognitive Cybersecurity for Artificial Intelligence: Guardrail Engineering with CCS-7

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Language models exhibit human-like cognitive vulnerabilities, such as emotional framing, that escape traditional behavioral alignment. We present CCS-7 (Cognitive Cybersecurity Suite), a taxonomy of seven vulnerabilities grounded in human cognitive security research. To establish a human benchmark, we ran a randomized controlled trial with 151 participants: a "Think First, Verify Always" (TFVA) lesson improved cognitive security by +7.9% overall. We then evaluated TFVA-style guardrails across 12,180 experiments on seven diverse language model architectures. Results reveal architecture-dependent risk patterns: some vulnerabilities (e.g., identity confusion) are almost fully mitigated, while others (e.g., source interference) exhibit escalating backfire, with error rates increasing by up to 135% in certain models. Humans, in contrast, show consistent moderate improvement. These findings reframe cognitive safety as a model-specific engineering problem: interventions effective in one architecture may fail, or actively harm, another, underscoring the need for architecture-aware cognitive safety testing before deployment.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

background 1

representative citing papers

Lexical Hints of Accuracy in LLM Reasoning Chains

cs.CL · 2025-08-19 · reject · novelty 5.0

Hesitation words in reasoning chains are claimed to flag incorrect LLM answers, but the manuscript body is a different paper and contains no such study.

citing papers explorer

Showing 1 of 1 citing paper.

  • Lexical Hints of Accuracy in LLM Reasoning Chains cs.CL · 2025-08-19 · reject · none · ref 35 · internal anchor

    Hesitation words in reasoning chains are claimed to flag incorrect LLM answers, but the manuscript body is a different paper and contains no such study.