Defensive LLMs frequently intervene without identifying the right compromised trust component, and sometimes correctly diagnose a failure while still recommending no protective action.
IEEE Access , volume =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction
Defensive LLMs frequently intervene without identifying the right compromised trust component, and sometimes correctly diagnose a failure while still recommending no protective action.