The paper introduces TRADE, an online-only MCP attack, and RAG-Pref, a retrieval-based preference alignment method that together with DPO improves strict refusal of falsely benign MCP exploits from 6.7% to 24.1% on average.
Refusal in language models is mediated by a single direction
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
The paper introduces TRADE, an online-only MCP attack, and RAG-Pref, a retrieval-based preference alignment method that together with DPO improves strict refusal of falsely benign MCP exploits from 6.7% to 24.1% on average.