A survey of LLM safety alignment, spurious correlation mitigation, and membership inference defenses, with a self-cited perspective on robust safety.
In: ICML (2024)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
A survey of LLM safety alignment, spurious correlation mitigation, and membership inference defenses, with a self-cited perspective on robust safety.