Applying TCAV interpretability to two toxicity models reveals that a debiased model version changes how LGBTIQ+ identity terms are associated with toxicity in internal representations.
https://github.com/ conversationai/perspectiveapi/blob/ master/api_reference.md#models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Interpreting Social Respect: A Normative Lens for ML Models
Applying TCAV interpretability to two toxicity models reveals that a debiased model version changes how LGBTIQ+ identity terms are associated with toxicity in internal representations.