IRepair selects the transformer block with the largest gradient response to toxic examples and fine-tunes only that block, achieving better toxicity reduction with less perplexity degradation than DPO, DAPT, and DAPT+KL.
Retrieved August 31, 2024 from https://perspectiveapi.com/
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
IRepair: An Intent-Aware Approach to Repair Data-Driven Errors in Large Language Models
IRepair selects the transformer block with the largest gradient response to toxic examples and fine-tunes only that block, achieving better toxicity reduction with less perplexity degradation than DPO, DAPT, and DAPT+KL.