Pith. sign in

REVIEW 1 cited by

Finding Pareto Trade-offs in Fair and Accurate Detection of Toxic Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.07661 v3 pith:T4DPJMWM submitted 2022-04-15 cs.CL

classification cs.CL
keywords fairnessaccuracytrade-offsacrossarchitectureschallengescompetingdetection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Optimizing NLP models for fairness poses many challenges. Lack of differentiable fairness measures prevents gradient-based loss training or requires surrogate losses that diverge from the true metric of interest. In addition, competing objectives (e.g., accuracy vs. fairness) often require making trade-offs based on stakeholder preferences, but stakeholders may not know their preferences before seeing system performance under different trade-off settings. To address these challenges, we begin by formulating a differentiable version of a popular fairness measure, Accuracy Parity, to provide balanced accuracy across demographic groups. Next, we show how model-agnostic, HyperNetwork optimization can efficiently train arbitrary NLP model architectures to learn Pareto-optimal trade-offs between competing metrics. Focusing on the task of toxic language detection, we show the generality and efficacy of our methods across two datasets, three neural architectures, and three fairness losses.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages

    cs.CL 2025-06 reject novelty 5.0 of 10

    Filipino language models attribute gender and sexuality bias to concrete nouns such as people and objects, diverging from the action-heavy bias drivers observed in English models.

Pith tools