Pith. sign in

REVIEW 2 cited by

Explainability for fair machine learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.07389 v1 pith:PJ3IBUER submitted 2020-10-14 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords modelexplainabilityfairnesslearningmachineunfairnessalgorithmeven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the decisions made or influenced by machine learning models increasingly impact our lives, it is crucial to detect, understand, and mitigate unfairness. But even simply determining what "unfairness" should mean in a given context is non-trivial: there are many competing definitions, and choosing between them often requires a deep understanding of the underlying task. It is thus tempting to use model explainability to gain insights into model fairness, however existing explainability tools do not reliably indicate whether a model is indeed fair. In this work we present a new approach to explaining fairness in machine learning, based on the Shapley value paradigm. Our fairness explanations attribute a model's overall unfairness to individual input features, even in cases where the model does not operate on sensitive attributes directly. Moreover, motivated by the linearity of Shapley explainability, we propose a meta algorithm for applying existing training-time fairness interventions, wherein one trains a perturbation to the original model, rather than a new model entirely. By explaining the original model, the perturbation, and the fair-corrected model, we gain insight into the accuracy-fairness trade-off that is being made by the intervention. We further show that this meta algorithm enjoys both flexibility and stability benefits with no loss in performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Using concept vectors as bias representations, the framework semi-automatically labels and localizes spurious artifacts and partially unlearns them, though success varies by architecture and artifact type.

  2. Constructing Fair Latent Space for Intersection of Fairness and Explainability

    cs.LG 2024-12 conditional novelty 5.0 of 10

    An invertible module attached to a frozen pretrained generative model disentangles labels from sensitive attributes in latent space, improving fairness metrics and enabling counterfactual explanations.

Pith tools