Pith. sign in

REVIEW 2 cited by

Identifying Significant Predictive Bias in Classifiers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1611.08292 v2 pith:SVKAEW6O submitted 2016-11-24 stat.ML cs.LG

classification stat.MLcs.LG
keywords subgroupsbiasclassifiersubgroupdetectidentifymethodmethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a novel subset scan method to detect if a probabilistic binary classifier has statistically significant bias -- over or under predicting the risk -- for some subgroup, and identify the characteristics of this subgroup. This form of model checking and goodness-of-fit test provides a way to interpretably detect the presence of classifier bias or regions of poor classifier fit. This allows consideration of not just subgroups of a priori interest or small dimensions, but the space of all possible subgroups of features. To address the difficulty of considering these exponentially many possible subgroups, we use subset scan and parametric bootstrap-based methods. Extending this method, we can penalize the complexity of the detected subgroup and also identify subgroups with high classification errors. We demonstrate these methods and find interesting results on the COMPAS crime recidivism and credit delinquency data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SubROC: AUC-Based Discovery of Exceptional Subgroup Performance for Binary Classifiers

    cs.LG 2025-05 conditional novelty 7.0 of 10

    SubROC is a new subgroup-discovery method that finds interpretable subpopulations where a binary classifier has unusually high or low ROC/PR AUC, with provably tight search bounds.

  2. Bias Detection via Maximum Subgroup Discrepancy

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Maximum Subgroup Discrepancy is a sample-efficient, interpretable distribution distance for intersectional bias detection, provably linear in the number of protected attributes and computable to global optimality via ...

Pith tools