Pith. sign in

REVIEW 1 cited by

The Price of Fair PCA: One Extra Dimension

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.00103 v1 pith:FDDIPWTQ submitted 2018-10-31 cs.LG stat.ML

classification cs.LGstat.ML
keywords datafairalgorithmdifferentdimensionaldimensionalityfidelityreal-world
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We investigate whether the standard dimensionality reduction technique of PCA inadvertently produces data representations with different fidelity for two different populations. We show on several real-world data sets, PCA has higher reconstruction error on population A than on B (for example, women versus men or lower- versus higher-educated individuals). This can happen even when the data set has a similar number of samples from A and B. This motivates our study of dimensionality reduction techniques which maintain similar fidelity for A and B. We define the notion of Fair PCA and give a polynomial-time algorithm for finding a low dimensional representation of the data which is nearly-optimal with respect to this measure. Finally, we show on real-world data sets that our algorithm can be used to efficiently generate a fair low dimensional representation of the data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SOMtime the World Ain$'$t Fair: Violating Fairness Using Self-Organizing Maps

    cs.AI 2026-02 reject novelty 5.0 of 10

    High-capacity self-organizing maps recover age and income orderings from unsupervised tabular data with Spearman correlations up to 0.85, but the comparison is weakened by feature selection that uses the withheld attributes.

Pith tools