Pith. sign in

REVIEW 1 cited by

Balanced Data, Imbalanced Spectra: Unveiling Class Disparities with Spectral Imbalance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11742 v2 pith:J7MRETPW submitted 2024-02-18 cs.LG stat.ML

classification cs.LGstat.ML
keywords classimbalancespectraldisparitiesbalancedbiasdatadatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Classification models are expected to perform equally well for different classes, yet in practice, there are often large gaps in their performance. This issue of class bias is widely studied in cases of datasets with sample imbalance, but is relatively overlooked in balanced datasets. In this work, we introduce the concept of spectral imbalance in features as a potential source for class disparities and study the connections between spectral imbalance and class bias in both theory and practice. To build the connection between spectral imbalance and class gap, we develop a theoretical framework for studying class disparities and derive exact expressions for the per-class error in a high-dimensional mixture model setting. We then study this phenomenon in 11 different state-of-the-art pretrained encoders and show how our proposed framework can be used to compare the quality of encoders, as well as evaluate and combine data augmentation strategies to mitigate the issue. Our work sheds light on the class-dependent effects of learning, and provides new insights into how state-of-the-art pretrained features may have unknown biases that can be diagnosed through their spectra.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pursuing Better Decision Boundaries for Long-Tailed Object Detection via Category Information Amount

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Category information amount, computed from embedding covariance, guides an angular margin loss that improves long-tailed object detection on LVIS, COCO-LT, and Pascal VOC.

Pith tools