Pith. sign in

REVIEW 2 cited by

Self-supervised Learning is More Robust to Dataset Imbalance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.05025 v2 pith:NUYWIGSG submitted 2021-10-11 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords learningself-superviseddatasetsfeaturesimbalanceimbalancedlearnrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised learning (SSL) is a scalable way to learn general visual representations since it learns without labels. However, large-scale unlabeled datasets in the wild often have long-tailed label distributions, where we know little about the behavior of SSL. In this work, we systematically investigate self-supervised learning under dataset imbalance. First, we find out via extensive experiments that off-the-shelf self-supervised representations are already more robust to class imbalance than supervised representations. The performance gap between balanced and imbalanced pre-training with SSL is significantly smaller than the gap with supervised learning, across sample sizes, for both in-domain and, especially, out-of-domain evaluation. Second, towards understanding the robustness of SSL, we hypothesize that SSL learns richer features from frequent data: it may learn label-irrelevant-but-transferable features that help classify the rare classes and downstream tasks. In contrast, supervised learning has no incentive to learn features irrelevant to the labels from frequent examples. We validate this hypothesis with semi-synthetic experiments and theoretical analyses on a simplified setting. Third, inspired by the theoretical insights, we devise a re-weighted regularization technique that consistently improves the SSL representation quality on imbalanced datasets with several evaluation criteria, closing the small gap between balanced and imbalanced datasets with the same number of examples.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Object-based environment inference via kernel density estimates on learned object features achieves zero-shot room retrieval and beats scene-based CLIP.

  2. Mixture of Balanced Information Bottlenecks for Long-Tailed Visual Recognition

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A balanced information bottleneck loss, extended to a mixture over intermediate layers, improves reported accuracy on three long-tailed visual recognition benchmarks.

Pith tools