Pith. sign in

REVIEW 2 cited by

When does Bias Transfer in Transfer Learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.02842 v1 pith:6JU2JQE5 submitted 2022-07-06 cs.LG

classification cs.LG
keywords transferbiasmodelsourcetargetwhendownsideeven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Using transfer learning to adapt a pre-trained "source model" to a downstream "target task" can dramatically increase performance with seemingly no downside. In this work, we demonstrate that there can exist a downside after all: bias transfer, or the tendency for biases of the source model to persist even after adapting the model to the target class. Through a combination of synthetic and natural experiments, we show that bias transfer both (a) arises in realistic settings (such as when pre-training on ImageNet or other standard datasets) and (b) can occur even when the target dataset is explicitly de-biased. As transfer-learned models are increasingly deployed in the real world, our work highlights the importance of understanding the limitations of pre-trained source models. Code is available at https://github.com/MadryLab/bias-transfer

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Context Attribution Handles What the Model Already Knows

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Context attribution methods cannot disentangle in-context from in-weight knowledge and assign unfaithful scores under overlap; new metrics and WMDP-Cyber++ quantify the failure.

  2. Assessing Intersectional Bias in Representations of Pre-Trained Image Recognition Models

    cs.CV 2025-06 reject novelty 4.0 of 10

    Pre-trained ImageNet classifiers encode age more strongly than race or gender in their activations, but the evidence is limited because linear probes overfit and do not generalize to unseen faces.

Pith tools