{"id":"e3d917fa-e34b-49d7-b3c2-fcf88e2eb140","arxiv_id":"2502.07957","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 131 CLIP models, the choice of pretraining data predicts intrinsic bias more than architecture or scale, and bias correlates with downstream performance in many test categories.","lead":"A study of 131 CLIP vision-language models finds that the pretraining dataset, not the architecture or model size, is the strongest driver of measurable social bias, and that bias often rises as downstream task performance improves. The result matters because it suggests current data-filtering methods that boost accuracy can inadvertently amplify stereotypes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The regression supporting the dataset-family claim treats 3,406 EAT scores as independent although dataset family varies only across 131 models; without model-level random effects or clustered SEs, the reported p-values likely overstate significance.","rationale":"The paper is a valuable large-scale descriptive study, and I do not dispute the raw correlations or the observation that dataset identity is associated with bias in this sample. The load-bearing weak spot is statistical: the regression that establishes dataset family as the dominant predictor ignores the clustering of EAT observations within models. Dataset family varies between models, not between the 26 EAT tests of the same model, so treating all 3,406 observations as independent inflates the effective sample size for exactly the coefficients the headline is about. This is distinct from, though related to, the reader's concern about family heterogeneity. Even if each dataset family is perfectly homogeneous, the current model cannot produce trustworthy p-values for family-level fixed effects without accounting for model-level correlation. The proposed refit with a model-level random intercept or clustered standard errors is straightforward with the released code and data, and it directly tests whether the key coefficients survive. I therefore keep the conditional verdict but add this as a necessary condition before the dataset-family claim can be accepted as established.","tokens_in":19409,"tokens_out":4942,"duration_ms":49536,"concrete_test":"Refit the Section 4 mixed-effects model with the same fixed effects but add a random intercept for model ID, or compute cluster-robust standard errors clustered by model, and re-examine the Figure 3 dataset-family coefficients and p-values. If 'dfn', 'commonpool', 'merged2b', 'webli', or 'datacomp' no longer reach p<0.01, the headline claim is not supported by the current analysis. As a secondary check, run a permutation test that shuffles dataset-family labels across the 131 models while preserving each model's 26 EAT scores, and compare the observed beta coefficients to the permutation null distribution.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative support for the claim that pretraining dataset choice predicts intrinsic bias is the mixed-effects regression in Section 4 (Eq. 1) and Figure 3. In that model, j indexes modality/test groups, and the only random effects are u0j, u1j, and u2j; there is no random intercept for model or dataset. Dataset family is a model-level fixed effect: all 26 EAT values from one CLIP model share the same dataset family, yet the regression conditions only on group means and treats the remaining residuals as independent. Because there are 131 models and 3,406 observations, the effective sample size for the dataset-family coefficients is at most 131, while the Wald/z p-values are computed as though n were 3,406. This can substantially understate standard errors and confidence intervals, making fixed effects such as 'dfn' (beta=0.608) appear significant when they may not be. Appendix B's convergence-motivated collapsing into dataset families does not fix this; it may even worsen it, because a family such as 'dfn' can contain checkpoints from one training run, allowing the coefficient to absorb training-recipe and checkpoint-selection effects. The paper's statement that dataset choice matters 'independent of architecture and parameter count' addresses only two potential confounders, not the model-level clustering of the data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a large-scale empirical study of intrinsic social bias in 131 CLIP models. The authors measure bias using 26 EAT/SEAT/iEAT-style tests across four modality combinations, and relate the resulting 3,406 effect sizes to upstream pretraining factors (dataset, architecture, parameter count, dataset size) via a mixed-effects regression, and to downstream VTAB+ performance via Pearson correlations. They report that pretraining dataset family is the most significant upstream predictor of intrinsic bias, that datasets curated with filtering techniques aimed at downstream performance tend to be associated with higher bias, and that intrinsic bias often correlates with downstream performance. They also introduce controlled, human-grounded attribute stimuli for the EATs and release their code and data.","tokens_in":19630,"tokens_out":3830,"duration_ms":35631,"significance":"If the regression result were statistically sound, this would be a valuable and much-needed large-scale comparison, indeed the largest such analysis of CLIP bias to date. The paper's strengths include the breadth of models (131), the use of established bias tests with improved human-grounded stimuli, the public release of code and data, and the connection between upstream curation choices and downstream performance. However, the central dataset-family claim is currently supported by a model that does not account for model-level clustering, so the key quantitative conclusion is not yet established. The contribution is significant and the empirical corpus is a useful resource, but the headline 'predicted by pretraining data' claim requires a statistically defensible regression.","major_comments":[{"comment":"The mixed-effects model includes random intercepts and slopes only for modality-by-test-order groups, not for models or dataset families. Each of the 131 models contributes 26 EAT observations that share a dataset family, so the residuals are correlated within model and within dataset family; the effective sample size for the dataset-family fixed effects is at most 131, not 3,406. The Wald p-values reported in Figure 3 are therefore likely anti-conservative, and it is uncertain whether coefficients such as dfn (beta = 0.608) or commonpool (beta = 0.399) survive when standard errors are clustered at the model or dataset-family level, or under a permutation test that shuffles models between families. Please re-fit with model-level random intercepts, cluster-robust standard errors, or a model-level permutation test, and report the corresponding confidence intervals.","section":"Section 4, Eq. (1); Section 5, Figure 3"},{"comment":"The convergence-driven grouping of datasets into families, combined with the absence of model-level random effects, leaves the dataset-family effect not separately identified from training recipe, optimization budget, or checkpoint-selection effects. A family such as 'dfn' or 'merged2b' may be represented by checkpoints from a single training run, and the regression conditions only on architecture and parameter count. The claim in the abstract and Section 6 that dataset choice is significant 'independent of other upstream factors such as model architecture or parameter count' is too strong; at best the analysis controls for the listed covariates. Please report the number of distinct models per dataset family and per architecture family, and discuss the remaining confounding with training recipe and scale.","section":"Appendix B"},{"comment":"The title and abstract state that intrinsic bias is 'predicted' by pretraining data, but the analysis is an in-sample regression with no out-of-sample validation, no cross-validation, and no predictive metric. The results are associational, not predictive. Please either rephrase to 'associated with' and 'explained by' throughout, or add a proper out-of-sample prediction evaluation, such as holding out entire dataset families before quantifying predictive accuracy.","section":"Title and Abstract"}],"minor_comments":[{"comment":"The heading 'EA Ts as an Aggregate Measure of Bias' contains an unnecessary space between 'EA' and 'Ts'; it should read 'EATs'.","section":"Section 5 heading"},{"comment":"The sentence about the YFCC15M subset contains a duplicated 'whose': 'whose whose title contains natural language' should be 'whose title contains natural language.'","section":"Appendix A.1.3"},{"comment":"The word 'instrinsic' is misspelled; it should be 'intrinsic.'","section":"Ethical Considerations"},{"comment":"The sentence 'while with training datasets are curated using multilingual and multicultural sources such as webli' is grammatically incomplete; it should read 'while some training datasets are curated using multilingual and multicultural sources such as webli.'","section":"Section 8 (Limitations)"},{"comment":"In the reference to Goh et al. (2021), 'V oss' should be 'Voss'.","section":"References"},{"comment":"The text says 'All code and data used in this study will be made available publicly' while the Introduction states 'We release our code and data at https://github.com/kshitishghate/CLIP_bias/.' Please make the release status consistent and specify the license and version for the released artifacts.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper makes a valuable empirical contribution, and the main issue is statistical rather than conceptual. I would accept a revision that adds model-level clustering or a suitable permutation test, softens the 'predicted' language, and addresses the grouping confound in Appendix B. The current version's central dataset-family claim is not yet statistically well supported, but the issue is fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is the largest systematic study of intrinsic bias in CLIP encoders to date, and its central claim — that pretraining dataset choice matters more than architecture — is probably correct in broad strokes. But the regression's p-values are overconfident, because the data are clustered by model and the analysis doesn't account for that.\n\nWhat is actually new: the scale. Prior work covered 9 models or fewer; this covers 131 models, 26 EATs, 3,406 measurements, and connects bias to both upstream factors and downstream performance in one framework. The controlled stimuli drawn from OASIS and NRC-VAD, with a measured 4.8% variance reduction, are a genuine methodological improvement with quantitative evidence attached. The finding that neural-filtered datasets (DFN, β=0.608) sit above heuristic-filtered ones (DataComp, β=0.360) is substantive and actionable for the data-curation community. The modality-dependent results — age bias flips direction between image and text — are interesting. Code and data are released.\n\nThe soft spots, in proportion: the stress-test concern is real and worth taking seriously. The mixed-effects model in Section 4 puts random effects only on modality/test-order groups. Dataset family is a fixed effect that varies at the model level, so the 3,406 observations are not independent for those coefficients; the effective sample size is closer to 131 models, or roughly 9 dataset families. Clustered standard errors or a model-level random intercept would be the standard fix. This doesn't kill the finding — the coefficients are large and the ordering across families is coherent — but the reported Wald p-values likely overstate significance.\n\nSecond, the title's 'predicted' overstates an in-sample regression; there's no out-of-sample validation. Soften the language or add a holdout check. Third, the bias-performance correlations (r=0.3-0.8) don't control for dataset family, which is a plausible common cause of both bias and VTAB+ performance; partial correlations are needed. Fourth, the dataset families partially encode training recipe — merged2b is essentially EVA-CLIP's pipeline, webli is PaLI's — so 'dataset choice' and 'training pipeline' are not fully separated. Minor: EAT uncertainty is not quantified, and the paper's own limitations section acknowledges several of these points but not the clustering issue.\n\nNone of this is fatal. The descriptive findings should survive in broad strokes; the quantitative precision needs work.\n\nWho is this for: fairness/auditing researchers, anyone working on CLIP data curation, and benchmark builders. It deserves a serious referee — the fixes are standard regression hygiene, and the contribution justifies a revision cycle.","headline":"The dataset-dominance finding is likely right in broad strokes, but the regression's p-values are overconfident because the 3,406 measurements are clustered in 131 models and that clustering is not modeled.","tokens_in":20187,"tokens_out":3641,"would_cite":true,"duration_ms":31443,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Pretraining data, not architecture, predicts bias in CLIP models, and high-performing models tend to carry more bias.","keywords":["CLIP","intrinsic bias","embedding association tests","vision-language models","pretraining data","dataset curation","zero-shot performance","social bias"],"falsifier":"Train two sets of CLIP encoders with identical architecture, parameter count, and training recipe on the same raw image-text pool, one filtered with a performance-oriented pipeline and one with a low-bias heuristic pipeline, then rerun the 26 EATs; if the bias gap disappears, the 'dataset choice drives bias' claim fails. Alternatively, add training-recipe covariates such as batch size, epochs, or filtering network weights to the mixed-effects model and check whether the dataset-family coefficients survive.","tokens_in":19182,"feed_emoji":"⚖️","tokens_out":6826,"duration_ms":57978,"temperature":0.7,"pith_summary":"What this paper is trying to establish: the biases that show up inside vision-language encoders are primarily a property of the data they were pretrained on, not of the architecture or size of the model, and that those biases move together with downstream task performance. It reaches this by measuring 131 CLIP-style models with 26 association tests spanning images, text, and cross-modal combinations, producing 3,406 bias effect sizes. The strongest upstream predictor is the dataset family, and datasets that were aggressively filtered for accuracy like dfn, commonpool, and datacomp carry the highest bias. The paper also reports that bias and performance are often positively correlated ($0.3 \\le r \\le 0.8$), meaning the same training choices that make models better at benchmarks can make them more stereotyped. If true, this shifts the burden of bias mitigation onto data curation rather than model design alone.","feed_headline":"Pretraining data predicts CLIP bias more than architecture","feed_subtitle":"Across 131 vision-language encoders, performance-focused data filtering tracks stronger stereotype associations.","key_machinery":"The load-bearing measurement is the EAT effect size $d$, the standardized difference in mean cosine similarity between target-concept embeddings (e.g., flowers, instruments, women, European Americans, young people) and pleasant versus unpleasant attribute embeddings. The statistical engine is a mixed-effects regression with random intercepts and slopes that attributes variance in 3,406 $d$ observations to dataset family, architecture family, log-parameter count, and log-dataset size; Pearson correlations against the VTAB+ benchmark then link bias to downstream zero-shot performance. The paper also introduces human-grounded attribute stimuli (OASIS images and NRC-VAD words) to reduce noise in the EAT estimates.","core_discovery":"Studying 131 CLIP encoders with 26 Embedding Association Tests (EATs) across four modality combinations, the paper's central claim is that the pretraining dataset family is the dominant upstream predictor of intrinsic bias, with statistically significant positive coefficients for performance-focused families like 'dfn' ($\\beta = 0.608$), 'commonpool' ($\\beta = 0.399$), and 'merged2b' ($\\beta = 0.396$) relative to the CC12m baseline, while no architecture family shows a significant effect. The same analysis finds that stronger intrinsic bias often accompanies better downstream zero-shot performance, with correlations between $0.3$ and $0.8$ for non-human associations and negative correlations for gender/valence in some modality settings. The authors interpret this as evidence that optimizing models for performance, especially through automated data filtering, can inadvertently amplify representational stereotypes, and that bias is modality-dependent rather than uniform across text and image.","pith_inferences":["If dataset curation is the dominant lever, then interventions like balanced resampling or demographic parity filtering are likely to be more cost-effective than architectural changes or scaling alone, a testable prediction the paper does not make.","The positive bias-performance correlation suggests intrinsic bias measures could be used as cheap, representation-level signals during training, but only if the correlation is causal rather than a shared confound with data quality, something the paper's correlational design cannot separate.","The modality-dependence of bias raises a route to mitigation: aligning text and image representations for a social category may dampen cross-modal stereotypes, since opposite-sign age associations appear in text versus image."],"forward_implications":["Bias audits of CLIP-style models should report the pretraining dataset family; models from the same family will likely cluster in bias regardless of architecture.","Performance-oriented data filtering pipelines need fairness constraints built in; the current approach of filtering for benchmark accuracy alone appears to raise intrinsic bias.","Correlations between bias and zero-shot performance mean benchmark rankings can systematically favor more stereotyped models, so accuracy and fairness should be evaluated together.","Modality-specific results imply unimodal audits, whether text-only or image-only, understate cross-modal bias; evaluation should cover all four modality combinations.","Data curation choices such as hypernymizing names to a generic '[PERSON]' token may lower bias, consistent with the low-bias CC12m reference family."],"supporting_citations":[{"why":"Introduces the CLIP training objective and the original WebImageText models that anchor the study's model set.","marker":"(Radford et al., 2021)"},{"why":"Establishes the Embedding Association Test (EAT) methodology for measuring intrinsic bias from embeddings.","marker":"(Caliskan et al., 2017)"},{"why":"Defines the Image EAT (iEAT) used for vision-modality bias measurement.","marker":"(Steed and Caliskan, 2021)"},{"why":"Defines the Sentence Encoder Association Test (SEAT) used for text-modality bias measurement and supplies template stimuli.","marker":"(May et al., 2019)"},{"why":"Provides the LAION-5B pretraining data and the VTAB+ zero-shot benchmark used for downstream performance correlations.","marker":"(Schuhmann et al., 2022)"},{"why":"Source of the reproducible OpenCLIP models spanning LAION subsets in the study.","marker":"(Cherti et al., 2023)"},{"why":"Defines the DataComp benchmark and CommonPool datasets whose filtering strategies are associated with higher bias.","marker":"(Gadre et al., 2024)"},{"why":"Introduces Data Filtering Networks (dfn), the dataset family with the largest positive bias coefficient in the regression.","marker":"(Fang et al., 2023b)"},{"why":"Describes the CC12M dataset and its hypernymization strategy, the low-bias baseline family.","marker":"(Changpinyo et al., 2021)"},{"why":"Prior comparison of bias across nine CLIP models; its claim that larger datasets reduce bias is contradicted by the present study.","marker":"(Berg et al., 2022)"}],"fun_headline_variants":["Pretraining data, not architecture, sets CLIP bias levels","Data curation for performance inflates CLIP stereotypes","CLIP bias correlates with downstream accuracy","Architecture hardly matters; pretraining data sets bias","Performance-focused data curation amplifies CLIP bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that dataset choice, not architecture or scale, drives bias rests on the assumption that grouping 131 models into dataset families and architecture families leaves each family homogeneous enough in training recipe and compute that the regression's dataset-family coefficients isolate the effect of data composition rather than capturing differences in optimization or checkpoint availability.","fun_headline_variants_meta":{"raw":{"variants":["Pretraining data, not architecture, sets CLIP bias levels","Data curation for performance inflates CLIP stereotypes","CLIP bias correlates with downstream accuracy","Architecture hardly matters; pretraining data sets bias","Performance-focused data curation amplifies CLIP bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00063,"raw_usage":{"total_tokens":2934,"prompt_tokens":994,"completion_tokens":1940,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":1867}},"tokens_in":610,"tokens_out":1940,"duration_ms":14481,"temperature":1.0,"reasoning_tokens":1867,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T11:17:18.385226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train two sets of CLIP encoders with identical architecture, parameter count, and training recipe on the same raw image-text pool, one filtered with a performance-oriented pipeline and one with a low-bias heuristic pipeline, then rerun the 26 EATs; if the bias gap disappears, the 'dataset choice drives bias' claim fails. Alternatively, add training-recipe covariates such as batch size, epochs, or filtering network weights to the mixed-effects model and check whether the dataset-family coefficients survive.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Image EAT (iEAT) used for vision-modality bias measurement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the LAION-5B pretraining data and the VTAB+ zero-shot benchmark used for downstream performance correlations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the reproducible OpenCLIP models spanning LAION subsets in the study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the CC12M dataset and its hypernymization strategy, the low-bias baseline family."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior comparison of bias across nine CLIP models; its claim that larger datasets reduce bias is contradicted by the present study."}],"review_version":1}