{"id":"b0faa011-a4a5-4e10-9817-f1e54381b312","arxiv_id":"2607.28211","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Model scale loses predictive power for multi-attribute bias robustness in VLMs, while training-data size and curation remain the more reliable levers.","lead":"A 194-model audit finds that bigger vision-language models stop helping—and can hurt—when biases involve multiple shortcuts at once. Training-data size and curation stay predictive, so model pickers who only chase ImageNet may choose fragile systems.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper’s strongest claim is an empirical regularity inside a clearly bounded audit, not a universal law of VLM bias. The evidence for that regularity is multi-legged (full-population Spearman tables, architecture/data-matched ablations, subgroup conflict analysis in Table 7, Pareto overlap). The main vulnerability is external validity—exactly what the reader flagged and what §5 already states—rather than an insecure step in the argument as written. Extending to more benchmarks or larger proprietary models would strengthen generality but is not required to accept the stated finding. Hence no verdict movement: ACCEPT / HIGH stands.","tokens_in":17030,"tokens_out":471,"duration_ms":10817,"concrete_test":"Re-run the 26 controlled size-only comparisons of Table 2 after recomputing UrbanCars WGA with the 80-template ImageNet prompt ensemble already used in the Discussion sensitivity check; if the mean Δ WGA remains negative (near −4%) and the ImageNet\to C elebA\to UrbanCars ρ decay is preserved, the central empirical claim is stable under the paper’s own robustness protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest-assumption note (external validity of CelebA/UrbanCars + OpenCLIP ≤3.6B as proxies for “bias complexity”) is a real scope limit, but it is already disclosed in §5 and does not undermine the internal claim the paper actually makes: within this 194-model population and these two group-structured benchmarks, scale’s Spearman association decays (0.68\to0.48\to0.05) and matched size-only steps average +2.63% ImageNet / −4.22% UrbanCars WGA while data size/curation remain more stable. Population correlations, 11 size-matched groups (Table 2), dataset-matched groups, and the prompt-sensitivity check in Discussion all point the same way; there is no internal contradiction or hidden statistical failure that would overturn that scoped result.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper audits 194 public OpenCLIP VLMs (16 families, 63M–3.6B params, 24 training datasets) on ImageNet accuracy and worst-group accuracy (WGA) on CelebA (single spurious attribute) and UrbanCars (two simultaneous spurious attributes). It reports that Spearman correlation of model scale with performance decays from ρ=0.68 (ImageNet) to 0.48 (CelebA WGA) to 0.05 n.s. (UrbanCars WGA), while data size remains more stable (ρ=0.59→0.57→0.41) and curated data can improve UrbanCars WGA by up to ~25% at matched scale. Matched comparisons isolating scale, data, patch size, and resolution, plus a subgroup analysis on bias-aligned vs. bias-conflicting UrbanCars groups, support a “bias complexity sensitivity” framing and practical guidance favoring data quality and fine tokenization over pure scale.","tokens_in":17194,"tokens_out":1296,"duration_ms":41662,"significance":"If the scoped empirical pattern holds, the work is a useful corrective to scale-centric VLM selection: overall accuracy and parameter count are weak proxies for multi-attribute shortcut robustness, whereas training-data size and curation remain more informative. Strengths include the breadth of the public-checkpoint audit (194 models), explicit matched-group controls (Tables 2, 4, 5; dataset isolation in Fig. 4), subgroup decomposition (Table 7), a prompt-sensitivity check in §5, Pareto/safe-model tables, and released code. The contribution is primarily empirical and ecosystem-level rather than methodological, but it is timely for practitioners choosing open VLMs and for evaluation practice that still overweights ImageNet-style averages.","major_comments":[{"comment":"Table 2 and the abstract/Fig. 1 framing: the mean UrbanCars effect of scaling is −4.22%, but the sign split is 12 positive / 13 negative with large configuration-dependent swings (e.g., large drops in some SigLIP2 and MetaCLIP2 rows, large gains in some CLIPA rows). “Scaling reduces WGA by −4.2%” overstates a systematic negative effect; the load-bearing, better-supported claim is that scale is unreliable / loses predictive power and is not a robust lever once bias is multi-attribute. Please rephrase abstract, Fig. 1, §4.2, and Conclusion to lead with unreliability/no consistent benefit, and report the mean with dispersion (e.g., median and IQR or a simple sign test) so the average is not read as a near-universal reduction.","section":"Table 2; Abstract; Figure 1; §4.2 Model Scale"},{"comment":"§3.3 and the “bias complexity sensitivity” construct treat the ImageNet→CelebA→UrbanCars progression as increasing spurious-correlation complexity (M=0/1/2). That axis is confounded with domain (faces vs. synthetic car scenes), realism (natural CelebA vs. pasted UrbanCars), and spatial structure of cues. §5 discloses limited benchmarks, but the central named claim still leans on “complexity.” Either (i) add at least one additional group-structured multi-attribute or non-synthetic bias benchmark, or (ii) narrow the claim language throughout to “across these two established group-structured bias benchmarks of increasing attribute count,” and discuss the confound explicitly in §4.1/§5 so the title-level generalization is not stronger than the design.","section":"§3.3 Benchmarks; §4.1; §5 Discussion"}],"minor_comments":[{"comment":"Table 1 reports Tokens ρ=+0.27*** on UrbanCars while Parameters are n.s.; a brief partial-correlation or note on collinearity among parameters, tokens, patch size, and resolution would clarify how much “representation” is separable from scale.","section":"Table 1; §4.1"},{"comment":"§3.6 prompt construction for CelebA is asymmetric (more non-blonde descriptors plus brunette/bald). The 80-template check in §5 is helpful; consider reporting the primary WGA numbers under that protocol in an appendix table for reproducibility.","section":"§3.6; §5"},{"comment":"Figure 4 sorts datasets by UrbanCars mean but vertical spread is large for some corpora; stating n per dataset (models entering each mean) in Table 3/Fig. 4 would prevent over-reading small-n leaders.","section":"Figure 4; Table 3"},{"comment":"Minor clarity: define WGA at first use in the main text before Fig. 1’s caption reliance; fix “gopt” / naming consistency for SigLIP2 variants if space allows in tables.","section":"Figure 1; Table 2"},{"comment":"Related work could briefly situate UrbanCars’ multi-shortcut “whac-a-mole” finding [25] against the paper’s pretraining-design focus so the novelty boundary is sharper.","section":"§2"}],"recommendation":"minor_revision","confidential_remarks":"I agree with the reader’s high-level accept lean: internally the correlations and matched comparisons cohere, and the external-validity limit is largely disclosed. I stopped at minor_revision rather than accept only because the abstract/Fig. 1 “scaling reduces WGA by −4.2%” line is punchier than Table 2’s near-even sign split, and the complexity framing slightly outruns the two-benchmark design. Both are fixable in revision without new theory. Scope (OpenCLIP ≤3.6B, zero-shot only) is appropriate for an open-ecosystem audit and not a reject reason for this venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: on 194 public VLMs, parameter count tracks ImageNet (ρ=0.68) and still somewhat CelebA WGA (0.48), then collapses on UrbanCars multi-attribute WGA (0.05, n.s.). In 26 size-only matched steps the mean moves are +2.6% ImageNet and −4.2% UrbanCars WGA. Data size stays correlated (down only to 0.41), and curated vs unfiltered at matched ~2B scale shows gaps up to ~25% WGA. That is the result.\n\nWhat is new is not the slogan that data matters, but the ecosystem-scale measurement plus the controlled slices (11 scale groups, 7 data, 7 patch, 25 resolution) and the subgroup split on UrbanCars (scale helps bias-aligned samples, weakens as conflicts pile up). They ship the eval code and stick to public OpenCLIP weights, so the numbers are checkable. Citations sit on the right prior work (OpenCLIP scaling, DataComp/DFN, Fang et al. on data and CLIP shifts, single-attribute accuracy-on-the-line). No internal contradiction; the Discussion already flags the prompt check and the ≤3.6B / two-benchmark scope.\n\nSoft spots are real but proportionate. “Bias complexity sensitivity” is mostly a label for the observed decay across two group-structured benchmarks, one of them synthetic. Zero-shot prompts and matched-group inclusion are free parameters, though the extra 80-template check does not flip the story. External validity beyond OpenCLIP and these two tasks is unproven — and they say so. None of that undoes the scoped empirical claim.\n\nThis is for people who pick VLMs, build data recipes, or care about multi-shortcut failure modes. Not a methods paper. I would bring it to reading group, cite the controlled tables when arguing against scale-only leaderboards for robustness, and send it to referees without hesitation.","headline":"Large OpenCLIP audit shows scale stops predicting multi-attribute WGA while data size/curation hold up; scoped but solid and worth engaging.","tokens_in":17860,"tokens_out":494,"would_cite":true,"duration_ms":10100,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Scaling vision-language models improves average accuracy but fails against multi-attribute bias; training data size and curation hold up instead.","keywords":["spurious correlations","bias","vision-language models","scaling laws","worst-group accuracy","data curation","zero-shot robustness","shortcut learning"],"falsifier":"Find a large matched family of VLMs where increasing parameter count alone, with data and architecture fixed, reliably raises worst-group accuracy on a multi-attribute bias benchmark comparable to UrbanCars; or show that on several additional multi-attribute benchmarks the scale–WGA correlation stays strong rather than collapsing near zero.","tokens_in":17894,"feed_emoji":"📉","tokens_out":955,"duration_ms":24297,"temperature":0.7,"pith_summary":"Vision-language models are usually chosen by how well they score on broad tests like ImageNet, under the quiet assumption that bigger, stronger models are also less fooled by shortcuts. This paper tests that assumption across 194 public models by comparing overall accuracy with worst-group accuracy on a single spurious cue (CelebA hair color vs gender) and on two simultaneous cues (UrbanCars car type vs background and co-occurring object). The link between parameter count and performance collapses as the bias gets more complex, and controlled size-only comparisons even show multi-attribute worst-group accuracy falling when models get larger. Dataset size and especially curation quality keep a clearer relationship with robustness, with curated data beating size-matched unfiltered data by large margins. The practical message is that practitioners who care about shortcut resistance should treat data quality as a first-class lever and stop treating scale as a reliable fix.","feed_headline":"Bigger VLMs do not fix multi-attribute bias","feed_subtitle":"Across 194 models, scale stops predicting worst-group accuracy; curated data still helps by up to 25%.","key_machinery":"Bias complexity sensitivity: the progressive decay, across a ladder of benchmarks with more simultaneous spurious attributes, of the usual scaling relationship between model size and performance, measured by worst-group accuracy on the group structure induced by label times spurious attributes.","core_discovery":"As evaluation moves from overall ImageNet accuracy to single-attribute then multi-attribute worst-group accuracy, the Spearman correlation of model scale with performance falls from 0.68 to 0.48 to a non-significant 0.05, while training-data size stays meaningfully correlated and curated datasets improve worst-group accuracy by up to about 25 percent over uncurated alternatives at matched scale. In matched comparisons that vary only backbone size, scaling averages +2.6 percent ImageNet accuracy but −4.2 percent UrbanCars worst-group accuracy.","pith_inferences":["If scale mainly improves context exploitation, larger models may look better on bias-aligned slices while quietly worsening the fully conflicting minority groups that WGA cares about.","The same pattern may appear in other foundation-model settings (audio-language, video-language) whenever multiple non-causal cues can be exploited at once.","Procurement and safety review for deployed VLMs may need mandatory multi-shortcut audits rather than relying on demographic single-attribute probes alone.","Data-curation methods that optimize average CLIP score may still leave multi-attribute holes; filters may need explicit worst-group or conflict-aware objectives."],"forward_implications":["Model cards and leaderboards that rank VLMs only by ImageNet-style averages will systematically mis-rank models for shortcut robustness.","At fixed compute, investing in data filtering and curation is a more reliable path to worst-group gains than simply enlarging the backbone.","Multi-attribute bias suites should become standard release criteria alongside average zero-shot accuracy.","Token granularity (patch size and resolution) should be tuned to the spatial nature of the expected shortcuts, not only to average accuracy.","Pareto-strong models in this audit share curated or large data, finer tokens, and moderate—not maximal—scale."],"fun_headline_variants":["VLM scale correlation collapses on multi-attribute bias","Curated data lifts worst-group accuracy up to 25% over scale","Bigger backbones gain ImageNet but lose on UrbanCars bias","Data size tracks bias robustness; model scale does not","From 0.68 to 0.05: scale stops predicting VLM bias results"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That two zero-shot group-structured benchmarks—one face attribute with a single demographic cue, one synthetic car scene with two controlled shortcuts—plus the public OpenCLIP checkpoint set are enough to stand in for how scale behaves under real multi-attribute bias.","fun_headline_variants_meta":{"raw":{"variants":["VLM scale correlation collapses on multi-attribute bias","Curated data lifts worst-group accuracy up to 25% over scale","Bigger backbones gain ImageNet but lose on UrbanCars bias","Data size tracks bias robustness; model scale does not","From 0.68 to 0.05: scale stops predicting VLM bias results"]},"model":"grok-4.5","effort":"low","cost_usd":0.004312,"raw_usage":{"total_tokens":1295,"prompt_tokens":810,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":43124000,"prompt_tokens_details":{"text_tokens":810,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":407,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":810,"tokens_out":78,"duration_ms":7404,"temperature":1.0,"reasoning_tokens":407,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T14:30:50.201818+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Find a large matched family of VLMs where increasing parameter count alone, with data and architecture fixed, reliably raises worst-group accuracy on a multi-attribute bias benchmark comparable to UrbanCars; or show that on several additional multi-attribute benchmarks the scale–WGA correlation stays strong rather than collapsing near zero.","supporting_citations":[],"review_version":1}