{"id":"a35bbbd0-21fa-44a5-bfc7-5adfbb282421","arxiv_id":"2506.08227","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Blind text-only likelihood models match or exceed CLIP on many compositionality benchmarks because positives and negatives differ systematically in length, plausibility, or image style.","lead":"This paper audits 17 vision-language compositionality benchmarks and shows that text-only models, ignoring images entirely, match or beat CLIP on several of them. It traces this to a distributional asymmetry between positive and negative captions and images built into how the benchmarks are constructed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Blind-baseline accuracy is reported for only 3 of 17 benchmarks, and 5 benchmarks are absent from both the likelihood analysis and accuracy tables; the abstract's blanket claim requires these to be representative.","rationale":"Read in good faith: the paper makes a specific, testable claim, and its strongest evidence is concrete. The likelihood-score definitions are explicit, LLaMA-2-13B is a public model, BlindCap is described with architecture and training budget, and Tables 2–4 show substantial blind performance on three benchmarks. That is real evidence and is not undermined by the fact that no code is released. However, the abstract's central claim is about 17 benchmarks. The evidence supports the three detailed cases and suggests, via Lik-Diff, that many others have distributional asymmetry; it does not establish that every benchmark in Table 1 is equally hackable. In particular, the five unexamined benchmarks span different constructions (e.g., CC-Neg uses negations; SVO-Probes uses templated captions; EqBen-Video is video-based), so representativeness cannot be assumed. This is a scope/strength-of-inference issue, not an internal inconsistency. It matches the reader's weakest-assumption analysis, so no change to the conditional verdict is needed; a targeted evaluation would either strengthen the paper into an unconditional accept or force a more qualified claim.","tokens_in":9660,"tokens_out":5440,"duration_ms":65950,"concrete_test":"Run the same LLaMA-13B normalized/unnormalized log-likelihood protocol (and BlindCap where feasible) on the six Table 1 benchmarks absent from the current analysis: What's-Up, CC-Neg, SugarCREPE++, SVO-Probes, EqBen-Video, and CounterCurate. Adopt a pre-specified threshold: if the best blind baseline fails to match or beat the reference CLIP/appropriate single-model baseline on a majority of these benchmarks, or if its average accuracy is within random-chance noise, then the claim should be narrowed from 'these benchmarks' to 'the three benchmarks tested in detail plus those with validated likelihood asymmetry.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline conclusion—that 'these benchmarks' (all 17 in Table 1) do not effectively measure compositional understanding—rests on two layers of evidence: detailed blind-baseline accuracy tables for SugarCREPE, VALSE, and VL-Checklist (Tables 2–4), and likelihood-difference scores for 12 datasets (Figure 2). Five benchmarks in Table 1 (What's-Up, CC-Neg, SugarCREPE++, SVO-Probes, EqBen-Video, CounterCurate) are never scored by a blind baseline, and Figure 2 even includes Flickr30K-Positions, which is not listed in Table 1. For the nine non-tabulated datasets in Figure 2, a large Lik-Diff is shown but not converted into task accuracy, so it is not demonstrated that those benchmarks are solvable by the heuristics. The load-bearing assumption is therefore that the three detailed benchmarks are representative of the full set and that likelihood asymmetry transfers to all remaining cases. This assumption is plausible for the widely used COCO-based benchmarks, but it is not established for text-to-image and video-derived benchmarks such as ColorSwap, COCO-Counterfactuals, and EqBen-Video. If those benchmarks are not shortcut-prone, the abstract overstates the scope of the finding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper audits 17 vision-language compositional benchmarks, arguing that their construction procedures induce a distributional asymmetry between positive and negative samples, and that this asymmetry allows blind, text-only heuristics to match CLIP on several benchmarks. The authors report log-likelihood baselines (LLaMA-13B) and two BlindCap variants (with and without COCO fine-tuning) on SugarCREPE, VALSE, and VL-Checklist, report likelihood-difference scores on 12 datasets, and derive recommendations for future benchmark construction.","tokens_in":9911,"tokens_out":5178,"duration_ms":59262,"significance":"If correct, the finding is significant: it would imply that recent progress on these benchmarks may partly reflect length, plausibility, COCO-domain, or synthetic-image artifacts rather than compositional understanding. The paper's systematic table of 17 benchmarks, use of multiple external probes (LLaMA-13B, BlindCap, and COCO fine-tuning), and concrete recommendations are useful contributions. However, the evidence is uneven across the benchmark set, so the scope of the conclusion needs to be tightened.","major_comments":[{"comment":"The abstract's blanket claim about 'these benchmarks' is not supported for the full set of 17. Detailed blind-baseline accuracy is provided for only SugarCREPE, VALSE, and VL-Checklist, and six benchmarks in Table 1 (What's-Up, CC-Neg, SugarCREPE++, SVO-Probes, EqBen-Video, and CounterCurate) appear in neither the accuracy tables nor Figure 2. For the nine non-tabulated datasets in Figure 2, Lik-Diff is not converted into task accuracy, so it remains possible that the likelihood asymmetry does not translate into solvability for those benchmarks. The authors should either add accuracy experiments for the missing benchmarks or explicitly restrict the conclusion to the evaluated subset.","section":"Abstract; Section 2, Tables 2–4, Figure 2"},{"comment":"Lik-Diff is defined with unnormalized log-likelihoods, which grow with caption length; this conflates length bias with plausibility bias and makes the reported asymmetry partly a restatement of the token-count difference. Reporting the normalized version L_norm alongside L_unnorm, or a length-controlled stratification, would make the 'distributional asymmetry' evidence cleaner and would separate the two biases that the paper otherwise treats as distinct.","section":"Section 2, Figure 2"},{"comment":"The tables report no error bars or significance tests, which weakens specific comparative claims. In VALSE, for example, the best blind baseline average exceeds CLIP by only 1.4 percentage points (65.4 vs. 64.0), while several category-level numbers flip depending on normalization and model variant; without uncertainty estimates, the 'on par or better' claim for VALSE is fragile. The authors should at least report bootstrap confidence intervals or multiple seeds for the trained BlindCap models.","section":"Section 2, Tables 2–4"}],"minor_comments":[{"comment":"References [41] and [42] are the same Tschannen et al. paper; the duplicate citation should be removed and the in-text citations renumbered.","section":"References"},{"comment":"The label 'VLChecklist' is inconsistent with 'VL-Checklist' in Table 1, and 'colorswap' should be 'ColorSwap'. More importantly, Flickr30K-Positions appears in Figure 2 but is not described in Table 1; the authors should clarify its origin and relationship to the 17-benchmark set.","section":"Figure 2"},{"comment":"The BlindCap training description is underspecified: 'using a ViT-B-32 vision-encoder and text-decoder with the same shape as the encoder, except half the depth' is ambiguous about whether the text decoder is a transformer, and the optimizer, learning rate, and evaluation details are missing. This hampers reproducibility.","section":"Section 2, BlindCap description"},{"comment":"The recommendation against single-model filtering is supported by only three model rows on two SugarCREPE subsets and the 7-average; reporting sample sizes and a broader set of filtering models would make the recommendation more persuasive.","section":"Section 3, Table 5"},{"comment":"The claim that '10 of the 17 benchmarks tested use COCO images directly' is not substantiated in the text; the authors should provide the list of these benchmark-dataset pairs in a table or appendix.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and important question for the vision-language benchmarks community, and the core direction is sound. The main issue is that the abstract and conclusion claim scope over all 17 benchmarks while the detailed accuracy evidence covers only three. This is fixable by adding experiments or narrowing the claims, so I recommend major revision rather than rejection. The duplicated reference [41]/[42] and inconsistent naming in Figure 2 suggest the final manuscript pass was rushed; these should be corrected in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nThis paper is worth your time. It is a systematic audit of 17 compositional VLM benchmarks, and the core finding is real: simple blind heuristics — token length, unnormalized LM likelihood, and text-only captioners — match or beat CLIP B/32 on SugarCREPE, VALSE, and VL-Checklist. The distributional-asymmetry explanation is convincing for those benchmarks, and the demonstration that SugarCREPE's own filtering model is still hackable by other models in the same family (Table 5) is a nice, non-obvious result. The COCO fine-tuning experiment is a clean external check that the bias is about the source distribution, not the probes.\n\nWhat's new: prior work poked at individual benchmarks. This paper connects them under one root cause and turns it into concrete construction advice — same-distribution negatives, multiple negatives, group matching. Those recommendations are sensible and actionable.\n\nThe soft spot is scope. Detailed blind-baseline accuracy is given for only 3 benchmarks. For the other 12 in Figure 2 they report likelihood differences, not task accuracy; a big likelihood gap suggests a shortcut but doesn't prove the benchmark is solvable by a blind heuristic. Five benchmarks in Table 1 (What's-Up, CC-Neg, SugarCREPE++, SVO-Probes, EqBen-Video, CounterCurate) appear in neither the accuracy tables nor the likelihood analysis. So the abstract's blanket claim about 'these benchmarks' is stronger than the evidence. I'd believe it for the COCO-derived ones, but text-to-image and video-derived benchmarks are genuinely untested here. Also: no error bars or significance tests, no code/checkpoints released, and Figure 2 includes Flickr30K-Positions, which isn't in Table 1. Minor: Tschannen et al. is cited twice as [41] and [42]. These are fixable in revision, not load-bearing flaws.\n\nThe central claim is not circular — the probes are external models. The direction of the finding is likely correct. This paper deserves a serious referee. I'd want the authors to either add accuracy for the missing benchmarks or explicitly narrow the conclusion, release code and checkpoints, and report variance. For now, treat the headline as 'most common benchmarks we checked have the problem,' not 'all 17 do.'\n\nWould bring to reading group; would cite as a cautionary reference.","headline":"A systematic, mostly convincing audit of compositionality benchmarks whose blanket scope claim runs ahead of the evidence.","tokens_in":10439,"tokens_out":3125,"would_cite":true,"duration_ms":37880,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Seventeen popular vision-language composition benchmarks are solvable by text-only heuristics at levels matching or beating CLIP, because benchmark construction introduces a systematic asymmetry between positives and negatives.","keywords":["vision-language models","compositional understanding","benchmark bias","distributional asymmetry","blind baselines","shortcut learning","negative sampling","CLIP"],"falsifier":"Run the same blind LLaMA and BlindCap baselines on the five surveyed benchmarks that lack detailed accuracy tables (including COLA, SVO-Probes, ColorSwap, COCO-Counterfactuals, and EqBen-Video) and compute their Lik-Diff values; if any of these shows near-zero likelihood asymmetry with blind baselines at chance while CLIP stays well above chance, the broad conclusion would not hold for that benchmark. A complementary test is to build a deliberately balanced benchmark—matching token length, plausibility, and image source between positives and negatives—and check that blind baselines collapse to chance while CLIP remains above chance.","tokens_in":9489,"feed_emoji":"🧩","tokens_out":8859,"duration_ms":88445,"temperature":0.7,"pith_summary":"The paper asks whether 17 widely used benchmarks for compositional understanding in vision-language models actually measure compositionality or can be solved through shortcuts. It shows that blind, image-free baselines—token-length counting, log-likelihood under a language model, and captioner models trained without image inputs—match or beat CLIP on several of these benchmarks. The cause is a distributional asymmetry between positive and negative examples: negatives are produced by different procedures (LLM rewriting, handcrafted swaps, diffusion generation) than the real positives, so the two classes differ systematically in length, plausibility, and image source. The paper concludes that these benchmarks do not effectively measure compositional understanding, and offers four recommendations for building more robust benchmarks. A reader should care because much of the reported progress in compositional reasoning may reflect dataset artifacts rather than model capability.","feed_headline":"Blind text-only baselines match CLIP on 17 composition benchmarks","feed_subtitle":"Leaderboard gains may reflect caption length and plausibility artifacts rather than compositional skill.","key_machinery":"The load-bearing mechanism is the distributional asymmetry between positive and negative samples, quantified by Lik-Diff, the average difference in unnormalized log-likelihoods that a language model assigns to positive versus negative captions. This statistic operationalizes four concrete biases—length bias, plausibility bias, COCO bias, and synthetic-image bias—and predicts when a blind text-only baseline can solve the task. The baselines that carry the argument are LLaMA-2-13B likelihood scoring (normalized and unnormalized) and BlindCap captioner models trained on DataComp-1B with and without COCO fine-tuning, all fed only the two captions at test time.","core_discovery":"The central discovery is that the 17 surveyed compositionality benchmarks harbor a distributional asymmetry between positive and negative examples that is an artifact of construction rather than a signal of compositionality. For image-to-text tasks, negative captions generated by rules, language models, or diffusion models turn out to be systematically longer, less plausible, or both; for text-to-image tasks, negative images produced by text-to-image models are detectable as synthetic. The authors demonstrate that LLaMA-2-13B scoring on unnormalized and normalized caption log-likelihoods, plus 'BlindCap' captioners trained without image inputs, reach accuracies on SugarCREPE, VALSE, and VL-Checklist that match or exceed CLIP-B/32—on average 9.5%, 1.4%, and 7.6% above CLIP for the best blind baselines, respectively—despite never seeing an image. A likelihood-difference statistic (Lik-Diff) computed between positive and negative captions is far from zero on 12 benchmarks, quantifying the asymmetry. The paper's claim is that leaderboard numbers on these benchmarks track this construction-induced asymmetry, not compositional reasoning.","pith_inferences":["Any future benchmark that generates negatives from a different distribution than its positives—via LLM rewriting, rule-based swaps, or diffusion models—should be audited with Lik-Diff before release, because the same asymmetry is likely to recur.","The Lik-Diff statistic could serve as a cheap pre-registration check: a benchmark designer could refuse to ship a dataset where a single language model cleanly separates positives from negatives.","The same blind-baseline methodology transfers to other multimodal abilities evaluated with synthetic negatives, such as spatial reasoning, ordering, and counterfactual understanding, where reported gains may be similarly confounded.","Because same-family models with different training data still exploit residual asymmetries, the authors' recommendations imply that robust benchmark construction may require adversarial co-training of negatives rather than one-off filtering."],"forward_implications":["Results on SugarCREPE, VALSE, and VL-Checklist should be reinterpreted as reflecting shortcut exploitation rather than compositional understanding, since blind baselines match or exceed CLIP on them.","Fine-tuning a blind captioner on COCO produces large accuracy gains on COCO-sourced benchmarks, showing that exposure to COCO statistics alone explains a substantial portion of performance.","Single-model filtering of negatives, as done for SugarCREPE with a plausibility model, is not a robust fix: other text-only models from the same family still achieve high accuracies on the filtered subsets.","Benchmarks built with multiple positives and negatives per image and group-based bidirectional matching, along the lines of the Winoground Group score, would be substantially harder to attack with blind heuristics.","Adopting the paper's four recommendations would produce evaluation setups in which the baseline chance level is lower and blind text-only attacks should lose their advantage."],"supporting_citations":[{"why":"SugarCREPE, the main benchmark attacked; its 'add' subset exhibits extreme length bias and its filtering is shown to be insufficient.","marker":"[15]"},{"why":"VALSE, one of three benchmarks with detailed blind-baseline tables where blind models outperform CLIP on 5 of 10 categories.","marker":"[31]"},{"why":"VL-Checklist, the third detailed benchmark where blind baselines match CLIP on attributes and relations.","marker":"[49]"},{"why":"LLaMA-2-13B, the language model used to compute log-likelihoods, normalized/unnormalized, and the Lik-Diff statistic.","marker":"[40]"},{"why":"CLIP-B/32, the reference vision-language model whose performance the blind baselines are compared against.","marker":"[33]"},{"why":"MS-COCO, the dominant data source whose statistics drive the COCO bias and the gains from COCO fine-tuning.","marker":"[25]"},{"why":"VERA, the single plausibility model used to filter SugarCREPE negatives, which the paper shows remains hackable (Table 5).","marker":"[26]"},{"why":"Winoground, the source of the Group score recommended for bidirectional image-text matching.","marker":"[38]"},{"why":"TextAttack, the grammar-based family used to construct SugarCREPE negatives; same-family models still exploit asymmetries.","marker":"[29]"},{"why":"The Cap-model architecture (Tschannen et al.) used to train the BlindCap baselines that achieve high accuracy without images.","marker":"[42]"}],"fun_headline_variants":["Text-only baselines rival CLIP on 17 composition benchmarks","Blind models match CLIP, but for the wrong reasons","Composition benchmarks skewed by caption length and plausibility","No vision needed: text heuristics top CLIP on composition","Leaderboard scores reflect artifacts, not compositional skill"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The blanket conclusion that all 17 benchmarks fail to measure compositional understanding assumes that the three benchmarks with detailed blind-baseline accuracy tables (SugarCREPE, VALSE, VL-Checklist) and the twelve benchmarks in the likelihood-asymmetry figure are representative of the full set.","fun_headline_variants_meta":{"raw":{"variants":["Text-only baselines rival CLIP on 17 composition benchmarks","Blind models match CLIP, but for the wrong reasons","Composition benchmarks skewed by caption length and plausibility","No vision needed: text heuristics top CLIP on composition","Leaderboard scores reflect artifacts, not compositional skill"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000283,"raw_usage":{"total_tokens":1662,"prompt_tokens":927,"completion_tokens":735,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":654}},"tokens_in":543,"tokens_out":735,"duration_ms":8728,"temperature":1.0,"reasoning_tokens":654,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:16:51.638501+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same blind LLaMA and BlindCap baselines on the five surveyed benchmarks that lack detailed accuracy tables (including COLA, SVO-Probes, ColorSwap, COCO-Counterfactuals, and EqBen-Video) and compute their Lik-Diff values; if any of these shows near-zero likelihood asymmetry with blind baselines at chance while CLIP stays well above chance, the broad conclusion would not hold for that benchmark. A complementary test is to build a deliberately balanced benchmark—matching token length, plausibility, and image source between positives and negatives—and check that blind baselines collapse to chance while CLIP remains above chance.","supporting_citations":[{"cited_title":"Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.NeurIPS,","cited_arxiv_id":null,"evidence_quote":"SugarCREPE, the main benchmark attacked; its 'add' subset exhibits extreme length bias and its filtering is shown to be insufficient."},{"cited_title":"Microsoft coco: Common objects in context","cited_arxiv_id":null,"evidence_quote":"MS-COCO, the dominant data source whose statistics drive the COCO bias and the gains from COCO fine-tuning."},{"cited_title":"Vera: A general- purpose plausibility estimation model for commonsense statements","cited_arxiv_id":null,"evidence_quote":"VERA, the single plausibility model used to filter SugarCREPE negatives, which the paper shows remains hackable (Table 5)."},{"cited_title":"Winoground: Probing vision and language models for visio- linguistic compositionality","cited_arxiv_id":null,"evidence_quote":"Winoground, the source of the Group score recommended for bidirectional image-text matching."},{"cited_title":"Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024","cited_arxiv_id":null,"evidence_quote":"The Cap-model architecture (Tschannen et al.) used to train the BlindCap baselines that achieve high accuracy without images."}],"review_version":1}