{"id":"3422bb85-69f1-44b2-9015-89feee0732ce","arxiv_id":"2507.07985","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Using a controllable synthetic dataset, the authors show that common data properties, especially saliency bias, prevent CLIP models from learning object-attribute binding, and that fixing the data achieves near-perfect binding.","lead":"This paper trains CLIP-style models on a carefully controlled synthetic dataset and finds that ordinary properties of image-caption data, such as captions that mention only some objects or describe mainly the most salient object, strongly limit whether the model can bind attributes like color to the correct object. The result suggests that the well-known bag-of-words failure of CLIP may be more a data problem than a model or training-loss problem.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The saliency-bias result is confounded: salient objects are always centered and always mentioned first in captions, so the drop in binding accuracy may reflect a spurious center/order cue rather than saliency itself; this weakens the claim that saliency is a key limiting factor on real data.","rationale":"The paper is a careful controlled study: the four data properties are cleanly manipulated, trends are validated across batch size and embedding size, OOD checks indicate compositional rather than memorized representations, and code is released. The reader's weakest assumption—that TinyCLIP results transfer to real-scale CLIP and that the realistic parameters estimated from 100 annotated CC12M images are representative—is a genuine concern and is explicitly acknowledged in the paper's limitations paragraph. However, I find a more precise soft spot in the internal validity of the saliency manipulation itself. The operationalization of saliency bundles center position, guaranteed caption inclusion, and deterministic first-mention order into a single variable. The paper's own footnote admits the text encoder can detect the salient object from word order alone. This means the drop in binding under high saliency could be explained by the model exploiting a deterministic positional/order shortcut, not by human-like saliency bias. Natural captions are far more stochastic: the salient object is not always centered and is not always mentioned first. Consequently, the strong real-data conclusion about saliency is the least secure part of the central claim. The 100-image annotation is secondary but compounds this: the saliency estimate of about 0.9 is used to argue that real data lies in the detrimental regime, and if the confound drives the effect, that argument collapses. The other data-property results—attribute density and caption completeness—are not similarly confounded and remain credible. I therefore do not propose changing the reader's CONDITIONAL verdict; I would only sharpen the required follow-up to include a deconfounded saliency ablation before accepting the saliency-specific claim.","tokens_in":18255,"tokens_out":12069,"duration_ms":138017,"concrete_test":"Run the saliency experiment with the confound removed: in MADMAN, keep the salient object centered and always in the caption, but randomize whether it appears first or second in the caption (or, alternatively, keep caption order fixed while randomizing the salient object's grid position). If binding accuracy at p(saliency)=0.9 recovers toward the ideal-data level, the saliency effect in Fig. 4b is an artifact of the center/first confound; if it stays near chance, saliency itself, as modeled, is harmful. A second check: evaluate models trained with saliency on test images that also contain a centered salient object. If binding is high in that case, the effect reflects a train/test distribution shift rather than an inability to bind.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 defines saliency bias in MADMAN such that a salient object is always in the center cell, always mentioned in the caption, and always mentioned first. The footnote in Section 4.2 concedes that the text encoder can infer which object is salient precisely because it appears first. Thus the saliency manipulation simultaneously changes three correlated factors: spatial center bias, guaranteed caption inclusion, and deterministic word order. The observed decline in binding accuracy as p(saliency) increases (Fig. 4b) could be caused by the model learning a shortcut—align the first noun phrase with the center cell—rather than by saliency as a real-world captioning tendency. Natural image-caption data has no such deterministic positional/order rule; the salient object is often but not always central and is not always mentioned first. The claim that 'saliency bias' is a key limiting factor on real data therefore rests on an unvalidated operationalization. This is load-bearing because the abstract and conclusion single out saliency as a critical factor, and Table 6 shows that replacing ideal data with realistic saliency causes the largest single-factor drop. If the effect is an artifact of the confound, the paper's headline list of harmful data properties is partially wrong.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MADMAN, a synthetic dataset built from transformed MNIST digits in a grid, to study how four data properties affect object-attribute binding in contrastive vision-language models. The properties are p(two-obj-img), p(two-obj-cap), expected number of attributes per object in a caption, and a saliency bias. The authors train small CLIP-like transformers on controlled variants and measure zero-shot binding accuracy with recognition filtering. They report that settings calibrated to 100 CC12M image-caption pairs produce near-chance binding, while a manually chosen 'ideal' setting produces almost perfect binding; they also report that batch size, embedding size, and hard negatives (NegCLIP) do not close the gap, and that the trends persist on out-of-distribution object-attribute combinations. They conclude that CLIP's binding failures are largely caused by common properties of the training data, with saliency bias singled out as a key factor.","tokens_in":18429,"tokens_out":5430,"duration_ms":74920,"significance":"The study is well-executed within its synthetic regime: three seeds with confidence intervals, hyperparameter sweeps across batch and embedding sizes, OOD checks, and unfiltered ablations are included, and code is provided. A data-centric testbed for binding is useful to the community and could redirect attention from loss and architecture fixes toward data curation. If the transfer assumption holds, the work would motivate filtering and re-captioning as concrete remedies. The main unvalidated link is the extrapolation from TinyCLIP on MADMAN to CLIP on natural data, and the saliency operationalization conflates several cues. These issues affect the headline causal attribution but not the internal consistency of the controlled trends.","major_comments":[{"comment":"The saliency manipulation is confounded. In MADMAN, a salient object is always placed in the center cell, is always mentioned in the caption, and is always mentioned first. The footnote in Section 4.2 concedes that the text encoder can infer the salient object from first-mention order. Thus increasing p(saliency) simultaneously increases center-position predictability, caption-inclusion rate, and word-order determinism, and the decline in binding accuracy in Fig. 4b and the large drop in the 'Ideal except Saliency bias' row of Table 6 cannot be uniquely attributed to saliency as a real-world captioning tendency. Natural image-caption data does not contain these deterministic positional and order cues. Please add ablations that disentangle the three factors, for example, randomize grid placement while keeping first mention and guaranteed inclusion, or randomize order and omission while keeping center placement, and report the binding curves for each sub-manipulation separately.","section":"Section 3.1, Fig. 4b, Table 6"},{"comment":"The headline conclusion that 'CLIP models are able to learn object-attribute binding but properties of the data inhibit this' is stated for CLIP, while the experiments train a small transformer (TinyCLIP, Appendix Table 4) on a synthetic dataset. The bridge is the assertion in the Introduction that 'the learning algorithm does not change its behavior fundamentally in larger scale training on more complex data,' plus the realistic parameter estimates from 100 manually annotated CC12M images (Appendix 5.2). A 100-image sample gives wide confidence intervals, and the annotation procedure for saliency in CC12M differs from the synthetic definition: in CC12M an image is labeled as having salient objects if there are any non-salient objects, whereas in MADMAN saliency means center placement plus guaranteed first mention. The paper should provide direct evidence for transfer, such as a real-data re-captioning or filtering experiment, or at least a larger-scale synthetic scaling study, otherwise the title and abstract generalize beyond the evidence.","section":"Section 4.1 and Introduction"},{"comment":"The 'ideal data' setting is selected by reading off the best-performing value of each property from the same binding-accuracy curves that are then used to report the ideal result in Table 1. This makes the high binding in Table 1 a post-hoc optimum rather than an independent validation of the prescription. If the paper claims that 'only when the data expresses our identified data properties does CLIP learn almost perfect binding,' an independent check is needed, for example, choosing the best parameters on a development split and reporting performance on a held-out split, or performing a small random search over the four-dimensional property grid and reporting the out-of-sample best. Without such a check, the 'only when' phrasing overstates what the one-at-a-time sweeps establish.","section":"Section 4.2 and Table 1"}],"minor_comments":[{"comment":"The term 'CLIP' is used for a small transformer trained from scratch on MADMAN; this is only clarified as TinyCLIP in Appendix Table 4. State this at first use in the main text to avoid ambiguity.","section":"Throughout main text"},{"comment":"The note that the search was performed 'by chance partially in the optimal setup' is confusing; it appears that the base setup already equals the best setting for three of the four properties, so Table 5 does not isolate the individual contribution of each property. Please rewrite the note to explain what was and was not varied.","section":"Appendix 5.3, Table 5"},{"comment":"The figure shows two distributions labeled E[na] = 3.50 with different shapes; one of the labels is likely a typo and should be corrected.","section":"Appendix Figure 7"},{"comment":"The realistic saliency annotation defines an image as having salient objects if there are any non-salient objects, which is not the same as the MADMAN parameter p(saliency). The mapping between the binary annotation and the synthetic probability should be explained explicitly.","section":"Appendix 5.2"},{"comment":"The recognition-filtering threshold of 1.1 times chance level is arbitrary and is not sensitivity-analyzed; please report how many samples are retained per attribute and condition, and whether the main trends change with a stricter or looser threshold.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical study with careful controls inside the synthetic setup. The saliency confound and the synthetic-to-real transfer assumption are both fixable with additional experiments or substantially softened claims; they are not fatal to the value of the dataset and the per-property trends. I do not see a novelty disclosure concern, but the authors should make sure the contribution is positioned relative to existing data-centric CLIP analyses such as Xu et al. (2024)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line up front: this is a careful, controlled study that makes a real contribution to the CLIP binding literature. The authors build a synthetic dataset (MADMAN) and show that four data properties—fraction of multi-object images, fraction of multi-object captions, attributes per object, and saliency bias—each move binding accuracy. The inverse U-shaped effect of attribute density is new and believable. So is the result that hard negatives alone don't get you there. The experiments are well executed: three seeds, confidence intervals, batch/embedding size sweeps, OOD checks, and they report unfiltered versions of the binding curves as well. The citation pattern is fair; they engage with the hard-negative and architecture-centric prior work without overclaiming novelty. Code is released. I'd trust the within-synthetic-regime claims.\n\nThe soft spots are at the boundary between the synthetic regime and real data. The authors want to conclude that data properties, not architecture or loss, limit CLIP's binding on natural data. That conclusion leans on two things: a stated but unverified transfer assumption (TinyCLIP on MADMAN behaves like CLIP at scale), and realistic parameter estimates from 100 manually annotated CC12M images. They acknowledge the transfer gap in their limitations paragraph, which is honest, but the paper's abstract and conclusion lean on it more than the evidence supports.\n\nThe bigger issue, which I think a careful referee should push on, is the saliency result. Salient objects in MADMAN are always centered, always mentioned in the caption, and always mentioned first. The footnote in Section 4.2 even says the text encoder can infer which object is salient because it appears first. So the drop in binding accuracy as p(saliency) increases could come from a spurious center/order shortcut rather than from saliency as a captioning tendency. Natural captions don't have that deterministic positional rule. Since the paper singles out saliency as a key limiting factor—Table 6 shows it causes the largest single-factor drop—this confound cuts the headline claim. The other three data properties don't have this problem, and the multi-object and attribute-density results stand on their own.\n\nMinor point: the 'ideal data' setup is chosen post hoc from the same curves used to demonstrate its success. That's not circular in a damaging way—they're identifying good data conditions, not testing a prediction—but it means the near-perfect numbers in Table 1 are a selected best case, not an independent confirmation.\n\nWho's this for? People who work on CLIP and compositional reasoning, and anyone designing data filtering or re-captioning pipelines. It deserves a serious refereeing. My own recommendation: send it out, but make the authors untangle the saliency confound—e.g., vary saliency without changing caption order or spatial position—and either temper the real-data conclusion or add a check on a larger model.","headline":"Careful synthetic study that convincingly shows data properties control CLIP binding, but the saliency claim is confounded and the real-data transfer is asserted rather than shown.","tokens_in":19040,"tokens_out":3559,"would_cite":true,"duration_ms":33847,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that CLIP's object-attribute binding failure is caused by properties of the training data—sparse captions, few multi-object images, and captioners' saliency bias—rather than by the model architecture or loss.","keywords":["CLIP","object-attribute binding","bag-of-words representation","synthetic dataset","saliency bias","contrastive learning","compositionality","data properties"],"falsifier":"Take a large web-scale corpus, automatically re-caption it so that every visible object is mentioned, the expected number of attributes per object is around 3.5, and no single object is systematically favored, then retrain a CLIP-scale model and measure binding accuracy on SugarCrepe-style swapped-caption tests. If binding accuracy does not rise well above the current near-chance level, the paper's transfer from MADMAN to real data fails.","tokens_in":17956,"feed_emoji":"🧩","tokens_out":8465,"duration_ms":94230,"temperature":0.7,"pith_summary":"This paper tries to establish that CLIP's well-known failure to bind attributes to objects—confusing “a yellow submarine and a blue bus” with “a blue submarine and a yellow bus”—is caused mainly by properties of the training data, not by the contrastive loss or the architecture. Using a fully synthetic dataset called MADMAN, the authors vary four data properties one at a time and measure a two-alternative forced-choice binding accuracy. With data properties set to values estimated from real web captions, small CLIP models stay near chance even when batch size and model size are scaled up; with an \"ideal\" data configuration, the same models reach near-perfect binding. The paper concludes that CLIP can learn object-attribute binding, but common properties of natural image-caption data—especially saliency bias, sparse captions, and few multi-object images—prevent it. This matters because it redirects the search for a fix from architecture and loss design toward data curation.","feed_headline":"Data, not architecture, explains CLIP's binding failures","feed_subtitle":"Saliency bias and sparse captions, not model size or hard negatives, decide whether CLIP binds attributes to objects.","key_machinery":"The central object is MADMAN (Multi-Attribute and Digit for Multi-Attribute biNding), a synthetic dataset built from transformed MNIST digits placed in a 3x3 grid, two per image, with template captions that chain attributes before each digit class. It makes four data properties independently controllable: the probability that an image contains two objects, the probability that a two-object image has both objects captioned, the distribution of the number of attributes per object in the caption, and the presence of a saliency bias (a center-placed object that is always captioned and mentioned first). The load-bearing evaluation is a forced-choice binding-accuracy score: for each attribute that is recognized in both objects, the authors swap that attribute between the two objects' captions and check that the model assigns higher similarity to the ground-truth caption. Realistic parameter values come from manual annotation of 100 CC12M image-caption pairs, which yields the estimates of about 95% multi-object images, 60% two-object captions, expected attributes per object of 0.57, and saliency in about 90% of images.","core_discovery":"The central claim, stated in Section 4.1, is that \"CLIP models are able to learn object-attribute binding but properties of the data inhibit this.\" Trained on a realistic MADMAN configuration—two objects in about 95% of images, both objects captioned in about 60% of two-object cases, an expected 0.57 attributes per object in captions, and saliency bias in roughly 90% of images—the model's binding accuracy sits at or near chance for all six attributes, and raising batch size to 256 or embedding size to 256 changes little. Varying the properties individually shows a steady rise in binding as multi-object images and multi-object captions become more frequent; an inverted-U curve in attributes per object, with a peak near 3.5; and a steep drop as saliency bias exceeds 0.75. Combining all optimal settings yields above 90% binding for most attributes and near 100% with a larger batch and model. Hard-negative training (NegCLIP) improves binding on realistic data but remains far below simply switching to ideal data. Out-of-distribution tests on unseen object-attribute combinations show the same trends, indicating the model learned a compositional representation rather than memorized combinations.","pith_inferences":["The authors do not test this on real data, but their account implies that re-captioning a large web corpus to mention every object, keep attribute density near 3.5 per object, and break saliency bias should produce a measurable jump in CLIP's binding accuracy on compositional benchmarks.","A corollary they leave implicit is that original human captions are not neutral descriptions: because humans preferentially caption salient objects, even perfectly matched image-caption pairs carry a bias that teaches the model to ignore binding.","Another extension: the inverse-U predicts an optimal caption length, so caption-enrichment methods should be tuned to a target density rather than made as complete as possible.","If data properties dominate, comparisons between CLIP-like models trained on different corpora are confounded by differences in these four properties, and controlling them would isolate architectural contributions."],"forward_implications":["If the central claim is right, then improving CLIP's binding on real data is mainly a data-curation problem: re-captioning or filtering web-scraped corpora to describe all objects, include a moderate number of attributes per object, and reduce saliency bias should transfer the gains seen on MADMAN.","Batch-size scaling and hard-negative mining are not the cure: the paper shows these produce at best small gains on realistic data, so future fixes should focus on data rather than only on loss design.","Saliency bias is a candidate primary cause on natural data: at realistic levels above 0.75, adding saliency to otherwise ideal data drops binding accuracy from about 90–95% to roughly 51–62%, and it diminishes the influence of the other three properties.","Because the same trends appear on held-out object-attribute combinations, the learned representations are compositional rather than memorized class labels, so the data-property effects are expected to generalize beyond the synthetic vocabulary.","The inverted-U in attribute density predicts an optimal caption length on real data; captions that are too sparse (like most web captions) or too dense (a bag-of-words can already disambiguate) both hurt binding."],"supporting_citations":[{"why":"Defines the CLIP contrastive architecture and training objective that the paper scales down to TinyCLIP and probes for binding.","marker":"[22]"},{"why":"Proposes the bag-of-words diagnosis and hard-negative (NegCLIP) remedy that this paper tests and finds insufficient.","marker":"[29]"},{"why":"Provides the binding-accuracy evaluation style and the evidence that earlier hard-negative gains were overestimated.","marker":"[13]"},{"why":"Supplies the 100 annotated image-caption pairs used to set realistic values for the four data properties.","marker":"[7]"},{"why":"Provides the MAD digit dataset whose transformed digits become the objects and attributes of MADMAN.","marker":"[25]"},{"why":"Represents the architecture-centric fix (object-centric cross-modal interaction) that the data-centric findings argue is unnecessary for binding.","marker":"[1]"},{"why":"Documents the original observation that CLIP fails object-attribute binding, which this paper sets out to explain.","marker":"[17]"}],"fun_headline_variants":["CLIP's binding failures traced to data properties, not model size","Saliency bias and sparse captions undermine CLIP object binding","Attribute density, captions, saliency: data factors that break CLIP","Hard negatives don't fix CLIP binding; ideal data does","Why CLIP can't bind attributes: common data quirks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a small CLIP trained on synthetic images behaves like a large CLIP trained on real web data, and that four data-property estimates taken from 100 manually annotated web images represent natural data; if either fails, the claim that data properties are the main cause of CLIP's binding problem may not hold for real models.","fun_headline_variants_meta":{"raw":{"variants":["CLIP's binding failures traced to data properties, not model size","Saliency bias and sparse captions undermine CLIP object binding","Attribute density, captions, saliency: data factors that break CLIP","Hard negatives don't fix CLIP binding; ideal data does","Why CLIP can't bind attributes: common data quirks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000365,"raw_usage":{"total_tokens":2014,"prompt_tokens":1044,"completion_tokens":970,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":880}},"tokens_in":660,"tokens_out":970,"duration_ms":9877,"temperature":1.0,"reasoning_tokens":880,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:27:27.137830+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a large web-scale corpus, automatically re-caption it so that every visible object is mentioned, the expected number of attributes per object is around 3.5, and no single object is systematically favored, then retrain a CLIP-scale model and measure binding accuracy on SugarCrepe-style swapped-caption tests. If binding accuracy does not rise well above the current near-chance level, the paper's transfer from MADMAN to real data fails.","supporting_citations":[{"cited_title":"In: International conference on machine learning","cited_arxiv_id":null,"evidence_quote":"Defines the CLIP contrastive architecture and training objective that the paper scales down to TinyCLIP and probes for binding."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proposes the bag-of-words diagnosis and hard-negative (NegCLIP) remedy that this paper tests and finds insufficient."},{"cited_title":"Advances in neural information processing systems36, 31096–31116 (2023)","cited_arxiv_id":null,"evidence_quote":"Provides the binding-accuracy evaluation style and the evidence that earlier hard-negative gains were overestimated."},{"cited_title":"In: CVPR (2021)","cited_arxiv_id":null,"evidence_quote":"Supplies the 100 annotated image-caption pairs used to set realistic values for the four data properties."},{"cited_title":"In: International Conference on Learning Representations (ICLR) (2025)","cited_arxiv_id":null,"evidence_quote":"Provides the MAD digit dataset whose transformed digits become the objects and attributes of MADMAN."},{"cited_title":"In: Graham, Y., Purver, M","cited_arxiv_id":null,"evidence_quote":"Documents the original observation that CLIP fails object-attribute binding, which this paper sets out to explain."}],"review_version":1}