{"id":"73739515-998b-480f-8300-911e742ce05c","arxiv_id":"2504.16801","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"DeGLA fine-tunes CLIP with LLM-generated hard negatives plus EMA self-distillation, improving compositional reasoning benchmarks by 1.9 to 4.9 points over CE-CLIP while staying within 2.3 points of the original CLIP on zero-shot classification.","lead":"This paper introduces DeGLA, a fine-tuning method that improves CLIP's understanding of compositional captions while largely retaining its general capabilities. It combines LLM-generated hard negative captions, two local contrastive losses, and an EMA self-distillation teacher to prevent catastrophic forgetting.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DeGLA's headline gain is driven by order-type benchmarks while it loses to CE-CLIP on ARO relation/attribute; the 'more effective balance' claim is over-scoped.","rationale":"The reader's conditional verdict is reasonable, but the weakest assumption identified by the reader (LLM-generated negative quality) is not the most load-bearing issue. The relation/attribute regression is directly visible in the paper's own Table 4 and acknowledged in Section 4.2, so no new experiments are needed to see that the headline average is driven by two order subtasks. The self-distillation contribution is supported by Table 7 (SD adds +1.8 zero-shot points at -0.1 ARO), and the data pipeline receives at least partial validation from the ablations in Figure 6c, so rejection is not warranted. However, the 'more effective balance' claim should be conditioned on how much weight relation/attribute errors receive; if those errors are weighted equally with order errors, DeGLA may not beat CE-CLIP. The proposed restricted-aggregate check settles this directly. The reader's negative-quality concern remains a legitimate secondary issue, hence partial agreement with the reader's weakest_assumption.","tokens_in":20663,"tokens_out":13205,"duration_ms":125119,"concrete_test":"Use the released DeGLA and CE-CLIP checkpoints (or rerun both) and compute a relation/attribute-focused aggregate: ARO Relation + ARO Attribute + VALSE spatial relations + VALSE actions + SugarCrepe REPLACE-Relation + SWAP-Attribute + ADD-Attribute, using the same averaging convention as the paper. If DeGLA's aggregate is not above CE-CLIP's, restrict the headline claim to order/permutation compositionality. Optionally add Winoground as an out-of-distribution relation/attribute check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Figure 5 rests on averaging VALSE, SugarCrepe, and ARO. That average hides a systematic weakness: on ARO, DeGLA is below CE-CLIP on Relation (81.6 vs 83.9) and Attribute (74.3 vs 76.4), and its ARO advantage comes almost entirely from COCO-order (+12.9) and Flickr-order (+11.0) in Table 4. The paper itself concedes this in Section 4.2, attributing it to negative types that emphasize word-order and noun/adjective substitution rather than relation/attribute foils (Table 1). For the 'more effective balance' claim to hold as stated, order-type compositionality must be assumed to be as important as relation/attribute compositionality, and the three benchmarks must be assumed commensurable when averaged equally. Neither assumption is defended. The paper's conclusion is therefore over-scoped: what is demonstrated is a favorable trade-off on permutation/order tasks plus zero-shot retention, not a general compositional improvement over CE-CLIP.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DeGLA, a fine-tuning framework for CLIP that aims to improve compositional understanding while mitigating catastrophic forgetting of general capabilities. The method has three components: (1) an LLM-driven pipeline (ChatGPT-4-Turbo for 200 seed examples per type, manual selection of 50, then Llama-3.1-8B-Instruct to generate about 2.07M hard negative captions from COCO 2014 train captions, covering five subtypes under intra-sentence reshuffling and minimal semantic substitution); (2) a global alignment loss (Eq. 4-5) that incorporates hard negative texts into the CLIP contrastive loss; (3) a self-distillation regularizer (Eq. 7) using an EMA teacher (image and text encoders) to preserve pretrained knowledge; and (4) two local alignment losses, Image-Grounded Contrast (IGC, Eq. 8) and Text-Grounded Contrast (TGC, Eq. 9), with the latter using the frozen EMA text embedding as anchor. The total loss is Eq. (10). Experiments are conducted on CLIP ViT-B/32, fine-tuned for 5 epochs on COCO, and evaluated on VALSE, SugarCrepe, ARO, zero-shot classification on 11 datasets, linear probe, and zero-shot retrieval on MSCOCO/Flickr30k. The headline claims are an average +3.5% over CE-CLIP on VALSE/SugarCrepe/ARO and +13.0% over CE-CLIP on zero-shot classification, leading to the stated 'more effective balance' between compositional and general capabilities.","tokens_in":20837,"tokens_out":4873,"duration_ms":35855,"significance":"If the reported results hold, the paper would make a useful contribution to the line of work on hard-negative fine-tuning of CLIP: the decoupled global-local formulation with an EMA self-distillation teacher is a plausible mechanism for retaining pretrained zero-shot capability while still gaining on compositionality benchmarks. The LLM-driven negative generation pipeline with five explicit subtypes is straightforward and applicable to other models, and the paper provides a clear ablative breakdown (Table 7) attributing the general-capability retention mainly to the distillation term. The paper also ships code and includes detailed prompt templates in the appendix. The significance is moderate: the gains are on established but narrow benchmarks, the method is incremental over CE-CLIP (which already used ranking/image-grounded hard negatives and intra-modal hard negatives), and the claim of an 'optimal balance' rests on averages that mask a regression on ARO relation/attribute subsets.","major_comments":[{"comment":"The central claim that DeGLA achieves 'a more effective balance between compositional reasoning and general comprehension capabilities' is over-scoped. The +3.5% average over VALSE, SugarCrepe, and ARO hides the fact that DeGLA is below CE-CLIP on ARO Relation (81.6 vs 83.9) and ARO Attribute (74.3 vs 76.4) in Table 4, and its ARO lead comes almost entirely from the Order subsets. The paper itself acknowledges this in Section 4.2 ('it still trails CE-CLIP and Structure-CLIP in the domains of relations and attributes'), but Figure 5 and the abstract are still phrased as an unconditional improvement. For the balance claim to be supported, the authors should either restrict the claim to order-type compositionality plus zero-shot retention, or provide an explicit justification for treating the three benchmarks as commensurable dimensions of a single 'compositional reasoning' average.","section":"Abstract, Figure 5, Section 4.4"},{"comment":"The assumption that the Llama-3.1-8B generated captions are genuinely hard negatives is not validated. The paper asserts semantic divergence via prompt constraints and manual selection of 50 ChatGPT exemplars, but no automatic or human evaluation is reported for the 2.07M generated negatives. Inspection of Figure 10 shows that several generated items are not hard negatives: e.g., the 'subtype 2' example 'Two young children playing around a fire hydrant' → 'Fire hydrant around playing a young two children' is a non-grammatical word salad that the model will trivially reject, while 'A gray cat is sitting on a wooden bench' → 'A goldfish is sitting on a concrete road' changes scene composition completely and is an easy negative rather than a hard one. The paper should quantify the difficulty and validity of the generated negatives, for example by reporting CLIP similarity distributions between positive and generated captions before and after post-filtering, and by reporting human-annotation agreement on a subsample. Without this, the contribution of the data-generation pipeline relative to rule-based or unmasking-based baselines is not cleanly established.","section":"Section 3.2 and Figure 10"},{"comment":"The hyperparameters lambda_1, lambda_2, lambda_3 and EMA alpha are tuned without a clearly separated validation protocol, and no error bars or multiple-seed results are reported anywhere in the paper. The ablation of lambda_1/lambda_2 in Figure 6a shows a rather flat landscape with an apparent optimum at (0.1, 0.1), but the selection procedure over these values (and its potential leakage into the benchmark average) is not described. Given that the headline 3.5%/13.0% improvements are point estimates over a single run, the authors should report at least two or three seeds with means and standard deviations for the main tables, and describe how the validation set was chosen for hyperparameter selection (e.g., a held-out subset of ARO or a separate validation benchmark).","section":"Section 4.1, Table 8"},{"comment":"The component ablation in Table 7 is reported only on ARO and zero-shot classification, and the row ordering makes the individual contribution of IGC ambiguous. The row 'CN + IGC' (84.8 ARO, 55.1 ZS) compared with 'CN' alone (80.8, 55.4) and 'CN + TGC' (82.2, 57.2) shows IGC and TGC each help ARO, but the subsequent 'CN + IGC + TGC' row (86.2, 56.9) shows a much larger ARO gain (+5.4 over CN+TGC) than the sum of individual gains, which is not discussed. More importantly, the table does not isolate the effect of the EMA self-distillation from the effect of the frozen teacher embeddings used inside TGC (Eq. 9); since TGC already uses the frozen EMA text embedding as its anchor, the 'SD' row that adds L_Distill may be partly redundant with the stabilizing effect already present in TGC. The ablation as presented therefore does not cleanly attribute general-capability retention to the distillation term.","section":"Table 7 and Section 4.4"},{"comment":"The connection between the five generated subtypes and the training-time pairing K=4 is under-specified. The paper states that K=4 'due to the merging of two subtypes, as detailed in the appendix', but the appendix does not actually explain which subtypes were merged, how the four negatives per text were sampled across the five subtypes, whether the distribution over subtypes was balanced, or whether each subtype was paired with its corresponding positive text in a controlled way. This makes it difficult to reproduce the training data construction exactly.","section":"Appendix A.4 and Section 4.1"}],"minor_comments":[{"comment":"There are inconsistencies in acronym usage: the Introduction and Figure 3 captions refer to 'ICC' and 'TCC', while the Method and experiments use 'IGC' and 'TGC'; the abstract and conclusion also give the variant spellings. The paper should standardize on one set of acronyms.","section":"Abstract and Section 3.3"},{"comment":"Numerous typos appear in the text, table headers, and appendix. These include 'liner probe' (Tables 6 and Section 4.3), 'VASLE' in the appendix A.3, 'Composionablity' in Table 11 header, 'Detatils' in A.2, a duplicate reference (Oord et al. appears as [45] and [46]), 'Preformance' in Figure 5 axis, and the unmatched parenthesis in Eq. (8). A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The implementation paragraph lists parameters 'beta_1 and beta_2' as 0.9 and 0.98, which is standard for AdamW, but the supplementary Table 8 does not list beta_1/beta_2; adding them to the table would make the configuration complete.","section":"Section 4.1"},{"comment":"In Table 3 the 'Avg.' column for the Replace/Relation group appears to be computed over the three replacement types while the table layout suggests a different grouping; in Table 4, CLIP's row shows Relation 59.2 and Attribute 62.9, but the average 57.4 below it is not the average of the visible four columns (it appears to be the average of the four ARO sub-tasks, which would be 57.65). The definition of the reported average should be stated in the captions.","section":"Tables 3 and 4"},{"comment":"The paper refers to 'ChatGPT4-Turbo' and 'ChatGPT-4-Turbo' with inconsistent hyphenation and no version/date information; specifying the exact model snapshot and sampling parameters would help reproducibility, especially since generation prompts appear in the appendix but the decoding temperature and number of samples are not reported.","section":"Section 3.2 and Figure 2"},{"comment":"The first column of Figure 10 is labeled 'Positive caption' and the second 'Hard negative examples', but the fourth row shows the original-positive caption repeated as a hard negative ('A red bus going down the road behind a blue bus' appears both as the input and as one of the generated examples). This suggests the filtering may not have removed near-duplicate outputs; the authors should clarify whether such cases were actually included in the training set and how they were handled.","section":"Appendix Figure 10"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for ACM MM and the core idea is reasonable, but the main claims are stated more strongly than the evidence supports: the ARO relation/attribute regression and the unvalidated 2M generated negatives are load-bearing. The lack of multi-seed reporting and a defined validation protocol for lambda selection are standard expectations for a paper that stakes a 'balanced trade-off' claim. I would not reject, but the authors need to substantially revise the claims and add the missing analyses. There is also a minor novelty-diligence point: CE-CLIP already contains ranking cross-modal and intra-modal losses, and the present contribution is mostly the EMA distillation term plus the LLM data pipeline, so the positioning sentences should be tightened. No concerns about misconduct; the paper is transparent about its release plan."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. DeGLA is a genuinely useful empirical result, with a caveat the authors mostly own up to. The new combination is EMA self-distillation plus two local contrastive losses (IGC and TGC) on top of LLM-generated hard negatives. It works: they get most of CE-CLIP's compositional improvement while keeping zero-shot classification at 58.7 vs CLIP's 61.0, where CE-CLIP drops to 45.7. The 13-point average gap over CE-CLIP on 11 zero-shot datasets is large and directionally consistent. The ablations in Table 7 support the role of each component; this is not a black box. And the paper is transparent in Section 4.2 that DeGLA trails CE-CLIP on ARO Relation (81.6 vs 83.9) and Attribute (74.3 vs 76.4).\n\nThe soft spots are real but proportionate. The headline claim of a 'more effective balance' is over-scoped, and the stress-test note is right: the compositional average is carried by order-type benchmarks. On ARO, DeGLA's entire advantage over CE-CLIP comes from COCO-order and Flickr-order. If relation and attribute understanding matter as much as word order, the trade-off is less one-sided. The paper should say 'order-type compositionality' where it says 'compositional understanding'. Second, there are no error bars or multiple seeds. Hyperparameters lambda1-3 were tuned without a clearly separated validation protocol, so the precise numbers could shift. Third, the 2.07M generated negatives were never checked for hardness or semantic divergence; the prompt instruction is plausible but unmeasured. If a chunk of these are hard positives, the local losses would be injecting noise. Fourth, the benchmarks overlap with the COCO fine-tuning data, so out-of-distribution gains are unknown—though that is equally true of the prior methods they compare against.\n\nWho is this for? Practitioners fine-tuning CLIP for compositional tasks who care about general zero-shot retention. The method is simple enough to adopt. For a paper at this venue, the evidence is sufficient to send to referees. I'd ask for code and data release, a sharper claim about which subtasks improve, and at least a couple of seeds or error bars before accepting. The central mechanism is credible, and the limitation is acknowledged rather than hidden.","headline":"DeGLA shows a real trade-off improvement for CLIP fine-tuning, but the compositional gain is mostly order-type; worth a serious referee.","tokens_in":21412,"tokens_out":2872,"would_cite":true,"duration_ms":25803,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DeGLA beats hard-negative baselines and keeps CLIP's old skills","keywords":["vision-language models","compositional understanding","CLIP","hard negatives","self-distillation","contrastive learning","zero-shot classification","LLM-generated captions"],"falsifier":"A random sample of 200 generated negative captions could be labeled by human annotators as hard negative, hard positive (same scene and meaning, only wording changed), or easy negative (semantics changed beyond one minimal edit); if hard positives or easy negatives make up a large share, the data premise of DeGLA is falsified. A cleaner experiment would rerun DeGLA with an automatic filter that keeps only captions whose frozen-CLIP similarity to the source falls inside a hard-negative band, then check whether the compositional gains shrink, stay, or grow.","tokens_in":20444,"feed_emoji":"🧩","tokens_out":10996,"duration_ms":96365,"temperature":0.7,"pith_summary":"Contrastive vision-language models compare whole images with whole captions, which leaves them blind to compositional structure: they cannot reliably tell the dog chases the cat from the cat chases the dog. This paper claims that fine-tuning on hard negatives can fix that blindness without the usual catastrophic forgetting, if global alignment is decoupled from local alignment and anchored to the pretrained model. The proposed DeGLA framework keeps a frozen exponential-moving-average teacher copy of the CLIP encoders, adds image-grounded and text-grounded contrast losses over LLM-generated hard negative captions, and reports an average gain of 3.5% over CE-CLIP on VALSE, SugarCrepe, and ARO while raising zero-shot classification by 13.0% on average across eleven datasets. The central claim is that hard-negative fine-tuning and knowledge retention are not opposing goals when the loss structure keeps the student close to the pretrained representation space.","feed_headline":"DeGLA beats hard-negative baselines and keeps CLIP's old skills","feed_subtitle":"Lifts composition benchmarks 3.5% and zero-shot accuracy 13.0% over CE-CLIP while keeping general skills.","key_machinery":"The load-bearing object is the decoupling itself: a global alignment channel, the standard image-to-text and text-to-image InfoNCE losses with four hard negative captions per text, is stabilized by an EMA self-distillation loss, while a separate local alignment channel sharpens compositionality through the Image-Grounded Contrast (IGC) loss and the Text-Grounded Contrast (TGC) loss. IGC uses the image as anchor and contrasts the positive caption against K negatives; TGC uses the frozen EMA text embedding as a stable positive anchor within the text modality, preventing the text encoder from drifting into a narrow fine-tuning space. The data feeding both channels comes from an LLM-driven negative-caption pipeline with five rewrite rules, whole-sentence reshuffle, noun swap, adjective swap, adjective replacement, and noun replacement, scaled to five negatives per COCO caption for about 2.07 million negatives in total. The machinery lets the model learn compositionality from minimal textual edits while the EMA anchor preserves the pretrained representation geometry.","core_discovery":"The paper's central discovery is that the trade-off previous hard-negative methods accept, compose better but forget more, is not forced. DeGLA keeps the global CLIP contrastive objective but augments it with hard negatives and with a self-distillation term that penalizes the squared distance between learnable and frozen EMA teacher embeddings of images, texts, and negative texts. On the local side, an Image-Grounded Contrast loss pulls the image embedding toward the correct caption and away from the generated negatives, while a Text-Grounded Contrast loss uses the frozen EMA text embedding as the positive anchor so the text encoder does not overfit to the fine-tuning distribution. The paper reports that this decoupling yields an average 3.5% compositional gain over CE-CLIP on VALSE, SugarCrepe, and ARO, and a 13.0% average zero-shot classification improvement across eleven datasets, with general capability largely retained: DeGLA averages 58.7% zero-shot accuracy versus CLIP's 61.0% and CE-CLIP's 45.7%.","pith_inferences":["The same EMA-anchored global-local split could apply to other narrow fine-tuning regimes, such as debiasing, domain adaptation, or retrieval distillation, since nothing in the loss is specific to compositionality.","The data premise could be tightened with an automatic hard-negative filter, for example keeping generated captions whose frozen-CLIP similarity to the source falls inside a hard band; the paper validates only 50 hand-picked seed examples per rewrite rule.","A relation-focused negative subtype that swaps grammatical roles rather than replacing or swapping nouns and adjectives is the obvious complement, because the ARO relation scores show where DeGLA leaves points on the table.","The zero-shot table also shows that retention is partial: DeGLA at 58.7% average remains 2.3 points below the original CLIP's 61.0%, so the honest framing is much less forgetting rather than no forgetting."],"forward_implications":["The paper's ablations show that each component pays: the LLM negatives add 23.4 points on ARO over CLIP, IGC adds 4.0, TGC adds 1.4, and the self-distillation term recovers 1.8 points of zero-shot accuracy while costing only 0.1 on ARO.","If the reported averages hold, DeGLA becomes the reference point for compositional fine-tuning of CLIP-sized models, beating CE-CLIP on every SugarCrepe group and on both ARO order tasks.","The remaining gap is explicit in the tables: DeGLA's ARO relation and attribute scores trail CE-CLIP, which the paper attributes to its negative-generation diversity being broader rather than relation- and attribute-focused.","Because the EMA anchor is a single squared-distance term, the mechanism can be dropped into existing hard-negative fine-tuning recipes without changing the pretraining or evaluation protocol."],"supporting_citations":[{"why":"Supplies the pretrained CLIP encoders that DeGLA fine-tunes and the zero-shot classification baseline that measures general-capability retention.","marker":"[52]"},{"why":"Defines the ARO benchmark and the hard-negative fine-tuning setup that DeGLA extends, providing the base global contrast loss.","marker":"[63]"},{"why":"CE-CLIP is the strongest prior compositional fine-tuning method and the main baseline behind the paper's 3.5% and 13.0% improvement claims.","marker":"[68]"},{"why":"VALSE provides one of the three compositional evaluation suites used in the headline average gain.","marker":"[47]"},{"why":"SugarCrepe provides the Replace, Swap, and Add evaluation tasks where DeGLA reports its largest compositional improvements.","marker":"[20]"},{"why":"Structure-CLIP is the compared hard-negative method whose relation and attribute scores define the gap DeGLA discusses.","marker":"[22]"},{"why":"The instruction-following LLM used to scale up the five types of hard negative captions to about 2.07 million training examples.","marker":"[13]"}],"fun_headline_variants":["DeGLA decouples alignment to lift composition 3.5% without skill loss","Composition up, skills kept: DeGLA's decoupled global-local trick","Zero-shot +13%, composition +3.5% — DeGLA beats trade-off","DeGLA: harder negatives, softer forgetting, better vision-language","Keep CLIP's wisdom, add composition: DeGLA's decoupled design"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the roughly 2.07 million LLM-generated captions are actually hard negatives, minimally altered sentences with genuinely different semantics, rather than easy negatives or hard positives; the paper hand-selects 50 seed examples per rewrite rule but does not evaluate the full generated set.","fun_headline_variants_meta":{"raw":{"variants":["DeGLA decouples alignment to lift composition 3.5% without skill loss","Composition up, skills kept: DeGLA's decoupled global-local trick","Zero-shot +13%, composition +3.5% — DeGLA beats trade-off","DeGLA: harder negatives, softer forgetting, better vision-language","Keep CLIP's wisdom, add composition: DeGLA's decoupled design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000326,"raw_usage":{"total_tokens":1893,"prompt_tokens":1084,"completion_tokens":809,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":700,"completion_tokens_details":{"reasoning_tokens":702}},"tokens_in":700,"tokens_out":809,"duration_ms":6782,"temperature":1.0,"reasoning_tokens":702,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:54:43.737032+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A random sample of 200 generated negative captions could be labeled by human annotators as hard negative, hard positive (same scene and meaning, only wording changed), or easy negative (semantics changed beyond one minimal edit); if hard positives or easy negatives make up a large share, the data premise of DeGLA is falsified. A cleaner experiment would rerun DeGLA with an automatic filter that keeps only captions whose frozen-CLIP similarity to the source falls inside a hard-negative band, then check whether the compositional gains shrink, stay, or grow.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SugarCrepe provides the Replace, Swap, and Add evaluation tasks where DeGLA reports its largest compositional improvements."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Structure-CLIP is the compared hard-negative method whose relation and attribute scores define the gap DeGLA discusses."}],"review_version":1}