{"id":"755f0e56-24d8-4b77-98fd-5d2c26aef305","arxiv_id":"2506.21237","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"DiMPLe separates CLIP's image and text features into invariant and spurious components and claims large OOD accuracy gains through cross-modal alignment of the invariant parts.","lead":"A new prompt-tuning method, DiMPLe, splits CLIP's image and text features into stable and spurious parts, aligning only the stable parts across modalities. The paper reports large gains over its CoOp-OOD baseline on 11 datasets, but the baseline scores are implausibly low and no code is released.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline gains are measured against a CoOp-OOD baseline that is far below expected performance; the paper's own controlled CoOp-OOD† baseline shrinks the novel-class gain from 44.31 to 7.75, so the central claim does not hold as stated.","rationale":"The Pith reader's weakest_assumption identifies the fidelity of the CoOp-OOD baseline, and I agree that this is the load-bearing point. The paper's entire quantitative argument is the 15.27/44.31 gain in Table 3a and the DG table showing superiority over CoOp-OOD. Without a faithful baseline, those numbers are not evidence about the real method. The paper's own Table 3a provides a nearly ideal control: CoOp-OOD†, which uses the same losses as CoOp-OOD and the same conditioned vision-language prompt architecture as DiMPLe, differing only in the multi-modal disentanglement. Against CoOp-OOD†, DiMPLe's gains shrink to 13.37 base and 7.75 novel, with DiMPLe still below CLIP on average novel accuracy (73.35 vs 74.22). The huge gap between CoOp-OOD (29.04 novel) and CoOp-OOD† (65.60 novel) under identical losses indicates an implementation problem in the baseline rather than a property of the method. The paper also lacks code, seeds, and error bars, and includes no state-of-the-art prompt-tuning baselines such as Co-CoOp, MaPLe, or CLIP-OOD. These are additional weaknesses, but the baseline issue alone is sufficient to reject the central claim. I therefore keep the reader's REJECT verdict unchanged.","tokens_in":15310,"tokens_out":6480,"duration_ms":69439,"concrete_test":"Obtain the official CoOp-OOD implementation from the authors of [29] (or an independently verified reimplementation) and run it under the paper's exact protocol: 16 shots per class, ViT-B/16, 5 epochs, batch size 4, on the 11 base-to-novel datasets. If the official baseline's average novel accuracy exceeds the reported 29.04 by more than about 10 points, then the headline gains of 15.27/44.31 are not established; Table 3a should be recomputed against the faithful baseline. A secondary confirmatory check: use the paper's own CoOp-OOD† row as the reference and report the resulting 13.37 base and 7.75 novel gains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that DiMPLe yields absolute gains of 15.27 base and 44.31 novel over CoOp-OOD (Table 3a). This claim is load-bearing on the fidelity of the CoOp-OOD baseline, and the paper's own numbers show that baseline is not the published method. Table 2 reports CoOp-OOD at 38.7% on ImageNet vs CLIP's 66.73%, and Table 3a reports an average novel accuracy of 29.04% vs CLIP's 74.22%. The paper's own CoOp-OOD variant CoOp-OOD†, which uses the same loss functions as CoOp-OOD but adds conditioned vision-language prompting, achieves 65.60% average novel accuracy; adding vision prompts alone cannot credibly produce a 36.56-point jump if the base implementation were faithful. The controlled comparison is therefore CoOp-OOD†, and against that baseline DiMPLe's gains are 13.37 base and 7.75 novel, not 15.27 and 44.31. This means the headline gain conflates the contribution of MaPLe-style conditioned vision-language prompting with the proposed multi-modal disentanglement. No official code, seeds, or error bars are provided, so the weak baseline cannot be checked; the Table 1 CoOp-OOD* result of 25.99 base with independent vision prompts, against DiMPLe* at 76.73, is another red flag. The appropriate conclusion is that the paper does not establish superiority over a correctly implemented CoOp-OOD.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DiMPLe, a prompt-tuning method for CLIP that jointly disentangles invariant and spurious features in both vision and language branches. The method combines three objectives: conditional mutual information minimization between invariant and spurious features, spurious-feature regularization, and a contrastive alignment loss on invariant features, along with deep vision-language prompting in which vision prompts are conditioned on language prompts. The authors report experiments on base-to-novel generalization over 11 datasets, domain generalization on four ImageNet variants, and cross-dataset transfer, comparing against CLIP, CoOp-OOD, and two self-created variants of CoOp-OOD.","tokens_in":15601,"tokens_out":5349,"duration_ms":63191,"significance":"The idea of extending feature disentanglement to the text modality is interesting and the objective in Eq. (10) is clearly stated. The paper also provides ablations on loss components, prompt depth, token count, and a CelebA worst-group analysis, which are useful. However, the empirical validation is not reliable: the CoOp-OOD baseline appears to be a severely weakened reimplementation, the paper's own controlled baseline (CoOp-OOD†) shrinks the headline novel-class gain from 44.31 to 7.75 points, and no comparisons are made against standard prompt-tuning methods such as CoOp, Co-CoOp, MaPLe, or CLIP-OOD. The absence of error bars and code further weakens the central empirical claims. If the method were evaluated properly against strong baselines and the claims were revised accordingly, the contribution could be of interest, but as presented the headline results are not supported.","major_comments":[{"comment":"The CoOp-OOD baseline is not a faithful implementation of the published method. On ImageNet it achieves only 38.7% accuracy versus 66.73% for CLIP (Table 2), and its average novel-class accuracy is 29.04% versus 74.22% for CLIP (Table 3a). Published CoOp-OOD results are substantially higher, and the paper's own CoOp-OOD† variant, which merely adds conditioned vision prompts to the same loss functions, jumps to 65.60% average novel accuracy. Adding vision prompts cannot plausibly yield a 36.56-point improvement if the base implementation were correct. Therefore the headline gains of 15.27 base and 44.31 novel reported in the abstract and §4.1 are computed against a broken baseline; against the paper's own controlled baseline CoOp-OOD†, the gains are 13.37 base and 7.75 novel (Table 3a).","section":"Table 2 / Table 3a"},{"comment":"The comparison between CoOp-OOD* and DiMPLe* in Table 1 is a further red flag. With identical independent vision-language prompting, DiMPLe* improves base accuracy from 25.99% to 76.73% and novel accuracy from 29.98% to 72.36% solely by adding the disentanglement losses. A 50-point base-accuracy jump cannot credibly be attributed to the proposed losses, confirming that the CoOp-OOD implementation is undertrained or missing essential components. This makes the central empirical claim of the paper untrustworthy.","section":"Table 1"},{"comment":"The paper never compares against actual prompt-tuning baselines such as CoOp, Co-CoOp, MaPLe, or CLIP-OOD, even though the deep prompting backbone in Eqs. (4)-(6) is directly based on MaPLe. Without these comparisons, it is impossible to determine whether the reported improvements come from the disentanglement losses or from adopting a stronger prompting architecture. The closest controlled comparison, CoOp-OOD†, shows only a 7.75-point novel-class gain for DiMPLe, which is not sufficient to establish the contribution of the proposed disentanglement mechanism.","section":"§4, Table 3"},{"comment":"Equation (3) defines Lcmi_v as I(zi,s; zi,s|Y), which is identically zero by the properties of mutual information; the intended quantity is I(zi,u; zi,s|Y). Additionally, the practical estimator is not described: the paper cites [14] for HSIC but does not explain how the conditional mutual information is computed, what kernels are used, or how the class variable enters the estimator. This makes the core objective non-reproducible.","section":"§3.1, Eq. (3)"},{"comment":"The paper states that results are averaged over 3 runs but provides no standard deviations or confidence intervals. Several reported differences are small (e.g., 61.2 vs. 60.83 on ImageNetV2 in Table 2), so without variance estimates these differences are not shown to be statistically meaningful. The paper also states that code will be released upon acceptance, but no code or seeds are provided, which precludes verification of the unusual baseline numbers and the claimed gains.","section":"Implementation Details, §4"}],"minor_comments":[{"comment":"The sentence 'Extensive experiments demonstrate DiMPLe demonstrates superior performance' contains a repeated verb; it should read 'Extensive experiments demonstrate that DiMPLe achieves superior performance...'.","section":"Abstract"},{"comment":"The model name is spelled 'DiMPle' in the Caltech101 panel; it should be 'DiMPLe' for consistency.","section":"Table 3(c)"},{"comment":"The capitalization of 'CoOp-OOD' is inconsistent (e.g., 'CoOP-OOD' in Table 1 versus 'CoOp-OOD' in the text).","section":"§4 and Table 1"},{"comment":"The cross-dataset evaluation refers to Fig. 3 for specific numbers such as OxfordPets 85.67%, but Fig. 3 is a radar plot and the underlying quantitative table is not provided, making these numbers difficult to verify.","section":"§4.3"},{"comment":"The equation is malformed: 'exp(sim(ztest, wc)/τ )PC' should be a fraction, presumably exp(sim(ztest, wc)/τ) divided by the sum over class similarities.","section":"Supplementary Eq. (11)"}],"recommendation":"reject","confidential_remarks":"The paper's experimental claims appear to rely on a severely undertrained CoOp-OOD baseline. The paper's own controlled baseline, CoOp-OOD†, reduces the novel-class gain from 44.31 to 7.75 points, and the 50-point jump between CoOp-OOD* and DiMPLe* in Table 1 is not credible. The absence of comparisons to MaPLe and other standard baselines, along with missing error bars and code, makes the central contribution unverifiable. A revision would require re-running all experiments with a correct CoOp-OOD implementation, adding standard baselines and variance estimates, and substantially revising the claims; this is beyond a local fix. I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The central claim is not supported by the evidence as presented. The advertised gains of 15.27 in base and 44.31 in novel accuracy over CoOp-OOD shrink to 13.37 and 7.75 when you compare against CoOp-OOD†, the paper's own variant with conditioned vision-language prompting. That is the right control, because the raw CoOp-OOD baseline scores 38.7 on ImageNet and 29.04 average novel accuracy—numbers far below published CoOp-OOD and even below CLIP zero-shot. That baseline looks broken, and the headline gains are largely an artifact of that breakage.\n\nWhat the paper does well: the idea of disentangling spurious text features and explicitly aligning them with spurious visual features is a sensible extension of CoOp-OOD's image-only disentanglement. The objective is clearly specified—conditional HSIC for mutual information minimization, KL regularization for spurious features, and contrastive alignment for invariant features—and they include ablations on loss components, prompt depth, token count, and early versus late disentanglement. The computational cost comparison between DiMPLe and DiMPLe-E is a useful practical detail.\n\nThe soft spot is load-bearing. There is no comparison against the published CoOp-OOD numbers, no MaPLe, Co-CoOp, or CLIP-OOD baselines, no code, no seeds, no error bars. The paper's own Table 1 shows CoOp-OOD* with independent vision-language prompting at 25.99 base accuracy—a number so low that it suggests the base implementation is missing something essential. Adding conditioned vision-language prompting alone (CoOp-OOD†) jumps novel accuracy from 29.04 to 65.60, an implausibly large effect for that architectural change if the base were faithful. Against the proper control, DiMPLe still improves harmonic mean by about 10 points on average, so the method is not empty, but the novel-class advantage is modest and the domain-generalization numbers are essentially tied with CLIP.\n\nThis paper is for a reader interested in prompt-tuning failure modes as a cautionary tale, but as evidence that DiMPLe improves OOD generalization it does not convince. The method may have merit, but it needs a rebuilt baseline evaluation, error bars, and real SOTA comparisons. I would not cite it in its current form and would desk-reject this version, though I would invite a resubmission with a properly implemented CoOp-OOD and a broader set of baselines.","headline":"Headline gains are an artifact of a broken CoOp-OOD baseline; the real improvement over the proper control is far smaller.","tokens_in":16162,"tokens_out":2728,"would_cite":false,"duration_ms":29096,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DiMPLe claims that out-of-distribution generalization in prompt-tuned CLIP requires disentangling spurious from invariant features in the text branch as well as the image branch.","keywords":["prompt learning","CLIP","out-of-distribution generalization","spurious correlation","feature disentanglement","mutual information","vision-language models","few-shot classification"],"falsifier":"Re-run CoOp-OOD with the official code under the paper's exact protocol (ViT-B/16, 16 shots, 5 epochs, learning rate 0.0035). If ImageNet base accuracy returns to the published range near 66 percent and average novel accuracy exceeds 60 percent, rather than the 38.7 and 29.04 reported here, the headline advantage is a baseline artifact.","tokens_in":15046,"feed_emoji":"🧩","tokens_out":6153,"duration_ms":62564,"temperature":0.7,"pith_summary":"This paper claims that out-of-distribution failures in prompt-tuned CLIP come from a source previous methods missed: spurious features in the text branch, and the ambiguous way they align with image features. The proposed method, DiMPLe, splits features in both modalities into invariant and spurious parts, minimizes the mutual information between the two parts within each modality, and aligns only the invariant parts across modalities. Across 11 datasets, the authors report average gains of 15.27 points on base classes and 44.31 points on novel classes over their CoOp-OOD baseline, with a harmonic-mean accuracy of 74.70. The result matters because it suggests that explicit cross-modal disentanglement, not just visual disentanglement, is what allows prompt-tuned models to generalize to unseen categories and distribution shifts.","feed_headline":"DiMPLe reports 44-point novel-class gain over CoOp-OOD","feed_subtitle":"Splitting spurious and invariant cues in both text and image lifts novel-class accuracy across 11 datasets.","key_machinery":"The load-bearing object is the disentangled multi-modal feature decomposition: image features $z_i^v$ and text features $z_i^t$ are each projected by separate linear heads into an invariant part ($z_{i,u}^v$, $z_{i,u}^t$) and a spurious part ($z_{i,s}^v$, $z_{i,s}^t$), and the training loss $L = L_{ce}^u + \\alpha L_{sp}^r + \\beta L_{cmi}$ enforces that the two parts carry separate information and that only invariant features drive classification. The conditional-mutual-information term $L_{cmi} = I(z_{i,u}; z_{i,s} \\mid Y)$ is computed per modality using a Hilbert-Schmidt independence criterion estimator, which is what gives the word 'disentangled' its operational meaning here. The coupling of vision prompts to language prompts through a learned linear map is what makes the disentanglement multi-modal rather than two independent single-modal decouplings.","core_discovery":"The central claim is that disentangling invariant from spurious features in both the vision and language streams, and mapping each to its counterpart in the other modality, removes the cross-modal ambiguity that makes image-only disentanglement (CoOp-OOD) collapse on novel classes. DiMPLe does this with three objectives: a conditional-mutual-information penalty, estimated with HSIC, that makes invariant and spurious features independent given the class label within each modality; a spurious-feature regularization that drives predictions based on spurious features toward a uniform distribution; and a cross-entropy alignment that matches only invariant image features to invariant text features. These are combined with deep multi-modal prompting in which vision prompts are generated as a linear projection of language prompts, following the MaPLe design. On the paper's reported numbers, the method reaches an average harmonic mean of 74.70 across 11 base-to-novel benchmarks, stays within about one point of zero-shot CLIP on four of five ImageNet shift datasets, and lifts CelebA worst-group accuracy from 31.11 to 70.0 without group labels.","pith_inferences":["The headline gains are relative to the paper's own CoOp-OOD numbers, and those numbers (ImageNet source accuracy 38.7, average novel accuracy 29.04) are far below published CoOp-OOD results; a faithful re-implementation would likely close most of the 44.31-point gap.","The HSIC-based conditional mutual information estimator is used with a batch size of 4, which is small for kernel-based independence tests; a testable extension is to check whether the three losses remain balanced at larger batch sizes or under different HSIC kernel widths.","The same cross-modal spurious alignment could be applied to other large vision-language models or to open-vocabulary detection, where the paper's central mechanism of aligning spurious text components with spurious image components is not specific to image classification."],"forward_implications":["If the reported numbers hold, a prompt-tuned CLIP can exceed zero-shot CLIP on the source domain (69.73 vs 66.73 on ImageNet) while staying competitive on ImageNetV2, ImageNet-Sketch, and ImageNet-R.","The 44.31-point average novel-class gain implies the disentangled invariant features transfer to classes never seen during training, which is the main failure mode CoOp-OOD was designed to fix.","The CelebA result suggests the method reduces reliance on spurious attributes such as gender for hair-color prediction without needing group annotations, improving worst-group accuracy by 38.9 points.","Because only 3.28 percent of parameters are trainable and training runs just five epochs, the claimed gains are attributed to the loss design rather than model scale or long fine-tuning."],"supporting_citations":[{"why":"CoOp-OOD, the image-only disentanglement baseline that DiMPLe extends and the reference against which the headline gains are measured.","marker":"[29]"},{"why":"CLIP, the frozen vision-language backbone whose encoders all prompt-tuning methods, including DiMPLe, adapt.","marker":"[21]"},{"why":"MaPLe, the source of deep multi-modal prompting and of the vision-prompt coupling through a linear projection.","marker":"[15]"},{"why":"Kalinke and Szabó, whose HSIC estimator implements the conditional mutual information penalty in the loss.","marker":"[14]"},{"why":"CoOp, the learnable-prompt predecessor whose few-shot protocol and prompt initialization the experiments follow.","marker":"[31]"}],"fun_headline_variants":["DiMPLe: disentangling spurious and invariant cues nets 44-point novelty gain","Spurious-invariant disentanglement across text and vision yields 44-point novel-class gain","Prompt-based disentanglement adds 44 points to novel-class accuracy","DiMPLe's three-way objective boosts OOD novel classes by 44 points","Separating spurious and invariant features in prompts yields 44-point OOD gain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claimed advantages are measured against the paper's own implementation of CoOp-OOD; if that implementation is weaker than the published method, the absolute gains of 15.27 and 44.31 points do not hold.","fun_headline_variants_meta":{"raw":{"variants":["DiMPLe: disentangling spurious and invariant cues nets 44-point novelty gain","Spurious-invariant disentanglement across text and vision yields 44-point novel-class gain","Prompt-based disentanglement adds 44 points to novel-class accuracy","DiMPLe's three-way objective boosts OOD novel classes by 44 points","Separating spurious and invariant features in prompts yields 44-point OOD gain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001336,"raw_usage":{"total_tokens":5425,"prompt_tokens":930,"completion_tokens":4495,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":4391}},"tokens_in":546,"tokens_out":4495,"duration_ms":38508,"temperature":1.0,"reasoning_tokens":4391,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:29:22.442492+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run CoOp-OOD with the official code under the paper's exact protocol (ViT-B/16, 16 shots, 5 epochs, learning rate 0.0035). If ImageNet base accuracy returns to the published range near 66 percent and average novel accuracy exceeds 60 percent, rather than the 38.7 and 29.04 reported here, the headline advantage is a baseline artifact.","supporting_citations":[{"cited_title":"Amend to alignment: De- coupled prompt tuning for mitigating spurious correlation in vision-language models","cited_arxiv_id":null,"evidence_quote":"CoOp-OOD, the image-only disentanglement baseline that DiMPLe extends and the reference against which the headline gains are measured."},{"cited_title":"Nystr ¨om m-hilbert- schmidt independence criterion, 2023","cited_arxiv_id":null,"evidence_quote":"Kalinke and Szabó, whose HSIC estimator implements the conditional mutual information penalty in the loss."}],"review_version":1}