{"id":"d77e5cf3-2d28-4cdf-b189-9d1afd75ec87","arxiv_id":"2505.17796","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A dual-branch CIR framework that pre-trains a detail-focused branch on InstructPix2Pix editing data, fuses global and detail features with an adaptive compositor, and reports state-of-the-art on CIRR and FashionIQ.","lead":"DetailFusion adds a detail-oriented branch to composed image retrieval, trained on an image editing dataset so the model notices small visual changes and atomic text edits. It reports state-of-the-art results on the CIRR and FashionIQ benchmarks at a modest extra compute cost.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stage 1 checkpoint is selected using CIRR and FashionIQ validation sets (Sec. 4.1); if this target-benchmark selection is removed, the reported SOTA margins may shrink, making the central comparison unfair.","rationale":"The reader's weakest assumption about transfer of detail priors is plausible, but it does not directly threaten the validity of the comparison. The checkpoint-selection step in Section 4.1 is a concrete, citable protocol detail: pretrained weights for the DI branch are chosen by performance on the same benchmarks used for final evaluation (FashionIQ validation) or for ablation and hyperparameter decisions (CIRR validation). This is a form of tuning on the target evaluation set that the baselines did not receive, so it could explain part of the reported gains without any detail-transfer story. The paper's own ablation (Table 2b) shows that mixing IPr2Pr and CIRR in Stage 2 hurts, which supports the reader's distribution-shift concern, but that does not address the selection bias. I agree with the reader's conditional verdict: the authors should disclose and control for this selection, report the variance across Stage 1 iterations, and ideally release code. The concern strengthens the condition but does not move the verdict; hence UNCHANGED. Credit is due for the extensive ablations and the fairness control against simply adding pretraining data to SPRC, but the checkpoint-selection issue remains unaddressed and is the most load-bearing threat to the central SOTA claim.","tokens_in":22684,"tokens_out":6929,"duration_ms":70071,"concrete_test":"Re-run the full pipeline with the Stage 1 DI branch checkpoint fixed to the last iteration (or to a checkpoint selected on a held-out split of IPr2Pr), keeping Stage 2 and Stage 3 settings identical. Compare CIRR validation Recall@1 and FashionIQ Avg Recall against Table 2(a) (56.83 vs 54.32 for SPRC) and Table 4 (66.50 vs 64.85). If either margin shrinks by more than ~1 point, the SOTA claim is substantially owed to target-benchmark checkpoint selection rather than to the method.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is in the experimental protocol for Stage 1, not the transfer mechanism per se. Section 4.1 (Implementation Details) states: 'the best-performing iteration on CIRR and FashionIQ is selected as the initialization for the next stage.' Stage 1 pretraining is on IPr2Pr, yet the checkpoint is chosen by evaluating the DI branch on the target benchmarks. For FashionIQ, the reported numbers are on the validation set (same as prior work), so this is selection on the evaluation set itself. For CIRR, the test set is separate, but the validation set is used for ablations and hyperparameter choices, and the same selection protocol applies. Baselines such as SPRC were not given this pretraining-selection step, so the comparison is not controlled. The paper does not report the range of performance across Stage 1 iterations or the result of using the final (or an IPr2Pr-held-out) checkpoint. Since the central claim is SOTA on these benchmarks, the gain attributed to detail enhancement could be partly a validation-selection artifact.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DetailFusion, a dual-branch framework for supervised composed image retrieval (CIR). A Detail-oriented Inference (DI) branch is pre-trained on the InstructPix2Pix (IPr2Pr) image-editing dataset with a contrastive loss that uses the reference image as a hard negative, then jointly fine-tuned with a Global Feature Matching (GM) branch on CIR datasets. An Adaptive Feature Compositor, trained in a third stage, dynamically fuses the global and detail features. The authors report state-of-the-art results on CIRR and FashionIQ, provide ablations supporting each design choice, and include zero-shot evaluations on CIRCO and GeneCIS, plus computational-cost comparisons. The paper claims that explicitly training one branch on fine-grained visual and textual variations improves supervised CIR at 10-20% additional training cost.","tokens_in":22851,"tokens_out":5370,"duration_ms":59636,"significance":"If the claimed results hold, the work is significant for supervised CIR: it would demonstrate that image-editing triplets, though distributionally different from retrieval data, can be used to train a detail-specialized branch that transfers to CIR benchmarks, and that a lightweight adaptive compositor can combine global and detail features effectively. The paper has several genuine strengths: the ablation of the SPRC-with-same-pretraining control in Table 2(a) directly targets the 'more data' confound, the three-stage training strategy is systematically ablated in Tables 2(b)-2(d), the loss-function variants are compared in Table 2(c), and the cross-domain zero-shot results in Tables 5-6 provide external evidence for the transferability of the detail prior. The main weakness is an experimental-protocol issue: the Stage-1 pre-training checkpoint is selected on the target benchmarks, which is not controlled for in the comparison against baselines, and all numbers are single-run with no error bars. These issues bear directly on the central 'state-of-the-art' claim.","major_comments":[{"comment":"The text states: 'the best-performing iteration on CIRR and FashionIQ is selected as the initialization for the next stage.' This is selection on the target benchmarks: for FashionIQ, Table 1 reports results on exactly this validation set, so checkpoint selection and evaluation share the same labels; for CIRR, the validation set is also used for ablation and hyperparameter choices. No comparable selection step is described for SPRC or the other baselines. This makes the reported SOTA comparison uncontrolled, and the gains attributed to detail enhancement could partly be a validation-selection artifact. Please report the performance range across Stage-1 iterations, use a selection criterion that does not touch the evaluation set (e.g., an IPr2Pr held-out split or the final checkpoint), and rerun the comparison under that protocol.","section":"Section 4.1, Implementation Details"},{"comment":"The 'Fairness of Using Image Editing Data for Pre-Training' ablation is the right control, but the text only says SPRC was pre-trained 'under the same dataset and training strategy.' It is not stated whether the same target-benchmark Stage-1 checkpoint selection was applied to SPRC. If it was not, the control does not rule out the selection artifact identified above. Please specify the exact protocol used for the SPRC-pretrained baseline; if the selection was applied, the numbers should be presented with the same detail.","section":"Section 4.3, Table 2(a)"},{"comment":"All reported results are from a single run with no standard deviations or confidence intervals. This is particularly concerning because some margins over strong baselines are small: on CIRR R@1, DetailFusion(Ours) is 54.55 versus SPRC† at 55.06, and on FashionIQ the average is 66.50 versus 66.41 for SPRC†. The 'outperforms previous methods' claim would be much more robust with repeated runs or, at minimum, an explicit statement of training variance. I would like to see either error bars over at least 3-5 seeds for the main comparisons, or a clear explanation of why such small differences should be treated as reliable.","section":"Tables 1 and 3"}],"minor_comments":[{"comment":"The denominator appears to contain a typo: 'L' is used where a plus sign is intended in the sum over S(D(Q(i)), D(I_t(j))) and S(D(Q(i)), D(I_r(j))). The definition 'S(A, BL C) := S(A,B) + S(A,C)' should be rewritten with an explicit plus sign and a short explanation.","section":"Equation (2)"},{"comment":"Equation (6) uses identical notation D(Q) = G(Q) = LinearH(EH(EI(Ir), Tm)), but the DI and GM branches have separate parameters according to Figure 3(a). Please use distinct symbols for the two branches' mapping and encoder parameters so that the shared-architecture but non-shared-parameter setup is unambiguous.","section":"Equation (6)"},{"comment":"The text says 'the DI branch also serves as the image encoder,' while Figure 3(a) says the image encoder shares parameters with the DI branch. Please clarify whether the image encoder is a separate frozen module or is literally the image side of the DI branch; this affects how Equations (6)-(7) should be interpreted.","section":"Section 3.4"},{"comment":"The row 'SPRC2* [24]' appears in Table 3 without a definition or a reference entry. If this is a variant from [24], it should be described in the table caption or in Appendix B.","section":"Table 3 and Appendix B"},{"comment":"The GitHub URL is given in the abstract and the conclusion promises public code. For a journal submission, please ensure that the repository is actually public and contains the training and evaluation code before acceptance; otherwise remove the URL from the abstract.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The core idea is interesting and the ablations are generally well designed, but I share the stress-test concern: the Stage-1 checkpoint selection on the target benchmarks is a real confound for the SOTA claim, and the absence of error bars makes the small margins against SPRC† hard to evaluate. The paper is within scope and the issue is fixable by rerunning with a fair selection protocol and reporting variance. I would also encourage the editor to require code release, since reproducibility is a key strength of the paper's empirical claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The core idea is new and reasonably motivated: a detail-oriented branch pretrained on InstructPix2Pix editing triplets with the reference image as a hard negative, then an adaptive compositor that fuses this with a global branch. The ablations are genuinely careful: they test the transfer mechanism directly, show that mixing IPr2Pr with CIRR data in Stage 2 hurts, and include a SPRC-with-same-pretraining control to address the 'more data' confound. If the numbers hold, the gains (+2.6 R@1 on CIRR, +1.65 avg on FashionIQ) are worth having for CIR practitioners at a modest training cost.\n\nThe problem is the experimental protocol. Section 4.1 says the Stage 1 checkpoint is chosen as 'the best-performing iteration on CIRR and FashionIQ.' For FashionIQ, the final numbers are on the validation set, so this is selection on the evaluation set itself. For CIRR, the test set is separate, but the validation set is still used for the selection and for many ablations and hyperparameters. SPRC and the other baselines were not given this selection step, so the comparison is not apples-to-apples. The fairness control with SPRC is a good instinct, but it does not remove the worry if the selection happened only for DetailFusion. This alone makes the SOTA claim premature.\n\nOther soft spots are minor by comparison: no error bars (single run), no code release despite the abstract promising it, and equation (2) has a garbled denominator where the 'L' appears instead of the intended plus. The 'first to use an image editing dataset for CIR training' claim also overreaches—CompoDiff and others generate or use editing-style triplets even if not exactly the same protocol.\n\nThe central design is sensible and the transfer mechanism is at least plausible. I would not desk-reject this. A serious referee should ask for: (1) a clean checkpoint-selection protocol (e.g., fixed iteration count or selection on a separate held-out split), (2) variance across 2-3 runs, and (3) code. If the gains survive that, this is a solid contribution to supervised CIR. If they don't, the paper is still a useful study of detail-oriented pretraining, just not a SOTA claim.","headline":"A solid, well-ablated CIR method whose SOTA numbers are compromised by a validation-set checkpoint selection that baselines don't get.","tokens_in":23441,"tokens_out":3120,"would_cite":false,"duration_ms":23825,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DetailFusion claims that pre-training one branch on image-editing triplets and adaptively fusing it with a global branch gives state-of-the-art composed image retrieval on CIRR and FashionIQ.","keywords":["composed image retrieval","detail enhancement","dual-branch framework","image editing pre-training","hard negative","adaptive feature compositor","CIRR","FashionIQ"],"falsifier":"Replace the IPr2Pr pre-training with an equal-size sample of the target CIR dataset's own triplets, keeping the same hard-negative loss and the same branch architecture; if Recallsubset@1 on CIRR does not drop, then the paper's attribution of the gain to image-editing priors is falsified.","tokens_in":22465,"feed_emoji":"🔍","tokens_out":6281,"duration_ms":41956,"temperature":0.7,"pith_summary":"Composed image retrieval lets a user find an image by pointing at a reference photo and describing a change, like making the cup brown instead of red. The paper argues that previous supervised methods fail on such queries because they fuse image and text at a coarse, global level and never learn to perceive fine-grained visual details or execute small textual modifications. It proposes DetailFusion, a two-branch framework: one branch is dedicated to global semantics, and one detail-oriented inference branch is pre-trained on the InstructPix2Pix image-editing dataset, where each query's reference image is deliberately used as a hard negative during training. A lightweight Adaptive Feature Compositor then learns to combine the two branches' features per query. The paper reports state-of-the-art Recall@K on CIRR and FashionIQ and claims the added detail capability transfers to unseen domains in zero-shot settings.","feed_headline":"Edit-trained detail branch lifts composed image retrieval","feed_subtitle":"Fusing it with a global branch sets top recall on CIRR and FashionIQ.","key_machinery":"The load-bearing object is the Detail-oriented Inference (DI) branch, trained with a detail-oriented contrastive loss in which the reference image itself is added as a hard negative to the denominator, forcing the model to ignore near-duplicate visual similarity and attend to the specific alteration described by the text. The second mechanism is the Adaptive Feature Compositor, which uses cross-attention between the [CLS] tokens of the two branches and their detailed token sets, then forms a convex combination of the global and detail features with a learned ratio and a bridging feature. The design rests on the premise that global and detail branches are complementary, as shown by branch-level ablations where the GM branch excels at global Recall@1 while the DI branch excels at subset retrieval.","core_discovery":"On the paper's own terms, the central claim is that the failure of current supervised CIR models to handle subtle modifications stems from missing targeted training for fine-grained details, and that this can be fixed by a three-stage recipe. First, pre-train a Detail-oriented Inference branch on the InstructPix2Pix image-editing dataset, whose triplets have near-identical reference and target images and short atomic edit texts, using a contrastive loss that treats the reference image as a hard negative. Second, fine-tune that branch together with a Global Feature Matching branch on CIR data, keeping the detail-oriented loss for the detail branch. Third, train a lightweight Adaptive Feature Compositor that fuses global and detail features through cross-attention and a convex combination. The paper claims this yields state-of-the-art results on CIRR and FashionIQ, with gains particularly on the CIRR subset metric that measures fine-grained discrimination, and that the detail capability transfers zero-shot to CIRCO and GeneCIS.","pith_inferences":["If the atomic detail priors transfer, one could pre-train the detail branch on any large paired image-edit dataset and expect similar gains on CIR benchmarks, suggesting a data-centric alternative to architecture search.","The adaptive compositor's convex combination with a learned bridging feature could generalize beyond CIR to other tasks where a query has two granularities, such as visual question answering or image-caption retrieval.","The hard-negative reference trick is a general recipe for contrastive retrieval: injecting a near-duplicate query's own reference as a negative forces finer discrimination, and the paper's ablations imply this is what preserves the DI branch's subset performance.","A testable extension is to evaluate on a benchmark with edit-type labels to see whether the DI branch learns disentangled atomic edit types, since the current aggregate subset metric does not reveal whether individual transformations are decoupled."],"forward_implications":["Supervised CIR models can be improved without new retrieval-specific annotations by borrowing an image-editing dataset, since editing triplets encode atomic text-image transformations.","Explicitly keeping a detail branch and a global branch separate, then fusing them adaptively, outperforms single-branch fusion; the DI branch alone wins on CIRR subset metrics while the GM branch wins on global Recall@1.","The detail capability transfers to unseen domains: zero-shot CIRCO and GeneCIS results improve over the same-backbone baseline, with the largest gains on 'change' subsets.","The extra training cost is modest (10-20% over single-stage methods) and inference is about 1.5x a single Q-Former and only 12% slower than SPRC, making the approach practical.","Directly mixing the editing dataset with CIR data in a single stage hurts performance; the three-stage coarse-to-fine schedule is necessary for the transfer to work."],"supporting_citations":[{"why":"Supplies the InstructPix2Pix image-editing dataset used to pre-train the detail branch.","marker":"[11]"},{"why":"Provides the BLIP-2 Q-Former hybrid-modal encoder that initializes both branches.","marker":"[10]"},{"why":"Provides the frozen EVA-CLIP ViT-G/14 vision encoder used for image features.","marker":"[35]"},{"why":"The main supervised CIR baseline compared against and used in the fairness pre-training ablation.","marker":"[24]"},{"why":"Scaling Positives and Negatives method whose variant is used to further boost the two branches and the compositor.","marker":"[37]"},{"why":"Defines the CIRR benchmark and its fine-grained subset evaluation.","marker":"[2]"},{"why":"Defines the FashionIQ benchmark used for fashion-domain evaluation.","marker":"[33]"},{"why":"Supplies the general feature combiner structure that the Compositor's fusion block adapts.","marker":"[14]"}],"fun_headline_variants":["DetailFusion boosts CIR with fine-grained training","Dual-branch fuses global and detail for better CIR","Editing data trains detail branch for sharper retrieval","Adaptive fusion of global and detail sets new CIR SOTA","Detail-aware CIR leverages edit priors and cross-attention"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole approach rests on the assumption that the kinds of small edits found in an image-editing dataset teach a model the same fine-grained distinctions that composed image retrieval queries ask for; if those editing cues do not transfer, the detail branch's extra training will not help retrieval.","fun_headline_variants_meta":{"raw":{"variants":["DetailFusion boosts CIR with fine-grained training","Dual-branch fuses global and detail for better CIR","Editing data trains detail branch for sharper retrieval","Adaptive fusion of global and detail sets new CIR SOTA","Detail-aware CIR leverages edit priors and cross-attention"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1296,"prompt_tokens":926,"completion_tokens":370,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":288}},"tokens_in":542,"tokens_out":370,"duration_ms":3837,"temperature":1.0,"reasoning_tokens":288,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:40:03.370810+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the IPr2Pr pre-training with an equal-size sample of the target CIR dataset's own triplets, keeping the same hard-negative loss and the same branch architecture; if Recallsubset@1 on CIRR does not drop, then the paper's attribution of the gain to image-editing priors is falsified.","supporting_citations":[{"cited_title":"Eva: Exploring the limits of masked visual representation learning at scale","cited_arxiv_id":null,"evidence_quote":"Provides the frozen EVA-CLIP ViT-G/14 vision encoder used for image features."},{"cited_title":"Sentence-level prompts benefit composed image retrieval","cited_arxiv_id":null,"evidence_quote":"The main supervised CIR baseline compared against and used in the fairness pre-training ablation."},{"cited_title":"Improving composed image retrieval via contrastive learning with scaling positives and negatives","cited_arxiv_id":null,"evidence_quote":"Scaling Positives and Negatives method whose variant is used to further boost the two branches and the compositor."},{"cited_title":"Fashion iq: A new dataset towards retrieving images by natural language feedback","cited_arxiv_id":null,"evidence_quote":"Defines the FashionIQ benchmark used for fashion-domain evaluation."}],"review_version":1}