{"id":"a71c09e8-5041-4262-953a-de6527859e4e","arxiv_id":"2607.02284","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"FlowCIR frames ZS-CIR as conditional flow matching transport on fixed VLM embeddings plus an inference-time Multi-Negative Steering fix for negation, reporting competitive benchmark results at far lower training cost.","lead":"FlowCIR casts zero-shot composed image retrieval as conditional semantic transport using flow matching on pre-extracted VLM embeddings. This trains only a small module without updating encoders and claims roughly 10x lower training cost than textual-inversion baselines.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Lightweight flow-matching transport on frozen VLM embeddings may fail to recover fine-grained target semantics","rationale":"The reader's weakest_assumption matches the load-bearing point exactly. Because the review was performed on the abstract, the full-text experiments would be needed to quantify whether the transport module actually compensates for the acknowledged VLM limitations; until then the UNVERDICTED status is appropriate.","tokens_in":1784,"tokens_out":284,"duration_ms":13841,"concrete_test":"Ablate the Multi-Negative Steering at inference and recompute recall@K on the negation-heavy subset of CIRR or FashionIQ; if the drop exceeds the gap to the strongest textual-inversion baseline, the frozen-embedding transport assumption does not hold for the hardest cases.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that a small conditional flow-matching network, trained only on fixed pre-extracted VLM embeddings and without any encoder gradients or task-specific triplets, can learn a transport field that accurately composes reference + instruction into a target-aligned query. This is the least secure step: the paper itself flags negation/removal as a major VLM failure mode and handles it with a separate inference-time heuristic, indicating that the embedding space already loses critical directional information that the transport module is then asked to recover without additional supervision or adaptation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes FlowCIR for zero-shot composed image retrieval (ZS-CIR), framing the task as conditional semantic transport via flow matching. A lightweight transport module is trained on fixed pre-extracted VLM embeddings to map an instruction representation (conditioned on a reference image) to a target-aligned query embedding, without updating encoders or using task-specific triplets. It claims this yields competitive performance on standard CIR benchmarks while requiring roughly 10× fewer training resources than textual-inversion baselines, and introduces an inference-only Multi-Negative Steering heuristic to address VLM limitations on negation/removal.","tokens_in":1879,"tokens_out":440,"duration_ms":22462,"significance":"If the performance and efficiency claims hold, the work offers a paradigm shift from textual inversion to flow-based transport on frozen embeddings, potentially lowering the barrier for ZS-CIR research. The explicit identification of negation as a VLM failure mode and the proposed mitigation are constructive contributions.","major_comments":[{"comment":"Abstract and method description: The central claim that a small conditional flow-matching network trained solely on fixed VLM embeddings can accurately recover fine-grained target semantics (including directional composition) is load-bearing, yet the paper itself identifies negation/removal as a major VLM failure mode and resorts to a separate inference-time heuristic; this indicates the learned transport field may not fully compensate for information lost in the embedding space without additional supervision or adaptation.","section":"Abstract / Method"},{"comment":"Abstract: The efficiency claim of 'roughly 10× fewer training resources' is presented without concrete metrics (e.g., parameter count of the transport module, GPU-hours, epochs, or side-by-side comparison tables), which is required to substantiate the load-bearing advantage over textual-inversion approaches.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract would benefit from a one-sentence description of the specific conditional flow-matching objective or network architecture to clarify how the transport field is parameterized.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and for recognizing the potential paradigm shift offered by FlowCIR. We address each major comment below with point-by-point responses.","responses":[{"response":"We agree that negation and removal constitute a notable VLM limitation, which is why the manuscript explicitly identifies this failure mode and introduces Multi-Negative Steering as a targeted inference-time mitigation. The conditional flow-matching transport is trained to learn directional semantic mappings on the fixed embeddings for general compositional instructions, and benchmark results indicate it recovers target semantics effectively in most cases. The steering heuristic specifically addresses residual negation handling issues that are not fully resolved in the VLM embedding space. We will revise the abstract and method sections to more clearly separate the scope of the learned transport from the additional steering strategy and to discuss this distinction as a limitation.","revision_made":"partial","referee_comment":"[Abstract / Method] Abstract and method description: The central claim that a small conditional flow-matching network trained solely on fixed VLM embeddings can accurately recover fine-grained target semantics (including directional composition) is load-bearing, yet the paper itself identifies negation/removal as a major VLM failure mode and resorts to a separate inference-time heuristic; this indicates the learned transport field may not fully compensate for information lost in the embedding space without additional supervision or adaptation."},{"response":"We acknowledge that the efficiency claim requires concrete supporting metrics to be fully substantiated. In the revised manuscript we will add explicit details on the transport module's parameter count, training epochs, approximate GPU-hours, and a side-by-side resource comparison against textual-inversion baselines.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The efficiency claim of 'roughly 10× fewer training resources' is presented without concrete metrics (e.g., parameter count of the transport module, GPU-hours, epochs, or side-by-side comparison tables), which is required to substantiate the load-bearing advantage over textual-inversion approaches."}],"tokens_in":1449,"tokens_out":429,"duration_ms":20204,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core move is treating ZS-CIR as learning a conditional transport field with flow matching instead of inverting the reference image into text tokens and concatenating. The model trains only a lightweight module on fixed embeddings, which the abstract says cuts training cost by roughly 10x. That efficiency claim is the clearest practical difference from prior work.\n\nIt also flags negation and removal as a persistent VLM weakness and counters it with a separate Multi-Negative Steering step at inference. That is a reasonable engineering patch, but it sits outside the learned transport.\n\nThe soft spot is verification. Without equations for the flow-matching objective, without ablation tables on the transport module size or conditioning, and without breakdowns on negation-heavy queries, it is impossible to tell whether the reported competitive numbers come from the flow field itself or from the steering heuristic and the underlying VLM. The weakest assumption in the abstract—that a small network on static embeddings can reconstruct fine-grained target semantics the VLM already dropped—remains untested in the supplied text.\n\nThis is a narrow but self-contained idea aimed at the ZS-CIR subfield. Readers already running retrieval experiments on CLIP-style embeddings could extract the method and check the efficiency numbers themselves. The work is coherent on its own terms and shows clear engagement with the limitations of current VLM composition, so it clears the bar for a serious referee even if the final verdict depends on the missing tables.","headline":"FlowCIR swaps textual inversion for a small conditional flow-matching transport on frozen VLM embeddings and adds an inference heuristic for negation, but the abstract supplies no equations, ablations, or error breakdowns to confirm the transport actually recovers the claimed semantics.","tokens_in":2373,"tokens_out":384,"would_cite":false,"duration_ms":13832,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"FlowCIR casts zero-shot composed image retrieval as conditional semantic transport learned by flow matching on fixed vision-language model embeddings.","keywords":["zero-shot composed image retrieval","flow matching","semantic transport","vision-language models","negation handling","image retrieval","conditional transport"],"falsifier":"A controlled experiment showing that FlowCIR retrieval accuracy on standard benchmarks drops below that of a textual-inversion baseline when both methods receive identical training compute and the same pre-trained encoders.","tokens_in":2683,"feed_emoji":"🔄","tokens_out":592,"duration_ms":23670,"temperature":0.7,"pith_summary":"The paper argues that converting a reference image and text instruction into a target query can be done by training a transport field that moves the instruction embedding toward the correct target embedding, conditioned on the reference. This replaces the textual-inversion step used in earlier methods, which the authors view as lossy for fine details. Because the transport module trains only on pre-extracted embeddings and leaves the encoders untouched, the approach requires far less compute than inversion-based training. The work further introduces an inference-time correction that steers away from negated concepts when the instruction contains removal or negation language.","feed_headline":"Flow matching replaces textual inversion for zero-shot image retrieval","feed_subtitle":"A small module learns to transport reference-conditioned instructions to target embeddings using ten times less training than earlier method","key_machinery":"Conditional flow matching transport field that maps the instruction representation toward a target-aligned query embedding conditioned on the reference image.","core_discovery":"Zero-shot composed image retrieval is reformulated as learning a conditional flow-matching transport field that maps an instruction representation, given the reference image, directly to a target-aligned query embedding; the resulting lightweight module produces competitive retrieval accuracy on standard benchmarks while using roughly ten times fewer training resources than textual-inversion baselines and incorporates a Multi-Negative Steering procedure to offset vision-language model weaknesses on negation.","pith_inferences":["The same transport formulation could be tested on other vision-language tasks that currently rely on token inversion or simple concatenation for composition.","Because the approach never updates the underlying encoders, it may allow reuse of the same transport module across different vision-language model backbones."],"forward_implications":["The method reaches strong performance on existing CIR benchmarks without requiring domain-specific triplet annotations.","Training cost is reduced by a factor of roughly ten compared with prior textual-inversion pipelines.","Multi-Negative Steering at inference improves results on queries that contain negation or removal instructions."],"fun_headline_variants":["Flow matching casts ZS-CIR as conditional semantic transport","Lightweight transport module uses ten times less training","Semantic transport via flow matching for zero-shot retrieval","Flow matching learns transport field for image retrieval"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"A lightweight transport module trained solely on fixed pre-extracted vision-language model embeddings can capture the fine-grained semantics needed for accurate target retrieval without any encoder updates.","fun_headline_variants_meta":{"raw":{"variants":["Flow matching casts ZS-CIR as conditional semantic transport","Lightweight transport module uses ten times less training","Semantic transport via flow matching for zero-shot retrieval","Flow matching learns transport field for image retrieval"]},"model":"grok-4.3","cost_usd":0.005735,"raw_usage":{"total_tokens":2764,"prompt_tokens":725,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":57349500,"prompt_tokens_details":{"text_tokens":725,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1981,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":725,"tokens_out":58,"duration_ms":15128,"temperature":1.0,"reasoning_tokens":1981,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-03T15:39:06.080601+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled experiment showing that FlowCIR retrieval accuracy on standard benchmarks drops below that of a textual-inversion baseline when both methods receive identical training compute and the same pre-trained encoders.","supporting_citations":[],"review_version":1}