Pith. sign in

REVIEW 4 cited by

Zero-shot Composed Text-Image Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07272 v2 pith:FV4MJEX2 submitted 2023-06-12 cs.CV

classification cs.CV
keywords datasetsmodelautomaticallycomposedconductimagesinformationproposed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we consider the problem of composed image retrieval (CIR), it aims to train a model that can fuse multi-modal information, e.g., text and images, to accurately retrieve images that match the query, extending the user's expression ability. We make the following contributions: (i) we initiate a scalable pipeline to automatically construct datasets for training CIR model, by simply exploiting a large-scale dataset of image-text pairs, e.g., a subset of LAION-5B; (ii) we introduce a transformer-based adaptive aggregation model, TransAgg, which employs a simple yet efficient fusion mechanism, to adaptively combine information from diverse modalities; (iii) we conduct extensive ablation studies to investigate the usefulness of our proposed data construction procedure, and the effectiveness of core components in TransAgg; (iv) when evaluating on the publicly available benckmarks under the zero-shot scenario, i.e., training on the automatically constructed datasets, then directly conduct inference on target downstream datasets, e.g., CIRR and FashionIQ, our proposed approach either performs on par with or significantly outperforms the existing state-of-the-art (SOTA) models. Project page: https://code-kunkun.github.io/ZS-CIR/

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    ZeroSight supplies a video-derived dataset and evaluation protocol for genuine zero-shot composed image retrieval plus the SC4CIR consistency method, demonstrating that prior benchmarks inflate reported performance ac...

  2. FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Disentangled fine-grained context then retrieval fine-tuning on an 87K auto-generated quintuple CIR dataset lifts a 4B MLLM past larger universal retrievers on complex zero-shot image search.

  3. Mixed-Modality Dual Face-Hair Retrieval

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Introduces DFHR task, DFHR-Bench with over 180K triplets, and MFHC framework for mixed-modality dual face-hair retrieval.

  4. Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval

    cs.CV 2026-01 conditional novelty 5.0 of 10

    By generating an edited "mental image" of a query and synthetic counterparts of database images, and matching in that synthetic space, Paracosm achieves state-of-the-art training-free zero-shot composed image retrieva...

Pith tools