Pith. sign in

REVIEW 6 cited by

ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.08101 v2 pith:AKXIBHV7 submitted 2022-03-15 cs.CV cs.IR

classification cs.CVcs.IR
keywords imageimagesretrievalcomplementaryelementsexamplefeaturesimplicit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An intuitive way to search for images is to use queries composed of an example image and a complementary text. While the first provides rich and implicit context for the search, the latter explicitly calls for new traits, or specifies how some elements of the example image should be changed to retrieve the desired target image. Current approaches typically combine the features of each of the two elements of the query into a single representation, which can then be compared to the ones of the potential target images. Our work aims at shedding new light on the task by looking at it through the prism of two familiar and related frameworks: text-to-image and image-to-image retrieval. Taking inspiration from them, we exploit the specific relation of each query element with the targeted image and derive light-weight attention mechanisms which enable to mediate between the two complementary modalities. We validate our approach on several retrieval benchmarks, querying with images and their associated free-form text modifiers. Our method obtains state-of-the-art results without resorting to side information, multi-level features, heavy pre-training nor large architectures as in previous works.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    FoCo learns composition for zero-shot CIR via text-anchored visual aggregation and context-conditioned semantic completion trained jointly with cross-instance contrastive loss, reporting SOTA on four benchmarks.

  2. Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    ZeroSight supplies a video-derived dataset and evaluation protocol for genuine zero-shot composed image retrieval plus the SC4CIR consistency method, demonstrating that prior benchmarks inflate reported performance ac...

  3. Mixed-Modality Dual Face-Hair Retrieval

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Introduces DFHR task, DFHR-Bench with over 180K triplets, and MFHC framework for mixed-modality dual face-hair retrieval.

  4. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    STiTch refines LLM captions via embedding transition and uses set-to-set bidirectional transportation alignment to improve training-free zero-shot composed image retrieval.

  5. Beyond Simple Edits: Composed Video Retrieval with Dense Modifications

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new benchmark with much longer, denser modification texts, plus a single-encoder fusion model, raises composed video retrieval Recall@1 by 3.4 points on its own test set.

  6. FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FAR-Net combines Q-Former cross-attention alignment with uncertainty-perturbed contrastive learning and reports up to 2.4 points higher Recall@1 on standard CIR benchmarks.

Pith tools