REVIEW 4 cited by
Image-text Retrieval: A Survey on Recent Research and Development
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the past few years, cross-modal image-text retrieval (ITR) has experienced increased interest in the research community due to its excellent research value and broad real-world application. It is designed for the scenarios where the queries are from one modality and the retrieval galleries from another modality. This paper presents a comprehensive and up-to-date survey on the ITR approaches from four perspectives. By dissecting an ITR system into two processes: feature extraction and feature alignment, we summarize the recent advance of the ITR approaches from these two perspectives. On top of this, the efficiency-focused study on the ITR system is introduced as the third perspective. To keep pace with the times, we also provide a pioneering overview of the cross-modal pre-training ITR approaches as the fourth perspective. Finally, we outline the common benchmark datasets and valuation metric for ITR, and conduct the accuracy comparison among the representative ITR approaches. Some critical yet less studied issues are discussed at the end of the paper.
Forward citations
Cited by 4 Pith papers
-
Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
A variance-based, retraining-free pruning framework for vision-language models that allocates per-layer sparsity and outperforms Wanda and SparseGPT at high sparsity.
-
Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
CAW adds a confidence-weighted KL loss and feature-alignment regularization to CLIP adversarial fine-tuning, raising average AutoAttack robust accuracy from 31.6% to 33.5% on 15 datasets.
-
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes
CM50, a 33k-meme dataset with GPT-4o-generated annotations, and mtrCLIP, a fine-tuned CLIP model, together improve meme-text retrieval on MemeCap over the original CLIP.
-
Robust image classification with multi-modal large language models
Multi-Shield rejects adversarial images by abstaining whenever a standard image classifier and a CLIP zero-shot classifier disagree, which raises robust accuracy noticeably under ordinary attacks and modestly under ad...
Discussion (0). Continue with ORCID to comment.