Pith. sign in

REVIEW 2 cited by

Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.14263 v3 pith:LAUXDK4H submitted 2023-08-28 cs.IR cs.MM

classification cs.IRcs.MM
keywords retrievalcross-modaldatamethodsacrosscomprehensivedirectionsfield
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the exponential surge in diverse multi-modal data, traditional uni-modal retrieval methods struggle to meet the needs of users seeking access to data across various modalities. To address this, cross-modal retrieval has emerged, enabling interaction across modalities, facilitating semantic matching, and leveraging complementarity and consistency between heterogeneous data. Although prior literature has reviewed the field of cross-modal retrieval, it suffers from numerous deficiencies in terms of timeliness, taxonomy, and comprehensiveness. This paper conducts a comprehensive review of cross-modal retrieval's evolution, spanning from shallow statistical analysis techniques to vision-language pre-training models. Commencing with a comprehensive taxonomy grounded in machine learning paradigms, mechanisms, and models, the paper delves deeply into the principles and architectures underpinning existing cross-modal retrieval methods. Furthermore, it offers an overview of widely-used benchmarks, metrics, and performances. Lastly, the paper probes the prospects and challenges that confront contemporary cross-modal retrieval, while engaging in a discourse on potential directions for further progress in the field. To facilitate the ongoing research on cross-modal retrieval, we develop a user-friendly toolbox and an open-source repository at https://cross-modal-retrieval.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new 400-document, 8,250-question benchmark measures how well vision-language models retrieve hidden text and image facts from long documents.

  2. Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A structured review of 81 text-to-video retrieval papers that leverage auxiliary information, organized by a taxonomy and compared on standard benchmarks.

Pith tools