Pith. sign in

REVIEW 2 cited by

ColBERT-XM: A Modular Multi-Vector Representation Model for Zero-Shot Multilingual Information Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15059 v1 pith:6CA6C347 submitted 2024-02-23 cs.CL cs.IR

classification cs.CLcs.IR
keywords languagesdataretrievalcolbert-xmmodelmodelsmodularmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-of-the-art neural retrievers predominantly focus on high-resource languages like English, which impedes their adoption in retrieval scenarios involving other languages. Current approaches circumvent the lack of high-quality labeled data in non-English languages by leveraging multilingual pretrained language models capable of cross-lingual transfer. However, these models require substantial task-specific fine-tuning across multiple languages, often perform poorly in languages with minimal representation in the pretraining corpus, and struggle to incorporate new languages after the pretraining phase. In this work, we present a novel modular dense retrieval model that learns from the rich data of a single high-resource language and effectively zero-shot transfers to a wide array of languages, thereby eliminating the need for language-specific labeled data. Our model, ColBERT-XM, demonstrates competitive performance against existing state-of-the-art multilingual retrievers trained on more extensive datasets in various languages. Further analysis reveals that our modular approach is highly data-efficient, effectively adapts to out-of-distribution data, and significantly reduces energy consumption and carbon emissions. By demonstrating its proficiency in zero-shot scenarios, ColBERT-XM marks a shift towards more sustainable and inclusive retrieval systems, enabling effective information accessibility in numerous languages. We publicly release our code and models for the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bilingual BSARD: Extending Statutory Article Retrieval to Dutch

    cs.CL 2024-12 conditional novelty 5.0 of 10

    The authors extend the French BSARD legal retrieval dataset to Dutch (bBSARD) and benchmark retrieval models, showing small fine-tuned language-specific models can outperform zero-shot proprietary embeddings.

  2. BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language

    cs.CL 2024-12 conditional novelty 4.0 of 10

    BEIR-NL is a Dutch-translated version of the BEIR benchmark with evaluations showing BM25 remains competitive against multilingual dense models.

Pith tools