Pith. sign in

REVIEW 2 cited by

Cross-lingual Information Retrieval with BERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.13005 v1 pith:QNFHKA3V submitted 2020-04-24 cs.IR cs.CLcs.LGstat.ML

classification cs.IRcs.CLcs.LGstat.ML
keywords bertmodelretrievalcross-lingualdocumentsenglishinformationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple neural language models have been developed recently, e.g., BERT and XLNet, and achieved impressive results in various NLP tasks including sentence classification, question answering and document ranking. In this paper, we explore the use of the popular bidirectional language model, BERT, to model and learn the relevance between English queries and foreign-language documents in the task of cross-lingual information retrieval. A deep relevance matching model based on BERT is introduced and trained by finetuning a pretrained multilingual BERT model with weak supervision, using home-made CLIR training data derived from parallel corpora. Experimental results of the retrieval of Lithuanian documents against short English queries show that our model is effective and outperforms the competitive baseline approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Anveshana is a new English-to-Sanskrit cross-lingual retrieval benchmark on Srimadbhagavatam chapters, with document translation outperforming direct and query-translation approaches.

  2. Multilingual Open QA on the MIA Shared Task

    cs.CL 2025-01 reject novelty 5.0 of 10

    Zero-shot question-generation reranking boosts Korean and Japanese retrieval but hurts Finnish and Bengali, and translated training data yields no consistent QA improvement.

Pith tools