Pith. sign in

HLTCOE at TREC 2023 NeuCLIR Track

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The HLTCOE team applied PLAID, an mT5 reranker, and document translation to the TREC 2023 NeuCLIR track. For PLAID we included a variety of models and training techniques -- the English model released with ColBERT v2, translate-train~(TT), Translate Distill~(TD) and multilingual translate-train~(MTT). TT trains a ColBERT model with English queries and passages automatically translated into the document language from the MS-MARCO v1 collection. This results in three cross-language models for the track, one per language. MTT creates a single model for all three document languages by combining the translations of MS-MARCO passages in all three languages into mixed-language batches. Thus the model learns about matching queries to passages simultaneously in all languages. Distillation uses scores from the mT5 model over non-English translated document pairs to learn how to score query-document pairs. The team submitted runs to all NeuCLIR tasks: the CLIR and MLIR news task as well as the technical documents task.

fields

cs.IR 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval cs.IR · 2025-01-31 · conditional · none · ref 60 · internal anchor

    The paper introduces mFollowIR, a multilingual instruction-following retrieval benchmark across Russian, Chinese, and Persian, and finds that English instruction-trained models transfer cross-lingually but struggle in multilingual settings.