{"id":"7459c04d-e7b0-4256-8c6c-8f806a63e040","arxiv_id":"2412.08329","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"BEIR-NL is a Dutch-translated version of the BEIR benchmark with evaluations showing BM25 remains competitive against multilingual dense models.","lead":"The authors translated 14 English information retrieval benchmarks into Dutch to create BEIR-NL, a zero-shot evaluation suite for Dutch IR models. They find that BM25 remains a strong baseline, while larger multilingual dense models perform best, and that translation quality measurably lowers retrieval scores.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's claim that BM25 is only outperformed by larger dense retrieval models is contradicted by Table 3: e5-small (118M, retrieval-finetuned) beats BM25 on FiQA-2018, ArguAna, CQADupstack, and DBPedia.","rationale":"The reader's weakest assumption was translation fidelity. While that is a legitimate concern, the most load-bearing issue is internal: the abstract's central performance claim contradicts the paper's own Table 3. This is independent of any external assumptions about translation quality or contamination flags. The dataset release may still be valuable, and the issue is fixable by rephrasing the claim, so a conditional acceptance remains appropriate. However, the current abstract and conclusions should not be published without correction.","tokens_in":13492,"tokens_out":6760,"duration_ms":63608,"concrete_test":"Recompute the nDCG@10 comparison on FiQA-2018, ArguAna, CQADupstack, and DBPedia using the released BEIR-NL data with the stated BM25 (Elasticsearch) and e5-multilingual-small configurations; if e5-small's scores remain above BM25, the abstract's exclusion of small retrieval-trained dense models is false and the claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, as stated in the abstract, is that 'BM25 remains a competitive baseline, and is only outperformed by the larger dense models trained for retrieval.' This is not supported by the paper's own results. In Table 3, e5-multilingual-small (118M parameters, fine-tuned for retrieval, and not flagged with a contamination dagger on FiQA-2018) scores 20.39 nDCG@10 on FiQA-2018 versus 18.73 for BM25; it also outperforms BM25 on ArguAna (44.76 vs 41.76), CQADupstack (28.51 vs 27.77), and DBPedia (25.89 vs 25.46). These are small, not larger, dense models. Section 5.1 uses the more modest phrasing 'in many cases,' but the abstract overstates the finding. The false claim is load-bearing because the paper's headline contribution includes this performance ranking; it is an internal inconsistency that does not depend on translation-quality assumptions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BEIR-NL, a Dutch-language zero-shot information retrieval benchmark created by automatically translating 14 publicly available BEIR datasets from English to Dutch, mainly with Gemini-1.5-flash. The authors evaluate BM25, eight multilingual dense ranking models, and three reranking models, reporting nDCG@10 and Recall@100 on all datasets plus Dutch mMARCO. They also report a small human quality check of the translations and a back-translation experiment on five datasets. The main claims are that BEIR-NL is a usable resource for Dutch IR evaluation and that BM25 remains competitive, being outperformed only by larger dense retrieval models, with BM25+reranking matching the best dense rankers.","tokens_in":13671,"tokens_out":4611,"duration_ms":48384,"significance":"If the benchmark holds up, this is a useful and timely resource for Dutch IR, a language that is underrepresented in retrieval evaluation. The authors release the data on Hugging Face, follow BEIR conventions, compare with BEIR-PL and BEIR, and are transparent about license inheritance and potential contamination. The evaluation is standard and the resource is likely to be reused. However, the paper's headline claims about the BM25 comparison are overstated, and the evidence for translation quality is thin; both issues need to be fixed before the paper can be relied upon as a benchmark paper.","major_comments":[{"comment":"The abstract's claim that BM25 'is only outperformed by the larger dense models trained for retrieval' is contradicted by the paper's own results. In Table 3, multilingual-e5-small (118M parameters, retrieval-finetuned, and not flagged with a contamination dagger on these datasets) outperforms BM25 on FiQA-2018 (20.39 vs 18.73), ArguAna (44.76 vs 41.76), CQADupstack (28.51 vs 27.77), and DBPedia (25.89 vs 25.46). Since this performance ranking is part of the paper's central contribution, the wording must be changed to reflect the actual pattern, for example by saying BM25 is outperformed by most or many retrieval-trained dense models, or by giving the exceptions explicitly.","section":"Abstract, Section 5.1, Table 3"},{"comment":"The evidence supporting the benchmark's reliability is currently too weak for the strength of the paper's claims. The translation quality check uses only 140 items (10 per dataset) with a single annotator, and the reported 2.2% major issues corresponds to just three items; this yields a very wide confidence interval. The back-translation experiment covers only 5 of the 14 datasets and only one dense model. I recommend reporting the confidence interval for the quality estimate, ideally adding a second annotator or a larger sample, and tempering statements such as 'almost 98% of the translated samples can be trusted' so that they do not overstate the precision of the estimate.","section":"Section 3.1, Section 5.3"},{"comment":"The translation prompts in Appendix B instruct the model to 'Translate to English', yet Section 3.1 states that the pipeline translates from English to Dutch. If this is a typographical error it should be corrected; if the prompts were actually used as written, the resulting data would not be Dutch. Either way, the appendix must be fixed because the prompt template is essential for reproducibility.","section":"Appendix B"},{"comment":"The paper repeatedly describes the evaluations as zero-shot even though Table 3 marks most of the top-performing dense models with a dagger indicating likely in-domain contamination, and the Limitations section acknowledges that these results may not be proper zero-shot. The text should more clearly separate contaminated from uncontaminated rows and specify that the benchmark itself is zero-shot for future models, while some of the reported numbers are not zero-shot evaluations. This distinction matters because the conclusion that larger dense models outperform BM25 rests substantially on daggered numbers.","section":"Section 5.1, Section 6, Limitations"}],"minor_comments":[{"comment":"The 'IR Finetuned' column contains the typo 'Y es' for several models; it should read 'Yes'.","section":"Table 2"},{"comment":"The first column header 'BEIR' is confusing because the table reports results on original BEIR and back-translated data; renaming it to 'BEIR (EN)' would match Table 4 and improve clarity.","section":"Table 6"},{"comment":"Footnote 8 says 'Assuming a uniform BM25 performance for different languages, which is not trivial'; this is an important caveat and should be moved into the main text rather than relegated to a footnote.","section":"Section 5.2"},{"comment":"The Hendrycks et al. reference for MMLU lacks a year and venue; please complete it. Also, several author names in the bibliography contain spacing artifacts such as 'Y ang', 'Y an', and 'T worek', which should be cleaned.","section":"References"},{"comment":"The phrase 'less than 450 Euro' should include the currency symbol and, ideally, a note on the exchange rate or date, to give readers a better sense of the cost.","section":"Section 3.1"},{"comment":"The table lists mMARCO as a dataset while the text says it is not translated in this work; a parenthetical note in the caption or table would avoid confusion for readers who only look at the table.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The benchmark release is genuinely useful and the evaluation is broadly sound, so I do not see grounds for rejection. The main issues are an overstatement of the BM25 comparison that is directly contradicted by Table 3, and a thin translation-quality validation that needs more careful reporting. Both are fixable within the manuscript's scope. The paper fits the journal's scope well."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The benchmark release is the real contribution here. BEIR-NL gives Dutch IR a standardized 14-dataset evaluation suite, it is publicly available, and the authors are honest about contamination and translation limitations. The back-translation control on five datasets is a sensible extra, and the comparison to BEIR-PL helps position the resource. This is exactly the kind of translated-benchmark work that mid-resource language communities need. It does not break new methodological ground, but it does fill a concrete gap.\n\nThe soft spots are real but mostly fixable. The biggest issue is the abstract's claim that BM25 'is only outperformed by the larger dense models trained for retrieval.' Table 3 shows multilingual-e5-small (118M parameters, retrieval-finetuned, not contamination-flagged on FiQA-2018) beating BM25 on FiQA-2018, ArguAna, CQADupstack, and DBPedia. Those are not larger models; they are smaller. The body text is more careful ('in many cases'), so the abstract overstates the head finding, and since the abstract is what most readers will see, that overstatement needs correcting. It is a load-bearing inconsistency because the title and abstract set up BM25's relative performance as a headline result.\n\nThe translation-quality check is 140 samples with a single annotator, reporting 2.2% major and 14.8% minor issues. That is a thin basis for the '98% can be trusted' gloss. The independent translation of queries and passages is flagged by the authors as a source of lexical mismatch, and the BM25 drop on back-translation (1.9 points) supports that concern. The paper would be stronger with a second annotator, more samples, and ideally a small native-Dutch human relevance judgment set for at least one dataset. Also, no error bars or significance tests on the model comparisons; for a resource paper that is acceptable, but worth stating.\n\nThe contamination daggers are a good practice and the authors deserve credit for flagging them rather than hiding them. The zero-shot framing is therefore conditional on those flags, which the body acknowledges, though the abstract still uses the unqualified phrase 'zero-shot.'\n\nOverall, the benchmark is credible and the experiments are reproducible in principle. The central resource stands. The abstract needs to be aligned with the actual numbers, and the translation validation should be strengthened. This paper deserves a serious referee and likely a conditional accept after revision.\n\nI would bring it to a reading group only if someone is working on multilingual IR or Dutch NLP; otherwise it is a solid resource paper rather than a conceptual advance.","headline":"Useful Dutch BEIR resource, but the abstract's BM25 claim is contradicted by the paper's own Table 3 and the translation validation is thinner than it looks.","tokens_in":14198,"tokens_out":1675,"would_cite":true,"duration_ms":19356,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces BEIR-NL, a machine-translated Dutch version of 14 BEIR datasets, and reports that BM25 remains a competitive zero-shot baseline, outperformed only by larger dense models trained specifically for retrieval.","keywords":["information retrieval","zero-shot evaluation","Dutch language","BEIR","BM25","dense retrieval","reranking","machine translation benchmark"],"falsifier":"A concrete check would be to have native Dutch speakers write natural Dutch queries for a subset of BEIR-NL topics and compare model rankings on these queries against rankings on the machine-translated queries; if the rankings diverge substantially or if a larger human-annotated sample finds major translation errors well above the reported 2.2%, the benchmark is measuring translation artifacts rather than Dutch retrieval ability.","tokens_in":13290,"feed_emoji":"🔍","tokens_out":9768,"duration_ms":86188,"temperature":0.7,"pith_summary":"The paper attempts to establish that a Dutch information retrieval benchmark can be built by automatically translating the English BEIR datasets, and that this translated benchmark is usable for zero-shot evaluation—gauging retrieval models on Dutch tasks they were not trained on. It translates the 14 publicly available BEIR datasets into Dutch, releases them as BEIR-NL, and evaluates a spread of lexical, dense, and reranking models. The main empirical claim is that BM25 keyword search remains a competitive baseline on the Dutch data, clearly outperformed only by the larger dense models trained specifically for retrieval; pairing BM25 with a multilingual reranker reaches performance comparable to the best dense rankers. A back-translation check on five datasets shows a small but consistent performance drop, which the authors take as evidence that translation itself reduces benchmark quality. The authors present BEIR-NL as a public resource for Dutch IR research while noting that native Dutch gold data is still needed for full fidelity.","feed_headline":"New Dutch IR benchmark puts BM25 above most dense models","feed_subtitle":"A machine-translated version of 14 BEIR datasets gives Dutch zero-shot IR evaluation; reranked BM25 matches the top dense models.","key_machinery":"The central object is BEIR-NL, a machine-translated Dutch mirror of the 14 publicly available datasets from the BEIR benchmark, spanning biomedical, Wikipedia, financial, scientific, argument, and question-answer retrieval tasks. The evaluation machinery is the standard BEIR zero-shot protocol: BM25 as the lexical baseline, dense bi-encoder models that score query-document pairs by cosine similarity on normalized embeddings, and cross-encoder rerankers applied to the top-100 documents retrieved by BM25, with nDCG@10 and Recall@100 as metrics. Translation is performed by a commercial LLM-based API with queries and documents translated independently, and quality is checked through a small human-annotated sample and a five-dataset back-translation control that isolates translation loss from model-language competence.","core_discovery":"The paper's discovery, stated on its own terms, is that a machine-translated benchmark can reproduce the shape of the English BEIR evaluation landscape in Dutch: on the ten overlapping datasets, BM25 reaches 35.9 average nDCG@10 against 41.9 for the original English BEIR, and the top spots go to the larger retrieval-trained dense models (notably multilingual-e5-large-instruct) and to BM25 combined with cross-encoder rerankers. The paper also establishes that translation carries a measurable cost: back-translating a five-dataset subset from Dutch to English lowers nDCG@10 by about 1.9 points for BM25 and 2.6 points for gte-multilingual-base, which the authors attribute to lexical mismatches caused by translating queries and passages independently. The overall pattern is that older sentence-embedding models trail BM25 on Dutch data, while only the new generation of retrieval-trained dense models clearly surpasses it.","pith_inferences":["If the translation penalty is roughly uniform across models, the model ordering found on BEIR-NL likely transfers to real Dutch IR, making the benchmark useful for model selection even if absolute scores are pessimistic.","The five-dataset back-translation protocol can serve as a reusable quality control for any future translated benchmark: a small or zero delta would indicate the translation pipeline is not the main source of performance loss.","The result that reranked BM25 matches the best dense models suggests that lexical recall in Dutch is not the bottleneck, so Dutch-specific rerankers or query-expansion methods may yield larger gains than scaling dense encoders.","A native Dutch gold benchmark built from the same relevance judgments would separate translation loss from model-language competence, which the paper explicitly leaves to future work."],"forward_implications":["Dutch IR models can be compared zero-shot across 14 tasks and multiple domains on a single public benchmark, filling a gap for a language with few native IR test collections.","For practical Dutch retrieval, BM25 followed by a multilingual reranker is a strong recipe that matches the best dense ranking models, so teams without the largest dense encoders are not at a major disadvantage.","Translated benchmarks carry a translation penalty: scores on BEIR-NL are several points lower than on English BEIR and drop further under back-translation, so cross-lingual numbers should not be read as exact native-language performance.","Evaluations on BEIR-NL need to account for training contamination, since several of the top dense models have likely seen BEIR data during training, which may inflate their zero-shot scores."],"supporting_citations":[{"why":"Defines the original BEIR benchmark and the zero-shot evaluation protocol that BEIR-NL translates and follows.","marker":"Thakur et al., 2021"},{"why":"Provides the Dutch-translated MSMARCO used in the evaluations and identifies the issue of independently translated queries and passages.","marker":"Bonifacio et al., 2021"},{"why":"Provides BEIR-PL, the parallel Polish translated benchmark used for cross-language comparison, and the BM25 setup the authors reuse.","marker":"Wojtasik et al., 2024"},{"why":"Defines the BM25 lexical retrieval method that serves as the paper's baseline.","marker":"Robertson et al., 1994"},{"why":"Introduces the multilingual E5 models, including multilingual-e5-large-instruct, which achieves the highest Recall@100 on half the datasets.","marker":"Wang et al., 2024"},{"why":"Introduces BGE-M3, one of the larger retrieval-trained dense models that outperform BM25 in the evaluation.","marker":"Chen et al., 2024"},{"why":"Introduces the gte-multilingual models used for the dense model comparison and for the back-translation experiment.","marker":"Zhang et al., 2024"}],"fun_headline_variants":["Dutch IR benchmark: BM25 beats most dense models","Machine-translated BEIR shows BM25 still strong for Dutch","BM25 outperforms many dense models on new Dutch IR test","Back-translation costs accuracy in Dutch IR benchmark","New Dutch benchmark: Reranked BM25 matches top dense models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's usefulness rests on the assumption that automatic translation preserves enough semantic fidelity that retrieval scores reflect Dutch-language competence rather than translation artifacts; the evidence offered is a 140-item human quality check and a five-dataset back-translation proxy, not native Dutch gold labels.","fun_headline_variants_meta":{"raw":{"variants":["Dutch IR benchmark: BM25 beats most dense models","Machine-translated BEIR shows BM25 still strong for Dutch","BM25 outperforms many dense models on new Dutch IR test","Back-translation costs accuracy in Dutch IR benchmark","New Dutch benchmark: Reranked BM25 matches top dense models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000463,"raw_usage":{"total_tokens":2317,"prompt_tokens":948,"completion_tokens":1369,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":1287}},"tokens_in":564,"tokens_out":1369,"duration_ms":11069,"temperature":1.0,"reasoning_tokens":1287,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:55:24.177271+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check would be to have native Dutch speakers write natural Dutch queries for a subset of BEIR-NL topics and compare model rankings on these queries against rankings on the machine-translated queries; if the rankings diverge substantially or if a larger human-annotated sample finds major translation errors well above the reported 2.2%, the benchmark is measuring translation artifacts rather than Dutch retrieval ability.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides BEIR-PL, the parallel Polish translated benchmark used for cross-language comparison, and the BM25 setup the authors reuse."}],"review_version":1}