ALEE generates AMR-based English minimal pairs with fine-grained semantic shifts, translates them, and evaluates embedding models on 275+ languages to expose cross-lingual gaps linked to training data and tokenization.
mGTE : Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
12 Pith papers cite this work, alongside 74 external citations. Polarity classification is still indexing.
years
2026 12representative citing papers
Introduces P-CHR AUC and CRR metrics to demonstrate that semantic caching model selection is limited by calibration quality rather than ranking performance.
HTEB introduces dynamic, multi-axis evaluation of text embedding robustness using LLM transformations, finding decoupled profiles across models and that scaling does not close all robustness gaps.
Dual-view training via polarity reversal improves instruction-following retrieval performance by 45% on the FollowIR benchmark using a 305M-parameter encoder.
XBCP benchmark shows deep research agents and multilingual retrievers lose accuracy, recall, calibration, and citation reliability when evidence is in non-English languages, even with gold evidence provided.
Optimal interpolation of query embeddings from parallel translations outperforms the best monolingual query in 88/105 cases on mMARCO, showing English-driven asymmetry and negative correlation with typological distance.
Inference-time attention calibration with partial strength reduces positional bias in dense retrievers on SQuAD-PosQ, FineWeb-PosQ and PosIR while preserving nDCG@10.
COMPASS uses semantic clustering on multilingual embeddings to select auxiliary data for PEFT adapters, outperforming linguistic-similarity baselines on multilingual benchmarks while supporting continual adaptation.
PETRA is a curated 1.36M-chunk petroleum-engineering retrieval dataset and pipeline that raises in-domain nDCG from 0.703 to 0.763 via score fusion and delivers 44% relative gain on an Earth Science benchmark through reranker adaptation on synthetic supervision.
Large-scale benchmarks of multilingual embeddings and QE models show no universal performer; direction-aware routing and calibration recommended for parallel data assessment.
An adaptive thresholding mechanism combined with sliding-window reranking retrieves a query-dependent number of tables from large corpora, improving retrieval and downstream text-to-SQL performance on Spider, BIRD, and Spider 2.0.
citing papers explorer
-
ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs
ALEE generates AMR-based English minimal pairs with fine-grained semantic shifts, translates them, and evaluates embedding models on 275+ languages to expose cross-lingual gaps linked to training data and tokenization.
-
Closing the Calibration Gap in Semantic Caching
Introduces P-CHR AUC and CRR metrics to demonstrate that semantic caching model selection is limited by calibration quality rather than ranking performance.
-
The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness
HTEB introduces dynamic, multi-axis evaluation of text embedding robustness using LLM transformations, finding decoupled profiles across models and that scaling does not close all robustness gaps.
-
Dual-View Training for Instruction-Following Information Retrieval
Dual-view training via polarity reversal improves instruction-following retrieval performance by 45% on the FollowIR benchmark using a 305M-parameter encoder.
-
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus
XBCP benchmark shows deep research agents and multilingual retrievers lose accuracy, recall, calibration, and citation reliability when evidence is in non-English languages, even with gold evidence provided.
-
When Does Mixing Help? Analyzing Query Embedding Interpolation in Multilingual Dense Retrieval
Optimal interpolation of query embeddings from parallel translations outperforms the best monolingual query in 88/105 cases on mMARCO, showing English-driven asymmetry and negative correlation with typological distance.
-
Attention Calibration for Position-Fair Dense Information Retrieval
Inference-time attention calibration with partial strength reduces positional bias in dense retrievers on SQuAD-PosQ, FineWeb-PosQ and PosIR while preserving nDCG@10.
-
COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling
COMPASS uses semantic clustering on multilingual embeddings to select auxiliary data for PEFT adapters, outperforming linguistic-similarity baselines on multilingual benchmarks while supporting continual adaptation.
-
PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation
PETRA is a curated 1.36M-chunk petroleum-engineering retrieval dataset and pipeline that raises in-domain nDCG from 0.703 to 0.763 via score fusion and delivers 44% relative gain on an Earth Science benchmark through reranker adaptation on synthetic supervision.
-
Model-Based Quality Assessment for Massively Multilingual Parallel Data
Large-scale benchmarks of multilingual embeddings and QE models show no universal performer; direction-aware routing and calibration recommended for parallel data assessment.
-
Retrieve Only Relevant Tables Whether Few or Many: Adaptive Table Retrieval Method
An adaptive thresholding mechanism combined with sliding-window reranking retrieves a query-dependent number of tables from large corpora, improving retrieval and downstream text-to-SQL performance on Spider, BIRD, and Spider 2.0.
- Comparing Architectures for Supervised Political Scaling