Hubness dominates as the causal driver of cross-lingual retrieval asymmetry in multilingual embeddings, with CSLS correction closing most of the reciprocity gap across five models and a 6518-item parallel corpus.
Language-agnostic BERT sentence embedding
9 Pith papers cite this work, alongside 196 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Machine translation from English to Polish preserves enough moral-semantic signal that classifiers trained on translated text reach near-English accuracy (AUROC gaps mostly 0.01–0.05).
Embedding model performance on MTEB tasks correlates strongly with nearest-neighbor overlap and ICA magnitude differences in their embedding spaces.
A latent variable IRT framework decouples four safety-driving factors across 61 model configurations and 10 languages using 1.9 million evaluations, revealing that safety is largely unidimensional and that high cross-lingual gaps cluster in physical harm prompts and lower-resource languages.
Concept Separation Curves provide a classifier-independent method to visualize and quantify how sentence embeddings distinguish conceptual meaning from syntactic variations across languages and domains.
MathNet compiles 30,676 multilingual Olympiad problems with solutions into a benchmark showing top LLMs score 69–78% while embedding retrievers rarely find mathematically equivalent problems at rank 1.
BLOOM is a 176B-parameter open-access multilingual language model trained on the ROOTS corpus that achieves competitive performance on benchmarks, with improved results after multitask prompted finetuning.
A new pre-training task that maps languages bidirectionally in embedding space improves machine translation by up to 11.9 BLEU, cross-lingual QA by 6.72 BERTScore points, and understanding accuracy by over 5% over strong baselines.
Larger 100K vocabularies in SPLADE models, especially those initialized with ESPLADE pretraining, improve retrieval effectiveness after pruning compared to 32K baselines while keeping similar efficiency.
citing papers explorer
-
Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models
Hubness dominates as the causal driver of cross-lingual retrieval asymmetry in multilingual embeddings, with CSLS correction closing most of the reciprocity gap across five models and a 6518-item parallel corpus.
-
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
Machine translation from English to Polish preserves enough moral-semantic signal that classifiers trained on translated text reach near-English accuracy (AUROC gaps mostly 0.01–0.05).
-
Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
Embedding model performance on MTEB tasks correlates strongly with nearest-neighbor overlap and ICA magnitude differences in their embedding spaces.
-
Why Do Safety Guardrails Degrade Across Languages?
A latent variable IRT framework decouples four safety-driving factors across 61 model configurations and 10 languages using 1.9 million evaluations, revealing that safety is largely unidimensional and that high cross-lingual gaps cluster in physical harm prompts and lower-resource languages.
-
Finding Meaning in Embeddings: Concept Separation Curves
Concept Separation Curves provide a classifier-independent method to visualize and quantify how sentence embeddings distinguish conceptual meaning from syntactic variations across languages and domains.
-
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
MathNet compiles 30,676 multilingual Olympiad problems with solutions into a benchmark showing top LLMs score 69–78% while embedding retrievers rarely find mathematically equivalent problems at rank 1.
-
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BLOOM is a 176B-parameter open-access multilingual language model trained on the ROOTS corpus that achieves competitive performance on benchmarks, with improved results after multitask prompted finetuning.
-
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
A new pre-training task that maps languages bidirectionally in embedding space improves machine translation by up to 11.9 BLEU, cross-lingual QA by 6.72 BERTScore points, and understanding accuracy by over 5% over strong baselines.
-
The Role of Vocabularies in Learning Sparse Representations for Ranking
Larger 100K vocabularies in SPLADE models, especially those initialized with ESPLADE pretraining, improve retrieval effectiveness after pruning compared to 32K baselines while keeping similar efficiency.