Pith. sign in

REVIEW 22 cited by

Arctic-Embed 2.0: Multilingual Retrieval Without Compromise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04506 v2 pith:LRVJHGFP submitted 2024-12-03 cs.CL cs.IRcs.LG

Arctic-Embed 2.0: Multilingual Retrieval Without Compromise

classification cs.CL cs.IRcs.LG
keywords retrievalarctic-embedmultilingualqualitydiscussionefficientembeddingquestions
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper presents the training methodology of Arctic-Embed 2.0, a set of open-source text embedding models built for accurate and efficient multilingual retrieval. While prior works have suffered from degraded English retrieval quality, Arctic-Embed 2.0 delivers competitive retrieval quality on multilingual and English-only benchmarks, and supports Matryoshka Representation Learning (MRL) for efficient embedding storage with significantly lower compressed quality degradation compared to alternatives. We detail the design and implementation, presenting several important open research questions that arose during model development. We conduct experiments exploring these research questions and include extensive discussion aimed at fostering further discussion in this field.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

    cs.CL 2026-06 unverdicted novelty 8.0

    SkMTEB is the first comprehensive text embedding benchmark for Slovak, and vocabulary-trimmed E5 adaptations achieve competitive performance with much smaller models.

  2. DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    cs.CL 2026-07 conditional novelty 7.0

    Openly trained late-interaction retrieval models transfer to unseen languages and scripts far better than dense models under identical translate-train data.

  3. Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders

    cs.IR 2026-07 conditional novelty 7.0

    Bekko a8m, with 7.7M active parameters, scores 56.2 on MMTEB Multilingual v2 Retrieval, beating mE5 models and BGE-M3, while a25m reaches 57.5, on par with gte-multilingual-base.

  4. Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

    cs.CR 2026-05 unverdicted novelty 7.0

    MemPoison enables stealthy memory poisoning in LLM agents via dialogue by using semantic relational bridges, entity masquerading, and joint embedding optimization to bypass selective extraction and rewriting, achievin...

  5. DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    cs.CL 2026-07 conditional novelty 6.0

    With matched open data and backbones, ColBERT-style late interaction turns English translate-train into multilingual generalization, while dense retrieval stays mostly inside the translated languages.

  6. A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

    cs.CL 2026-07 conditional novelty 6.0

    CUP, a 104-query Greek book retrieval benchmark, shows hybrid lexical-semantic retrieval outperforms both BM25 and dense-only methods in a real publisher catalog.

  7. Larch: Learned Query Optimization for Semantic Predicates

    cs.DB 2026-06 unverdicted novelty 6.0

    Larch uses a GNN-MDP formulation and a selectivity predictor plus dynamic programming to reorder semantic filter evaluation, cutting token usage 3x-19x versus prior systems on real and synthetic workloads.

  8. Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

    cs.CL 2026-05 unverdicted novelty 6.0

    Embedding model performance on MTEB tasks correlates strongly with nearest-neighbor overlap and ICA magnitude differences in their embedding spaces.

  9. Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs

    cs.LG 2026-05 unverdicted novelty 6.0

    LRD framework with Frenet, NRS, and GFMI metrics shows layer-wise structure in 31 models provides usable signal for model selection and pruning on MTEB tasks.

  10. MLAIRE: Multilingual Language-Aware Information Retrieval Evaluation Protocal

    cs.IR 2026-05 unverdicted novelty 6.0

    MLAIRE is a protocol that evaluates multilingual retrievers on both semantic accuracy and query-language preference using parallel passages and new metrics like LPR and Lang-nDCG, showing that standard metrics hide di...

  11. Identifier-Free Code Embedding Models for Scalable Search

    cs.CR 2026-05 unverdicted novelty 6.0

    A fine-tuned Qwen3-Embedding model with contrastive learning outperforms baselines on bidirectional source-to-decompiled code association and generalizes to constant-algorithm tasks.

  12. CORE-T: COherent REtrieval of Tables for Text-to-SQL

    cs.CL 2026-01 conditional novelty 6.0

    CORE-T uses LLM purpose metadata, a compatibility cache, and a single LLM call to select coherent joinable table sets, improving open-book multi-table text-to-SQL retrieval and execution accuracy.

  13. LLMs Meet Isolation Kernel: Lightweight, Learning-free Binary Embeddings for Fast Retrieval

    cs.IR 2026-01 unverdicted novelty 6.0

    IKE converts LLM embeddings into binary codes via Isolation Kernel for up to 16.7x faster retrieval and 16x lower memory with comparable accuracy.

  14. LLMs Meet Isolation Kernel: Lightweight, Learning-free Binary Embeddings for Fast Retrieval

    cs.IR 2026-01 conditional novelty 6.0

    Isolation-kernel binary hashing (IKE) compresses LLM embeddings to a few hundred bytes per point with retrieval accuracy near the original and large speedups in exhaustive and ANN search.

  15. LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations

    cs.IR 2025-09 conditional novelty 6.0

    LEAF distills teacher-aligned student embedding models that achieve new SOTA results on BEIR and MTEB for their size class while requiring only modest data and compute.

  16. Grounding Text Embeddings in Stakeholder Associations

    cs.CL 2026-05 unverdicted novelty 5.0

    The Stakeholder Grounding Exercise shows neural text embeddings are 19-26pp less reliable than human experts at capturing semantic distinctions, with misalignment strongly correlated to poorer clustering performance (...

  17. jina-embeddings-v5-text: Task-Targeted Embedding Distillation

    cs.CL 2026-02 unverdicted novelty 5.0

    A distillation-plus-task-contrastive training regimen yields compact embedding models that match or exceed state-of-the-art performance for their size while supporting 32k-token contexts and quantization.

  18. Generating consensus and dissent on massive discussion platforms with a semantic-vector model

    physics.soc-ph 2026-01 conditional novelty 5.0

    A semantic-vector O(N) model on a 2D lattice generates consensus (β>0) or maximum dissent (β<0) by local copying of neighbor responses.

  19. Retrofitting Small Multilingual Models for Retrieval: Matching 7B Performance with 300M Parameters

    cs.CL 2025-10 conditional novelty 5.0

    A 300M multilingual embedding model matches or exceeds 7B retrieval performance via optimized data scale, hard negatives, and task diversity over language diversity.

  20. MimirRAG: A Multi-Agent RAG Framework for Financial Data Retrieval with Metadata Integration

    cs.LG 2026-05 unverdicted novelty 4.0

    MimirRAG, a multi-agent RAG framework with metadata integration and table-aware chunking, reaches 89.3% accuracy on FinanceBench and outperforms prior baselines for financial document retrieval.

  21. Comparison of Modern Multilingual Text Embedding Techniques for Hate Speech Detection Task

    cs.CL 2026-04 unverdicted novelty 4.0

    Supervised models using embeddings like jina and e5 reach up to 92% accuracy on multilingual hate speech detection, substantially outperforming anomaly detection, while PCA to 64 dimensions preserves most performance ...

  22. Federated Learning for ICD Classification with Lightweight Models and Pretrained Embeddings

    cs.IR 2025-07 unverdicted novelty 4.0

    Lightweight federated learning with frozen embeddings and MLP heads reaches competitive micro and macro F1 scores for ICD-9 and ICD-10 coding on MIMIC-IV, nearly matching centralized training.