Pith. sign in

REVIEW 13 cited by

KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.01028 v4 pith:DBHOAQEY submitted 2025-01-02 cs.CL

KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

classification cs.CL
keywords embeddingmodelmodelsdatatraininggeneralkalm-embeddinglanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As retrieval-augmented generation prevails in large language models, embedding models are becoming increasingly crucial. Despite the growing number of general embedding models, prior work often overlooks the critical role of training data quality. In this work, we introduce KaLM-Embedding, a general multilingual embedding model that leverages a large quantity of cleaner, more diverse, and domain-specific training data. Our model has been trained with key techniques proven to enhance performance: (1) persona-based synthetic data to create diversified examples distilled from LLMs, (2) ranking consistency filtering to remove less informative samples, and (3) semi-homogeneous task batch sampling to improve training efficacy. Departing from traditional BERT-like architectures, we adopt Qwen2-0.5B as the pre-trained model, facilitating the adaptation of auto-regressive language models for general embedding tasks. Extensive evaluations of the MTEB benchmark across multiple languages show that our model outperforms others of comparable size, setting a new standard for multilingual embedding models with <1B parameters.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

    cs.AI 2026-06 unverdicted novelty 7.0

    RealMath-Eval benchmark shows LLM judges have an evaluation gap, performing worse on diverse real human math reasoning than on synthetic solutions due to greater error diversity and higher surprisal.

  2. SEA-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

    cs.CL 2026-06 unverdicted novelty 7.0

    SEA-Embedding is a fully open text embedding pipeline for Southeast Asian languages that achieves state-of-the-art performance on the SEA-BED benchmark by analyzing data composition, training objectives, and base enco...

  3. Prism-Reranker: Beyond Relevance Scoring -- Jointly Producing Contributions and Evidence for Agentic Retrieval

    cs.IR 2026-04 accept novelty 7.0

    Prism-Reranker models output relevance, contribution statements, and evidence passages to support agentic retrieval beyond scalar scoring.

  4. LMEB: Long-horizon Memory Embedding Benchmark

    cs.CL 2026-03 unverdicted novelty 7.0

    LMEB benchmark shows that embedding models' performance on traditional retrieval does not transfer to long-horizon memory tasks, larger models do not always perform better, and LMEB measures capabilities orthogonal to MTEB.

  5. KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking

    cs.CL 2026-06 accept novelty 6.0

    KaLM-Reranker-V1 uses encoder–decoder FBNL with Matryoshka pooling to match Qwen3-class reranking quality at substantially lower online cost.

  6. Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

    cs.CL 2026-05 unverdicted novelty 6.0

    Embedding model performance on MTEB tasks correlates strongly with nearest-neighbor overlap and ICA magnitude differences in their embedding spaces.

  7. LMEB: Long-horizon Memory Embedding Benchmark

    cs.CL 2026-03 conditional novelty 6.0

    LMEB is a new benchmark that evaluates embedding models on long-horizon memory retrieval and shows this skill is largely orthogonal to traditional passage-retrieval performance.

  8. LMEB: Long-horizon Memory Embedding Benchmark

    cs.CL 2026-03 conditional novelty 6.0

    LMEB is a new benchmark of 193 retrieval tasks spanning episodic, dialogue, semantic, and procedural memory on which top embedding models score about 61 NDCG@10, largely uncorrelated with MTEB.

  9. LMEB: Long-horizon Memory Embedding Benchmark

    cs.CL 2026-03 unverdicted novelty 6.0

    LMEB is a 22-dataset, 193-task zero-shot benchmark showing that long-horizon memory retrieval is hard, not solved by scale, and largely orthogonal to MTEB passage-retrieval skill.

  10. SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension

    cs.CL 2025-08 unverdicted novelty 6.0

    SitEmb-v1.5 uses a new training paradigm to produce context-situated embeddings for short chunks, outperforming larger models by over 10% on a curated book-plot retrieval benchmark.

  11. KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking

    cs.CL 2026-06 unverdicted novelty 5.0

    KaLM-Reranker-V1 introduces a fast but not late-interaction reranker that decouples passage pre-encoding from query processing via encoder-decoder architecture and cross-attention to achieve efficiency and competitive...

  12. LRanker: LLM Ranker for Massive Candidates

    cs.IR 2026-05 unverdicted novelty 5.0

    LRanker combines K-means candidate aggregation with graph-partitioned ensemble of query embeddings to improve LLM ranking accuracy and scalability on massive candidate pools, reporting 3-30% gains on RBench tasks up t...

  13. Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

    cs.IR 2026-05 unverdicted novelty 5.0

    Multi-task evaluation of 22 patent embedding models finds task-specific fine-tuning benefits and significant cross-landscape retrieval degradation that cannot be fixed by hybrid fusion.