Pith. sign in

Sonar: Sentence-level multimodal and language-agnostic representations

10 Pith papers cite this work. Polarity classification is still indexing.

10 Pith papers citing it
abstract

We introduce SONAR, a new multilingual and multimodal fixed-size sentence embedding space. Our single text encoder, covering 200 languages, substantially outperforms existing sentence embeddings such as LASER3 and LabSE on the xsim and xsim++ multilingual similarity search tasks. Speech segments can be embedded in the same SONAR embedding space using language-specific speech encoders trained in a teacher-student setting on speech transcription data. Our encoders outperform existing speech encoders on similarity search tasks. We also provide a text decoder for 200 languages, which allows us to perform text-to-text and speech-to-text machine translation, including for zero-shot language and modality combinations. Our text-to-text results are competitive compared to the state-of-the-art NLLB~1B model, despite the fixed-size bottleneck representation. Our zero-shot speech-to-text translation results compare favorably with strong supervised baselines such as Whisper.

citation-role summary

background 1 dataset 1

citation-polarity summary

years

2026 9 2024 1

representative citing papers

The Faiss library

cs.LG · 2024-01-16 · unverdicted · novelty 3.0

Faiss is a library offering indexing methods and primitives for efficient vector similarity search, a core need in vector databases for AI applications.

citing papers explorer

Showing 10 of 10 citing papers.