Current VLMs depend on tightly aligned curated data and cannot exploit the weakly-aligned egocentric video signals that dominate naturalistic infant input.
Sonar: Sentence-level multimodal and language-agnostic representations
10 Pith papers cite this work. Polarity classification is still indexing.
abstract
We introduce SONAR, a new multilingual and multimodal fixed-size sentence embedding space. Our single text encoder, covering 200 languages, substantially outperforms existing sentence embeddings such as LASER3 and LabSE on the xsim and xsim++ multilingual similarity search tasks. Speech segments can be embedded in the same SONAR embedding space using language-specific speech encoders trained in a teacher-student setting on speech transcription data. Our encoders outperform existing speech encoders on similarity search tasks. We also provide a text decoder for 200 languages, which allows us to perform text-to-text and speech-to-text machine translation, including for zero-shot language and modality combinations. Our text-to-text results are competitive compared to the state-of-the-art NLLB~1B model, despite the fixed-size bottleneck representation. Our zero-shot speech-to-text translation results compare favorably with strong supervised baselines such as Whisper.
citation-role summary
citation-polarity summary
representative citing papers
Four axioms (Causality, Minimality, Separability, Stability) are formalized for latent thought representations; audits of open LLMs on 23 tasks show none satisfy all four and representations add little beyond input embeddings.
A linguistic-invariant spoofing detection method using teacher-student gradient reversal and VIB achieves up to 36.2% relative EER reduction across nine DF Arena datasets.
Interventional contrastive learning applied after pre-training disentangles speaker and content information in speech foundation model embeddings, yielding better out-of-domain speaker verification.
VLMs recover reliable population-level trends in climate change visual discourse on social media even when per-image accuracy is only moderate.
A factorized log-linear model (FLiP) recovers 75-80% of lexical content from sentence embeddings and reveals English-centric biases in SONAR, LaBSE, and Gemini encoders.
HydraQE is a new end-to-end speech translation QE system using Qwen3-ASR backbone, sparsemax layer mixing, bidirectional Transformer, and multi-task curriculum training on human and pseudo labels that outperforms cascaded baselines.
The authors built and publicly released sentence-aligned simplification corpora for five languages by processing crowd-sourced data from comparable documents.
Faiss is a library offering indexing methods and primitives for efficient vector similarity search, a core need in vector databases for AI applications.
citing papers explorer
-
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
Current VLMs depend on tightly aligned curated data and cannot exploit the weakly-aligned egocentric video signals that dominate naturalistic infant input.
-
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Four axioms (Causality, Minimality, Separability, Stability) are formalized for latent thought representations; audits of open LLMs on 23 tasks show none satisfy all four and representations add little beyond input embeddings.
-
Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck
A linguistic-invariant spoofing detection method using teacher-student gradient reversal and VIB achieves up to 36.2% relative EER reduction across nine DF Arena datasets.
-
Learning task-specific subspaces via interventional post-training of speech foundation models
Interventional contrastive learning applied after pre-training disentangles speaker and content information in speech foundation model embeddings, yielding better out-of-domain speaker verification.
-
From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media
VLMs recover reliable population-level trends in climate change visual discourse on social media even when per-image accuracy is only moderate.
-
FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
A factorized log-linear model (FLiP) recovers 75-80% of lexical content from sentence embeddings and reveals English-centric biases in SONAR, LaBSE, and Gemini encoders.
-
HydraQE: OSU's Submission for the IWSLT 2026 Speech Translation Metrics Shared Task
HydraQE is a new end-to-end speech translation QE system using Qwen3-ASR backbone, sparsemax layer mixing, bidirectional Transformer, and multi-task curriculum training on human and pseudo labels that outperforms cascaded baselines.
-
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
The authors built and publicly released sentence-aligned simplification corpora for five languages by processing crowd-sourced data from comparable documents.
-
The Faiss library
Faiss is a library offering indexing methods and primitives for efficient vector similarity search, a core need in vector databases for AI applications.
- Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement