Pith. sign in

REVIEW 1 cited by

How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.17666 v2 pith:3E4IPOBQ submitted 2024-11-26 cs.CL

classification cs.CL
keywords speechtextmodelsrepresentationscross-modalcross-lingualdifferencesfoundation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal foundation models aim to create a unified representation space that abstracts away from surface features like language syntax or modality differences. To investigate this, we study the internal representations of three recent models, analyzing the model activations from semantically equivalent sentences across languages in the text and speech modalities. Our findings reveal that: 1) Cross-modal representations converge over model layers, except in the initial layers specialized at text and speech processing. 2) Length adaptation is crucial for reducing the cross-modal gap between text and speech, although current approaches' effectiveness is primarily limited to high-resource languages. 3) Speech exhibits larger cross-lingual differences than text. 4) For models not explicitly trained for modality-agnostic representations, the modality gap is more prominent than the language gap.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Text models encode linguistic taxonomies early and densely; speech models develop them later and less prominently, with multimodal models showing intermediate patterns.

Pith tools