Pith. sign in

REVIEW 1 cited by

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.06343 v1 pith:DIO3CPKC submitted 2025-06-01 cs.CL cs.AIcs.LGcs.SDeess.AS

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

classification cs.CL cs.AIcs.LGcs.SDeess.AS
keywords speechdataencodertesu-llmtextbuildingcomputationallanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advances in speech-enabled language models have shown promising results in building intelligent voice assistants. However, most existing approaches rely on large-scale paired speech-text data and extensive computational resources, which pose challenges in terms of scalability and accessibility. In this paper, we present \textbf{TESU-LLM}, a novel framework that enables training speech-capable language models using only text data. Our key insight is to leverage a unified encoder that maps semantically equivalent text and speech inputs to a shared latent space. By aligning the encoder output with the embedding space of a LLM via a lightweight projection network, we enable the model to generalize from text-only supervision to speech-based inference. Despite being trained exclusively on text, TESU-LLM achieves strong performance on various speech-related benchmarks, comparable to baseline methods trained with large-scale multimodal datasets and substantial computational resources. These results highlight the effectiveness and efficiency of our approach, offering a scalable path toward building speech LLMs without speech data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

    cs.CL 2026-07 conditional novelty 5.0

    MEUSLI, a released family of linear projectors, enables open-source LLM-based ASR across 28 European languages and supports few-hour adaptation to unseen languages and extra speech tasks.