PRIME-Speech adds low-latency speech output to frozen S2T LLMs by synchronizing a causal post-decoder with intermediate hidden states and using mixed conditioning plus turn-level KV-cache packing, preserving original S2T performance across translation, QA, and dialogue tasks.
Neural codec language models are zero-shot text to speech synthesizers
3 Pith papers cite this work. Polarity classification is still indexing.
fields
eess.AS 3years
2026 3representative citing papers
A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18-HQ tests.
Cross-feature knowledge distillation from Fbank teachers lets EnCodec-token ASV systems approach continuous-feature accuracy by better exploiting preserved speaker cues.
citing papers explorer
-
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation
PRIME-Speech adds low-latency speech output to frozen S2T LLMs by synchronizing a causal post-decoder with intermediate hidden states and using mixed conditioning plus turn-level KV-cache packing, preserving original S2T performance across translation, QA, and dialogue tasks.
-
Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models
A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18-HQ tests.
-
Text-Independent Speaker Verification Using Discrete Audio Tokens
Cross-feature knowledge distillation from Fbank teachers lets EnCodec-token ASV systems approach continuous-feature accuracy by better exploiting preserved speaker cues.