Pith. sign in

REVIEW 13 cited by

VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16107 v1 pith:TV7BUSLE submitted 2023-05-25 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modeltaskslanguagecodecspeechviolaconditionalcross-modal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we propose VioLA, a single auto-regressive Transformer decoder-only network that unifies various cross-modal tasks involving speech and text, such as speech-to-text, text-to-text, text-to-speech, and speech-to-speech tasks, as a conditional codec language model task via multi-task learning framework. To accomplish this, we first convert all the speech utterances to discrete tokens (similar to the textual data) using an offline neural codec encoder. In such a way, all these tasks are converted to token-based sequence conversion problems, which can be naturally handled with one conditional language model. We further integrate task IDs (TID) and language IDs (LID) into the proposed model to enhance the modeling capability of handling different languages and tasks. Experimental results demonstrate that the proposed VioLA model can support both single-modal and cross-modal tasks well, and the decoder-only model achieves a comparable and even better performance than the strong baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

    cs.CL 2026-07 conditional novelty 6.0 of 10

    PINT uses parallel utterances to train invariant speech tokens, cutting speaker probe accuracy from 93.1% to 1.2% and lowering LM perplexity by 27 to 30% relative to HuBERT and WavLM tokens.

  2. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  3. The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Text jailbreak prompts converted to audio match or beat dedicated audio jailbreaks on omni-models, and transfer success tracks how tightly the model aligns text and audio representations.

  4. NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.

  5. Probing the Robustness Properties of Neural Speech Codecs

    eess.AS 2025-05 conditional novelty 6.0 of 10

    DAC is the most noise-robust neural codec at high bitrates, but at 3 kbps EnCodec wins, and measured non-linearity correlates with robustness.

  6. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  7. Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

    cs.SD 2025-05 conditional novelty 6.0 of 10

    AJailBench is an open benchmark showing that large audio-language models can be jailbroken through TTS-converted text attacks and through subtle acoustic perturbations that preserve speech semantics.

  8. A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Across six speech-understanding tasks, continuous SSL features beat k-means discrete tokens in almost all cases when paired with a 0.5B instruction-tuned LLM.

  9. Towards Flow-Matching-based TTS without Classifier-Free Guidance

    eess.AS 2025-04 reject novelty 5.0 of 10

    Modifying the flow-matching training target lets F5-TTS synthesize speech without classifier-free guidance at inference, halving per-step cost and improving measured WER, SIM-O, and MOS.

  10. Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.

  11. Speech Separation using Neural Audio Codecs with Embedding Loss

    eess.AS 2024-11 conditional novelty 5.0 of 10

    Codecformer-EL replaces the waveform loss in codec-based speech separation with an MSE loss on frozen codec embeddings, cutting training cost by about half and improving perceptual metrics on WSJ0-2mix.

  12. Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy

    eess.AS 2025-09 conditional novelty 4.0 of 10

    Learning layer weights on continuous features and reusing them for discrete token extraction, plus fine-tuning XLS-R with extra data, yields 44% relative CER reduction on ML-SUPERB and tops the challenge's single-syst...

  13. Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Under matched settings, continuous SSL speech features generally outperform discrete tokens on six spoken language understanding tasks in SpeechLLMs.

Pith tools