Pith. sign in

REVIEW 36 cited by

SoundStorm: Efficient Parallel Audio Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.09636 v1 pith:RLM74DQB submitted 2023-05-16 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audiogenerationsoundstormmodelaudiolmefficientparallelseconds
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present SoundStorm, a model for efficient, non-autoregressive audio generation. SoundStorm receives as input the semantic tokens of AudioLM, and relies on bidirectional attention and confidence-based parallel decoding to generate the tokens of a neural audio codec. Compared to the autoregressive generation approach of AudioLM, our model produces audio of the same quality and with higher consistency in voice and acoustic conditions, while being two orders of magnitude faster. SoundStorm generates 30 seconds of audio in 0.5 seconds on a TPU-v4. We demonstrate the ability of our model to scale audio generation to longer sequences by synthesizing high-quality, natural dialogue segments, given a transcript annotated with speaker turns and a short prompt with the speakers' voices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 36 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

    eess.AS 2026-08 conditional novelty 7.0 of 10

    SemBridge supervises autoregressive states with discrete semantic tokens during training, improving content fidelity of continuous-latent speech generation without changing inference.

  2. ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A speech language model trained only on audiobooks with explicit word-level prosody tokens displays emerging abilities in prosody-controlled generation, emphasis and emotion understanding, and prosodic consistency acr...

  3. SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling

    eess.AS 2025-01 conditional novelty 7.0 of 10

    A single masked language model can both entropy-code audio tokens for efficient transmission and predict missing tokens for packet loss concealment, outperforming traditional codecs in tests.

  4. Efficient Generative Modeling with Residual Vector Quantization-Based Tokens

    cs.LG 2024-12 conditional novelty 7.0 of 10

    ResGen predicts cumulative vector embeddings of masked RVQ tokens, decoupling generative sampling cost from token depth and improving FID and TTS metrics over autoregressive baselines.

  5. Scaling Transformers for Low-Bitrate High-Quality Speech Coding

    eess.AS 2024-11 conditional novelty 7.0 of 10

    A scaled transformer codec with FSQ reaches state-of-the-art speech reconstruction at 400-700 bps, outperforming CNN/RVQ baselines.

  6. DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

    eess.AS 2026-07 conditional novelty 6.5 of 10

    A LoRA-plus-convolution adaptation converts an AR TTS backbone into a confidence-ordered discrete diffusion model that improves WER and speed on limited data.

  7. Luna-TTS Family Technical Report

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Luna-TTS shows that a 0.6B masked-diffusion text-to-speech model pretrained on 1M hours can match or beat autoregressive systems on standard benchmarks while decoding in parallel or in 1.28-second streaming blocks.

  8. A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

    eess.AS 2026-08 conditional novelty 6.0 of 10

    The deciding factors for modeling an audio latent are dependency horizon and conditional ambiguity, viewed jointly with the representation, not whether the latent is discrete or continuous.

  9. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A mobile-oriented 83M-parameter masked transformer with sparse phone-anchored temporal embeddings achieves RTF 0.08 and lower WER than MaskGCT/F5-TTS on Seed-TTS test sets.

  10. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Using audio-aware LLMs to judge event presence and temporal order as DPO rewards improves multi-event text-to-audio instruction following.

  11. Fr\'echet Distance Loss on Speech Representations for Text-to-Speech Synthesis

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Four-step SR-FD fine-tuning cuts VoxCPM2's Seed-TTS English WER from 2.23% to 1.41%, beating the ten-step baseline at 1.74% with no inference-time cost.

  12. DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

    cs.CL 2026-06 conditional novelty 6.0 of 10

    Block discrete diffusion over X-Codec2 tokens yields competitive zero-shot TTS with 0.6B parameters, 20K training hours, and a 0.15 real-time factor.

  13. LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    LaDA-Band generates complete vocal-aligned instrumental accompaniments with discrete masked diffusion, claiming simultaneous gains in audio fidelity, long-range coherence, and orchestration quality.

  14. SpectroStream: A Versatile Neural Codec for General Audio

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SpectroStream, a 2D time-frequency neural codec, reconstructs 48 kHz stereo music at 4-16 kbps with better ViSQOL and subjective quality than DAC.

  15. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  16. DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

    eess.AS 2025-06 conditional novelty 6.0 of 10

    DiffSoundStream uses a latent diffusion decoder conditioned on WavLM semantic tokens and coarse SoundStream acoustic tokens to match 100-token-per-second quality at 50 tokens per second.

  17. LinearVC: Linear transformations of self-supervised features through the lens of voice conversion

    eess.AS 2025-06 conditional novelty 6.0 of 10

    A linear projection of WavLM features, including a rank-100 SVD factorization, performs voice conversion on par with much larger neural systems.

  18. Learning to Upsample and Upmix Audio in the Latent Domain

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Lightweight networks trained only on autoencoder latent codes can do bandwidth extension and mono-to-stereo upmixing at a fraction of the FLOPS of raw-audio models, but match those models only when the baselines are a...

  19. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.

  20. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  21. A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A hierarchical speech enhancement pipeline that estimates semantic tokens first and acoustic tokens second, via a factorized codec and diffusion, improves DNSMOS and downstream TTS speaker similarity in noisy far-fiel...

  22. Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Preference alignment on the new INTP dataset improves intelligibility and quality of zero-shot TTS across diverse domains, with weak-to-strong generalization shown on CosyVoice 2 and Ints.

  23. SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A four-stage pipeline (separate, correct in text, re-synthesize, align) improves speech separation quality and out-of-domain noise robustness.

  24. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

    eess.AS 2025-02 conditional novelty 6.0 of 10

    GenSE enhances speech by first denoising semantic tokens with a language model and then generating acoustic tokens from a single-quantizer codec, reporting higher DNSMOS, speaker similarity, and lower WER than prior systems.

  25. Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Continuous autoregressive text-to-speech with a Gaussian-mixture codec matches or beats a discrete-codec VALL-E baseline with a fraction of the language model parameters.

  26. CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A new large-scale dataset and codec taxonomy show that codec re-synthesized speech, especially balanced by decoder type, trains detectors that catch codec-based deepfake speech better than traditional anti-spoofing training.

  27. SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SOLAMI is an end-to-end social vision-language-action model that takes a user's speech and body motion as input and generates a 3D character's spoken and gestural responses in one pass.

  28. Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A post-hoc Re-Bottleneck network trained only in latent space can impose ordering, semantic alignment, or equivariance on pre-trained audio autoencoder latents with little extra compute.

  29. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching

    eess.AS 2025-06 conditional novelty 5.0 of 10

    R-VC performs zero-shot voice conversion in two sampling steps while transferring the target speaker's rhythm, matching or exceeding prior systems in naturalness and intelligibility.

  30. FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A three-stage fusion of a sparse compression separator, an SSL-conditioned codec-token generator, and a fusion network achieves third place in URGENT 2025 and improves perceptual metrics with a modest fidelity trade-off.

  31. Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning the English F5-TTS model on small Indian-language datasets yields a near-human-quality polyglot TTS (IN-F5) with voice cloning, code-mixing, and zero-resource synthesis for Bhojpuri and Tulu.

  32. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.

  33. ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

    eess.AS 2025-01 conditional novelty 5.0 of 10

    ZSVC uses a speech codec and a latent diffusion model with a style prompt, plus an information bottleneck and adversarial training, to convert speaking style while preserving speaker identity in zero-shot settings.

  34. Overview of the Amphion Toolkit (v0.2)

    cs.SD 2025-01 conditional novelty 4.0 of 10

    Amphion v0.2 is an open-source toolkit for audio, music, and speech generation, adding a 101K-hour multilingual dataset, processing pipelines, and pretrained models.

  35. CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation

    cs.SD 2025-01 reject novelty 4.0 of 10

    CycleFlow applies cycle-consistency training and a two-stage flow matching decoder to non-parallel voice conversion, but the cycle loss as written is not a cycle and the similarity metric is the same encoder used for ...

  36. Watermarking across Modalities for Content Tracing and Generative AI

    cs.CR 2025-02 conditional novelty 3.0 of 10

    A thesis showing that invisible watermarks can be embedded across images, audio, text, and model weights, with statistical tests for tracing AI-generated content.

Pith tools