Pith. sign in

REVIEW 56 cited by

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12503 v3 pith:KSUBW4F6 submitted 2023-01-29 cs.SD cs.AIcs.MMeess.ASeess.SP

classification cs.SDcs.AIcs.MMeess.ASeess.SP
keywords audioldmaudiogenerationlatentsystemclapcomputationalembedding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-audio (TTA) system has recently gained attention for its ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a latent space to learn the continuous audio representations from contrastive language-audio pretraining (CLAP) latents. The pretrained CLAP models enable us to train LDMs with audio embedding while providing text embedding as a condition during sampling. By learning the latent representations of audio signals and their compositions without modeling the cross-modal relationship, AudioLDM is advantageous in both generation quality and computational efficiency. Trained on AudioCaps with a single GPU, AudioLDM achieves state-of-the-art TTA performance measured by both objective and subjective metrics (e.g., frechet distance). Moreover, AudioLDM is the first TTA system that enables various text-guided audio manipulations (e.g., style transfer) in a zero-shot fashion. Our implementation and demos are available at https://audioldm.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 56 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Analytic Distribution of Classifier-Free Guidance for Schedule Design

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Deterministic CFG samples from p_t0 times an exponential path integral of the score discrepancy, and the resulting schedule DG-CFG reduces sampling steps at high guidance.

  2. MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

    cs.CV 2026-06 conditional novelty 7.0 of 10

    MAVIN is a dual-tower diffusion framework that produces temporally aligned multi-shot audio-visual content from hierarchical captions with optional multi-identity image and audio references.

  3. Net-Ev$^2$: A Generative Simulator for Network Event Evolution

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Net-Ev² proposes a two-stage generative simulator with structure-guided masked pre-training and topology-aware diffusion using graph U-Net down/upsampling to model network event evolution from text inputs, plus a new ...

  4. Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    A per-timestep conditioned diffusion transformer generates realistic fMRI dynamics for unseen cognitive tasks by injecting compositional language and optional spatial priors in-context.

  5. Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    An entropy-based log-barrier on DiT outputs acts as an online curriculum in supervised diffusion fine-tuning, producing higher thematic development and textural diversity than standard training on MusicCaps.

  6. HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation

    cs.HC 2026-05 unverdicted novelty 7.0 of 10

    HapticLDM is the first latent diffusion model that generates vibrotactile signals directly from text, using dynamic text curation and global denoising to improve realism and semantic alignment over autoregressive baselines.

  7. Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    MixtureTT performs direct per-stem timbre transfer on polyphonic mixtures via a shared diffusion transformer, outperforming single-stem baselines on SATB choral data while eliminating cascaded separation errors.

  8. Latent Fourier Transform

    cs.SD 2026-04 unverdicted novelty 7.0 of 10

    LatentFT uses latent-space Fourier transforms and frequency masking in diffusion autoencoders to enable timescale-specific manipulation of musical structure in generative models.

  9. FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    FoleyDesigner generates spatio-temporally aligned stereo Foley audio for film clips via multi-agent analysis, diffusion models on video cues, and LLM mixing, supported by the new FilmStereo dataset.

  10. JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

    cs.GR 2026-01 unverdicted novelty 7.0 of 10

    JUST-DUB-IT adapts a joint audio-visual diffusion model via LoRA to generate high-quality dubbed videos with translated audio and lip-synced facial motion.

  11. AudioMoG: Guiding Audio Generation with Mixture-of-Guidance

    cs.SD 2025-09 unverdicted novelty 7.0 of 10

    AudioMoG is a mixture-of-guidance sampling technique that combines CFG and AG signals to outperform single-guidance baselines in text-to-audio generation at equivalent speed.

  12. A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions

    stat.ML 2025-08 conditional novelty 7.0 of 10

    A new analysis shows O~(d/epsilon) steps suffice for KL-close diffusion sampling under only L2 score error and finite second moment assumptions, improving the known O~(d/epsilon^2).

  13. Flow Matching Policy Gradients

    cs.LG 2025-07 conditional novelty 7.0 of 10

    FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.

  14. Amortized Moment Matching for Visual Generation

    cs.LG 2026-07 accept novelty 6.0 of 10

    Amortized Fréchet Distance uses neural nets to match conditional means and covariances, yielding stronger one-step visual generators than explicit FD-loss or multi-step teachers.

  15. RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

    cs.SD 2026-07 conditional novelty 6.0 of 10

    RPPNet generates melodies by planning variable-length perceptually grouped rhythm-pitch primitives first and then decoding them into notes, beating bar-level baselines in subjective structure and musicality ratings.

  16. FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

    cs.DC 2026-07 conditional novelty 6.0 of 10

    FlashDiff reduces diffusion serving latency by 30–97% and raises throughput 1.2–2.2× by adaptively skipping refinement of latent regions that no longer need it.

  17. Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Beat-guided contrastive alignment of pretrained MotionBERT/MERT features plus ControlNet conditioning of AudioLDM improves dance–music alignment on AIST++ over a MusicGen textual-inversion baseline while remaining com...

  18. SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.

  19. An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

    eess.AS 2026-07 unverdicted novelty 6.0 of 10

    Extends vLLM with delay-pattern de-interleaving, multi-stream sampling, and co-scheduled CFG to achieve 80% of non-CFG throughput for unified audio tasks while open-sourcing the pipeline.

  20. dMoE: dLLMs with Learnable Block Experts

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    dMoE aggregates token expert distributions to block level in dLLMs, cutting unique experts from 69.5 to 14.6, memory by 76-80%, and latency by 1.14-1.66x while retaining 99.11% performance.

  21. WavFlow: Audio Generation in Waveform Space

    cs.SD 2026-05 conditional novelty 6.0 of 10

    WavFlow performs direct waveform audio generation via flow matching on 2D token grids from raw patches plus amplitude lifting, matching latent-based methods on VGGSound and AudioCaps without intermediate compression.

  22. Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation

    cs.SD 2026-05 unverdicted novelty 6.0 of 10

    Caption poisoning attacks can steer retrieval-augmented text-to-music generation toward attacker-chosen targets by injecting crafted captions into the knowledge database.

  23. PoDAR: Power-Disentangled Audio Representation for Generative Modeling

    eess.AS 2026-05 unverdicted novelty 6.0 of 10

    PoDAR disentangles audio signal power from semantic content in latents using power augmentation and consistency objectives, yielding 2x faster convergence and gains of 0.055 speaker similarity and 0.22 UTMOS when appl...

  24. DiffATS: Diffusion in Aligned Tensor Space

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    DiffATS trains diffusion models directly on aligned Tucker tensor primitives that are proven to be homeomorphisms, delivering efficient unconditional and conditional generation across images, videos, and PDE data with...

  25. Stage-adaptive audio diffusion modeling

    cs.SD 2026-05 unverdicted novelty 6.0 of 10

    A semantic progress signal from SSL discrepancy slope enables three stage-aware mechanisms that improve training efficiency and performance in audio diffusion models over static baselines.

  26. Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Jointly training the watermark embedder/detector with the source separator enables ~1% bit-error-rate recovery of per-stem watermarks after mixing and separation, where independent training yields 15–35%.

  27. Diffusion Models Memorize in Training -- and Generalize in Inference

    cs.LG 2026-03 unverdicted novelty 6.0 of 10

    Diffusion models overfit denoising loss at intermediate noise but generalize in inference as model error smooths the flow field and sampling paths avoid memorized noisy training data.

  28. Dual-End Consistency Model

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    DE-CM reaches state-of-the-art one-step FID of 1.70 on ImageNet 256x256 by decomposing PF-ODE trajectories into three critical sub-trajectories and using flow matching plus N2N mapping for stability.

  29. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  30. Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Separating the text condition into a video caption and a visually grounded audio caption, and fusing the diffusion towers with dual cross-attention, gives the reported-best text-to-sounding-video quality and synchroni...

  31. Audio-Guided Visual Editing with Complex Multi-Modal Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...

  32. DGSNA: Dynamic Generative Scene-based Noise Addition method

    cs.SD 2024-11 unverdicted novelty 6.0 of 10

    DGSNA dynamically generates scene-specific noise via prompt-driven language models and text-to-audio diffusion, then mixes it with speech to improve recognition and keyword spotting robustness by up to 11.32%.

  33. Exploring Efficient Waveform Diffusion Models for Foley Sound Generation

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Dual-path attention over time-frequency representations lets a 3.26M-parameter waveform diffusion model match the quality of 50M+ parameter Foley generators.

  34. FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

    cs.DC 2026-07 conditional novelty 5.0 of 10

    FlashDiff cuts diffusion serving latency 30–97% and raises throughput 1.2–2.2× by selectively executing only active latent regions and rescheduling the reclaimed compute.

  35. ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    ALM2Vec learns unified audio embeddings from large audio-language models for text-audio retrieval, instruction-aware retrieval, and other tasks across domains.

  36. ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    ARIA adaptively focuses distillation updates on conditioning-space regions with high ongoing teacher-student misalignment while preserving the base objective.

  37. EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

    cs.SD 2026-05 unverdicted novelty 5.0 of 10

    EigeNet applies a cross-view alternate-attention transformer with geometry modulation for few-shot novel-view RIR prediction, reporting SOTA results on simulated and real data.

  38. Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation

    cs.SD 2026-05 unverdicted novelty 5.0 of 10

    A one-step text-to-audio model using energy-distance training and contextual distillation outperforms prior fast baselines on AudioCaps and achieves up to 8.5x faster inference than the multi-step IMPACT system with c...

  39. Woosh: A Sound Effects Foundation Model

    cs.SD 2026-04 accept novelty 5.0 of 10

    Woosh is a new publicly released foundation model optimized for high-quality sound effect generation from text or video, showing competitive or better results than open alternatives like Stable Audio Open.

  40. Dual-End Consistency Model

    cs.CV 2026-02 conditional novelty 5.0 of 10

    DE-CM trains a flow-map consistency model on three sub-trajectories (coupling, instantaneous, noise-to-noisy) and reports 1.70 FID one-step on ImageNet 256.

  41. Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

    cs.CL 2025-10 unverdicted novelty 5.0 of 10

    An adaptive CFG method that tunes guidance based on LLM-detected mismatch between emotion prompts and text semantics improves emotional expressiveness in AR TTS while preserving audio quality and intelligibility.

  42. Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

    cs.CL 2025-10 reject novelty 5.0 of 10

    An emotion TTS system adjusts Classifier-Free Guidance strength according to text-style semantic mismatch; it shows small emotion-accuracy gains, but headline baselines and subjective results are absent from the main text.

  43. Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

    cs.CL 2025-10 unverdicted novelty 5.0 of 10

    Introduces CCG-CFG with inconsistency-based dynamic scales and hard-sample mining distillation to boost emotional alignment in auto-regressive TTS, reporting up to 12% absolute gains in emotion recognition accuracy.

  44. Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    A bar-level symbolic-score song generator (BACH) is claimed to beat published systems and commercial Suno on human-rated quality, duration, and efficiency, but the supporting full text is corrupted and unverifiable.

  45. AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A single multimodal diffusion transformer generates video-synchronized general audio, speech, and song from flexible combinations of video, text, and lyrics inputs.

  46. Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A multi-view SSL framework with combined reconstruction and separation-based contrastive losses obtains disentangled pitch and instrument subspaces without the accuracy loss seen with contrastive-only training on NSynth.

  47. Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An integrated storytelling framework that generates text, scene graphs, images, and sound together reports higher expert-rated coherence than a sequential baseline, but the evaluation is small and qualitative.

  48. How Far Are We from Generating Missing Modalities with Foundation Models?

    cs.MM 2025-06 unverdicted novelty 5.0 of 10

    Evaluates 42 variants of foundation models across three formalized paradigms for missing modality reconstruction, identifies shortfalls in semantic extraction and validation, and introduces an agentic framework that r...

  49. MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    MAVIN proposes boundary-aware attention, ID-aware propagation, a multi-agent scripting pipeline, and the MAVINSet dataset as the first framework for multi-shot audio-visual generation with narrative control, claiming ...

  50. STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

    eess.AS 2026-06 unverdicted novelty 4.0 of 10

    STAR-VAE introduces topology-aware regularization to reshape VAE latent geometry for audio, claiming to resolve the Rate-Distortion-Regularity Trilemma and achieve SOTA reconstruction.

  51. ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis

    cs.SD 2026-04 unverdicted novelty 4.0 of 10

    ATRIE disentangles timbre and prosody in a Persona-Prosody Dual-Track model distilled from a large LLM to achieve strong identity preservation (EER 0.04) and emotional speech synthesis with SOTA results on an extended...

  52. Inference-time Scaling for Diffusion-based Audio Super-resolution

    cs.SD 2025-08 conditional novelty 4.0 of 10

    Generating 120 candidate super-resolved audios and choosing the best by task-specific verifiers improves speech, music, and sound effects over single-sample diffusion output, at 120x compute.

  53. AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

    cs.SD 2026-08 conditional novelty 3.0 of 10

    A narrative review of 30 recent papers classifies AI sound-effect generators by input modality and summarizes reported progress and remaining limitations in temporal sync, evaluation, and controllability.

  54. AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

    cs.SD 2026-04 unverdicted novelty 3.0 of 10

    AT-ADD introduces standardized tracks and datasets for evaluating audio deepfake detectors on speech under real-world conditions and on diverse unknown audio types to promote generalization beyond speech-centric methods.

  55. WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

    cs.SD 2025-08 conditional novelty 3.0 of 10

    WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.

  56. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

Pith tools