Pith. sign in

REVIEW 41 cited by

AudioGen: Textually Guided Audio Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.15352 v2 pith:CGLOUNM4 submitted 2022-09-30 cs.SD cs.CLcs.LGeess.AS

classification cs.SDcs.CLcs.LGeess.AS
keywords audiotextaudiogensamplesmultipleabilityannotationschallenges
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We tackle the problem of generating audio samples conditioned on descriptive text captions. In this work, we propose AaudioGen, an auto-regressive generative model that generates audio samples conditioned on text inputs. AudioGen operates on a learnt discrete audio representation. The task of text-to-audio generation poses multiple challenges. Due to the way audio travels through a medium, differentiating ``objects'' can be a difficult task (e.g., separating multiple people simultaneously speaking). This is further complicated by real-world recording conditions (e.g., background noise, reverberation, etc.). Scarce text annotations impose another constraint, limiting the ability to scale models. Finally, modeling high-fidelity audio requires encoding audio at high sampling rate, leading to extremely long sequences. To alleviate the aforementioned challenges we propose an augmentation technique that mixes different audio samples, driving the model to internally learn to separate multiple sources. We curated 10 datasets containing different types of audio and text annotations to handle the scarcity of text-audio data points. For faster inference, we explore the use of multi-stream modeling, allowing the use of shorter sequences while maintaining a similar bitrate and perceptual quality. We apply classifier-free guidance to improve adherence to text. Comparing to the evaluated baselines, AudioGen outperforms over both objective and subjective metrics. Finally, we explore the ability of the proposed method to generate audio continuation conditionally and unconditionally. Samples: https://felixkreuk.github.io/audiogen

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 41 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 54 citations worldwide. Full citation record

  1. Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books

    cs.HC 2026-08 conditional novelty 7.0 of 10

    Dramarrator introduces object-based audio editing, where characters and scenes are editable objects that propagate edits across all linked speech, sound effects, music, and ambience.

  2. Latent Swap Joint Diffusion for 2D Long-Form Latent Generation

    cs.SD 2025-02 conditional novelty 7.0 of 10

    A training-free latent swap method that replaces averaging with binary swapping in joint diffusion, improving long-form audio spectrum and panorama generation.

  3. Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Stem uses a conditional diffusion model to infer spatially resolved gene expression from H&E histology images, outperforming regression baselines on several datasets.

  4. VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

    cs.SD 2026-08 conditional novelty 6.0 of 10

    VoxAudio generates audio scenes with intelligible, temporally placed quoted speech by combining chunk-wise causal flow matching with multi-reward fine-tuning and a large transcript-annotated corpus.

  5. HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

    cs.CV 2026-08 reject novelty 6.0 of 10

    HarmoniDPO pairs global and frame-level video features with preference-style optimization to generate audio from silent video, reporting improved synchronization and quality metrics over prior V2A baselines.

  6. MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

    cs.SD 2026-08 conditional novelty 6.0 of 10

    MADBench introduces a component-aware audio-visual deepfake benchmark with independently manipulated speech and environmental audio, and shows environmental manipulation is easier to detect than synthetic speech.

  7. AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A structured soundscape benchmark with 25,707 binary semantic rubrics shows that rubric-based, audio-grounded evaluation tracks human semantic judgments better than CLAP-style global similarity for text-to-audio models.

  8. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Using audio-aware LLMs to judge event presence and temporal order as DPO rewards improves multi-event text-to-audio instruction following.

  9. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  10. SemanticAudio: Audio Generation and Editing in Semantic Space

    eess.AS 2026-01 conditional novelty 6.0 of 10

    SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...

  11. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  12. CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation

    cs.SD 2025-08 conditional novelty 6.0 of 10

    A multi-agent LLM pipeline autonomously constructs a music theory lexicon that, when used to enrich prompts, improves text-to-music generation across symbolic and audio models.

  13. Improving GANs by leveraging the quantum noise from real hardware

    quant-ph 2025-07 conditional novelty 6.0 of 10

    Using bitstrings from a 16-qubit entangling circuit as a Gaussian latent prior lowers FID in WGAN, SNGAN, and BigGAN on CIFAR-10 versus a standard Gaussian.

  14. AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

    cs.SD 2025-07 conditional novelty 6.0 of 10

    AudioBERTScore, a training-free metric combining max-norm and p-norm similarity of audio embeddings, correlates more strongly with human subjective scores for text-to-audio synthesis than conventional metrics.

  15. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  16. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  17. SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A ControlNet branch plus a frequency-aware feature aligner lets a pretrained masked generative TTA model produce video-synchronized foley, beating several from-scratch models on VGGSound.

  18. PAST: Phonetic-Acoustic Speech Tokenizer

    cs.SD 2025-05 conditional novelty 6.0 of 10

    PAST jointly optimizes an EnCodec-style codec with CTC and phoneme classification losses to produce hybrid phonetic-acoustic speech tokens that outperform SpeechTokenizer and X-Codec.

  19. AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

    cs.MM 2025-01 conditional novelty 6.0 of 10

    AGAV-Rater, an LMM fine-tuned in two stages, achieves state-of-the-art quality scores for AI-generated audio-visual content, text-to-audio, and text-to-music.

  20. ETTA: Elucidating the Design Space of Text-to-Audio Models

    cs.SD 2024-12 conditional novelty 6.0 of 10

    ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.

  21. VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.

  22. OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

    cs.MM 2024-12 conditional novelty 6.0 of 10

    A modular multi-modal rectified-flow model built on Stable Diffusion 3 achieves competitive text-to-image and text-to-audio generation while also supporting audio-to-image, image-to-text, and audio-to-text tasks.

  23. Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

    cs.SD 2026-07 conditional novelty 5.5 of 10

    A streaming encoder plus temporal and fully shared DiT-conditioned depth decoders converts semantic audio tokens to RVQ with constant memory and ~16× real-time on-device synthesis.

  24. Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems

    cs.CR 2025-08 unverdicted novelty 5.0 of 10

    The abstract claims a first-of-kind, outcome-linked comparison of four vulnerability scoring systems showing major ranking disagreements, but the submitted full text is an unrelated paper, leaving the study unevaluable.

  25. AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A single multimodal diffusion transformer generates video-synchronized general audio, speech, and song from flexible combinations of video, text, and lyrics inputs.

  26. SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A three-stage diffusion pipeline maps 3D Gaussian Splatting object representations to position-dependent impact sounds, trained first on text captions and then on real recordings.

  27. MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

    cs.AI 2025-07 reject novelty 5.0 of 10

    Fine-tuning MU-LLaMA on 3,371 pseudo-labeled video-music examples lets it produce scene captions that yield small subjective gains in video background music generation over music-only captions.

  28. UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

    cs.SD 2025-06 conditional novelty 5.0 of 10

    UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.

  29. Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Uncertainty-o estimates uncertainty in large multimodal models by perturbing prompts and computing entropy over semantically clustered answers, improving hallucination detection across five modalities.

  30. RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A reconstruction-perception-reinforcement-attention framework for audio deepfake detection reports state-of-the-art EERs on ASVspoof 2019/2021 and competitive results on sound and singing benchmarks.

  31. SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation

    cs.SD 2024-12 conditional novelty 5.0 of 10

    A caption-augmentation method that adds DSP-derived acoustic descriptors to text prompts gives a text-to-audio diffusion model controllable loudness, pitch, reverb, noise, brightness, fade, and duration.

  32. Efficient Text-to-Audio Generation via Pruning

    eess.AS 2026-07 conditional novelty 4.0 of 10

    L1-norm filter pruning of AudioLDM's U-Net removes 83% of parameters and 39% of MACs, with quality maintained after 1M-step finetuning, but the comparison is confounded by unequal finetuning budgets.

  33. Segment Transformer: AI-Generated Music Detection via Music Structural Analysis

    cs.SD 2025-09 conditional novelty 4.0 of 10

    A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.

  34. Effectively obtaining acoustic, visual and textual data from videos

    cs.MM 2025-09 conditional novelty 4.0 of 10

    A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.

  35. Text2Weight: Bridging Natural Language and Neural Network Weight Spaces

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A diffusion transformer generates the weights of a frozen-feature CLIP classifier head from text task descriptions, achieving moderate accuracy on unseen class subsets.

  36. DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

    cs.SD 2025-05 conditional novelty 4.0 of 10

    DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.

  37. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

  38. YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

    cs.SD 2024-12 reject novelty 4.0 of 10

    A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.

  39. AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

    cs.SD 2026-08 conditional novelty 3.0 of 10

    A narrative review of 30 recent papers classifies AI sound-effect generators by input modality and summarizes reported progress and remaining limitations in temporal sync, evaluation, and controllability.

  40. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

  41. Watermarking across Modalities for Content Tracing and Generative AI

    cs.CR 2025-02 conditional novelty 3.0 of 10

    A thesis showing that invisible watermarks can be embedded across images, audio, text, and model weights, with statistical tests for tracing AI-generated content.

Pith tools