REVIEW 41 cited by
AudioGen: Textually Guided Audio Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We tackle the problem of generating audio samples conditioned on descriptive text captions. In this work, we propose AaudioGen, an auto-regressive generative model that generates audio samples conditioned on text inputs. AudioGen operates on a learnt discrete audio representation. The task of text-to-audio generation poses multiple challenges. Due to the way audio travels through a medium, differentiating ``objects'' can be a difficult task (e.g., separating multiple people simultaneously speaking). This is further complicated by real-world recording conditions (e.g., background noise, reverberation, etc.). Scarce text annotations impose another constraint, limiting the ability to scale models. Finally, modeling high-fidelity audio requires encoding audio at high sampling rate, leading to extremely long sequences. To alleviate the aforementioned challenges we propose an augmentation technique that mixes different audio samples, driving the model to internally learn to separate multiple sources. We curated 10 datasets containing different types of audio and text annotations to handle the scarcity of text-audio data points. For faster inference, we explore the use of multi-stream modeling, allowing the use of shorter sequences while maintaining a similar bitrate and perceptual quality. We apply classifier-free guidance to improve adherence to text. Comparing to the evaluated baselines, AudioGen outperforms over both objective and subjective metrics. Finally, we explore the ability of the proposed method to generate audio continuation conditionally and unconditionally. Samples: https://felixkreuk.github.io/audiogen
Forward citations
Cited by 41 Pith papers
-
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books
Dramarrator introduces object-based audio editing, where characters and scenes are editable objects that propagate edits across all linked speech, sound effects, music, and ambience.
-
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
A training-free latent swap method that replaces averaging with binary swapping in joint diffusion, improving long-form audio spectrum and panorama generation.
-
Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images
Stem uses a conditional diffusion model to infer spatially resolved gene expression from H&E histology images, outperforming regression baselines on several datasets.
-
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching
VoxAudio generates audio scenes with intelligible, temporally placed quoted speech by combining chunk-wise causal flow matching with multi-reward fine-tuning and a large transcript-annotated corpus.
-
HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion
HarmoniDPO pairs global and frame-level video features with preference-style optimization to generate audio from silent video, reporting improved synchronization and quality metrics over prior V2A baselines.
-
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection
MADBench introduces a component-aware audio-visual deepfake benchmark with independently manipulated speech and environmental audio, and shows environmental manipulation is easier to detect than synthetic speech.
-
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation
A structured soundscape benchmark with 25,707 binary semantic rubrics shows that rubric-based, audio-grounded evaluation tracks human semantic judgments better than CLAP-style global similarity for text-to-audio models.
-
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Using audio-aware LLMs to judge event presence and temporal order as DPO rewards improves multi-event text-to-audio instruction following.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
SemanticAudio: Audio Generation and Editing in Semantic Space
SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...
-
Testing chatbots on the creation of encoders for audio conditioned image generation
All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.
-
CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation
A multi-agent LLM pipeline autonomously constructs a music theory lexicon that, when used to enrich prompts, improves text-to-music generation across symbolic and audio models.
-
Improving GANs by leveraging the quantum noise from real hardware
Using bitstrings from a 16-qubit entangling circuit as a Gaussian latent prior lowers FID in WGAN, SNGAN, and BigGAN on CIFAR-10 versus a standard Gaussian.
-
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
AudioBERTScore, a training-free metric combining max-norm and p-norm similarity of audio embeddings, correlates more strongly with human subjective scores for text-to-audio synthesis than conventional metrics.
-
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.
-
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.
-
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
A ControlNet branch plus a frequency-aware feature aligner lets a pretrained masked generative TTA model produce video-synchronized foley, beating several from-scratch models on VGGSound.
-
PAST: Phonetic-Acoustic Speech Tokenizer
PAST jointly optimizes an EnCodec-style codec with CTC and phoneme classification losses to produce hybrid phonetic-acoustic speech tokens that outperform SpeechTokenizer and X-Codec.
-
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
AGAV-Rater, an LMM fine-tuned in two stages, achieves state-of-the-art quality scores for AI-generated audio-visual content, text-to-audio, and text-to-music.
-
ETTA: Elucidating the Design Space of Text-to-Audio Models
ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.
-
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.
-
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
A modular multi-modal rectified-flow model built on Stable Diffusion 3 achieves competitive text-to-image and text-to-audio generation while also supporting audio-to-image, image-to-text, and audio-to-text tasks.
-
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
A streaming encoder plus temporal and fully shared DiT-conditioned depth decoders converts semantic audio tokens to RVQ with constant memory and ~16× real-time on-device synthesis.
-
Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems
The abstract claims a first-of-kind, outcome-linked comparison of four vulnerability scoring systems showing major ranking disagreements, but the submitted full text is an unrelated paper, leaving the study unevaluable.
-
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
A single multimodal diffusion transformer generates video-synchronized general audio, speech, and song from flexible combinations of video, text, and lyrics inputs.
-
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
A three-stage diffusion pipeline maps 3D Gaussian Splatting object representations to position-dependent impact sounds, trained first on text captions and then on real recordings.
-
MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation
Fine-tuning MU-LLaMA on 3,371 pseudo-labeled video-music examples lets it produce scene captions that yield small subjective gains in video background music generation over music-only captions.
-
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.
-
Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models
Uncertainty-o estimates uncertainty in large multimodal models by perturbing prompts and computing entropy over semantically clustered answers, improving hallucination detection across five modalities.
-
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
A reconstruction-perception-reinforcement-attention framework for audio deepfake detection reports state-of-the-art EERs on ASVspoof 2019/2021 and competitive results on sound and singing benchmarks.
-
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
A caption-augmentation method that adds DSP-derived acoustic descriptors to text prompts gives a text-to-audio diffusion model controllable loudness, pitch, reverb, noise, brightness, fade, and duration.
-
Efficient Text-to-Audio Generation via Pruning
L1-norm filter pruning of AudioLDM's U-Net removes 83% of parameters and 39% of MACs, with quality maintained after 1M-step finetuning, but the comparison is confounded by unequal finetuning budgets.
-
Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.
-
Effectively obtaining acoustic, visual and textual data from videos
A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.
-
Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
A diffusion transformer generates the weights of a frozen-feature CLIP classifier head from text task descriptions, achieving moderate accuracy on unseen class subsets.
-
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.
-
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.
-
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.
-
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
A narrative review of 30 recent papers classifies AI sound-effect generators by input modality and summarizes reported progress and remaining limitations in temporal sync, evaluation, and controllability.
-
ASAudio: A Survey of Advanced Spatial Audio Research
A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.
-
Watermarking across Modalities for Content Tracing and Generative AI
A thesis showing that invisible watermarks can be embedded across images, audio, text, and model weights, with statistical tests for tracing AI-generated content.
Discussion (0). Continue with ORCID to comment.