REVIEW 21 cited by
Long-form music generation with latent diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure.
Forward citations
Cited by 21 Pith papers
-
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
A training-free latent swap method that replaces averaging with binary swapping in joint diffusion, improving long-form audio spectrum and panorama generation.
-
MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
A text-to-music system that first writes an explicit structural layout and then generates audio from it shows better long-range boundary agreement than a matched model without the layout.
-
On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs
A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.
-
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
In-the-wild Audio Spatialization with Flexible Text-guided Localization
A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.
-
Fast Text-to-Audio Generation with Adversarial Post-Training
ARC post-training speeds up text-to-audio generation to near-real-time speeds on GPUs and a few seconds on phones, without distillation or classifier-free guidance.
-
Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
MUFFIN is a neural audio codec that quantizes separate latent frequency bands with dedicated codebooks, claiming SOTA reconstruction quality and a competitive 12.5 Hz ultra-low-rate variant.
-
Bridge-SR: Schr\"odinger Bridge for Efficient SR
Bridge-SR applies tractable Schrödinger bridge models to waveform-domain speech super-resolution, and with 1.7M parameters reports the lowest log-spectral distance on VCTK while matching diffusion quality at 4 sampling steps.
-
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.
-
ETTA: Elucidating the Design Space of Text-to-Audio Models
ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.
-
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
FolAI predicts an editable RMS envelope from silent video and uses it, with semantic embeddings, to condition a Stable Audio diffusion model for 44.1 kHz stereo foley generation.
-
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
A language-model song generator adapted for multi-task editing, supporting segment-wise and track-wise modifications alongside full-song generation.
-
Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation
Boomerang sampling, applied to a pretrained music diffusion model, creates audio variations that improve beat tracking when training data is scarce and can change instruments via text prompts.
-
WAKE: Watermarking Audio with Key Enrichment
WAKE embeds and decodes multiple 32-bit audio watermarks with separate 8-bit keys using an invertible neural network, avoiding the overwriting problem in existing systems.
-
VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features
An adaptation of MusicGen that adds semantic conditioning from CLIP global features and rhythmic conditioning from CLIP local inter-frame similarity, trained in two stages, generates video-aligned background music.
-
Compression of Higher Order Ambisonics with Multichannel RVQGAN
A multichannel RVQGAN with a covariance loss compresses 16-channel third-order Ambisonics to 16 kbps and outperforms Opus at 160 kbps in a MUSHRA listening test on ambient scenes.
-
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
USC is a low-bitrate codec whose first token stream keeps content, prosody, and sentiment while removing speaker identity, enabling privacy-preserving speech LLM training.
-
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
A narrative review of 30 recent papers classifies AI sound-effect generators by input modality and summarizes reported progress and remaining limitations in temporal sync, evaluation, and controllability.
-
ASAudio: A Survey of Advanced Spatial Audio Research
A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.
-
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.
Discussion (0). Continue with ORCID to comment.