Pith. sign in

REVIEW 21 cited by

Long-form music generation with latent diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.10301 v2 pith:AGWJTH2U submitted 2024-04-16 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords musiclatentcoherentfull-lengthgenerativelong-formmodelproduce
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Latent Swap Joint Diffusion for 2D Long-Form Latent Generation

    cs.SD 2025-02 conditional novelty 7.0 of 10

    A training-free latent swap method that replaces averaging with binary swapping in joint diffusion, improving long-form audio spectrum and panorama generation.

  2. MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A text-to-music system that first writes an explicit structural layout and then generates audio from it shows better long-range boundary agreement than a matched model without the layout.

  3. On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.

  4. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  5. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  6. In-the-wild Audio Spatialization with Flexible Text-guided Localization

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.

  7. Fast Text-to-Audio Generation with Adversarial Post-Training

    cs.SD 2025-05 conditional novelty 6.0 of 10

    ARC post-training speeds up text-to-audio generation to near-real-time speeds on GPUs and a few seconds on phones, without distillation or classifier-free guidance.

  8. Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding

    cs.SD 2025-05 conditional novelty 6.0 of 10

    MUFFIN is a neural audio codec that quantizes separate latent frequency bands with dedicated codebooks, claiming SOTA reconstruction quality and a competitive 12.5 Hz ultra-low-rate variant.

  9. Bridge-SR: Schr\"odinger Bridge for Efficient SR

    cs.SD 2025-01 conditional novelty 6.0 of 10

    Bridge-SR applies tractable Schrödinger bridge models to waveform-domain speech super-resolution, and with 1.7M parameters reports the lowest log-spectral distance on VCTK while matching diffusion quality at 4 sampling steps.

  10. TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.

  11. ETTA: Elucidating the Design Space of Text-to-Audio Models

    cs.SD 2024-12 conditional novelty 6.0 of 10

    ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.

  12. FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

    cs.SD 2024-12 conditional novelty 6.0 of 10

    FolAI predicts an editable RMS envelope from silent video and uses it, with semantic embeddings, to condition a Stable Audio diffusion model for 44.1 kHz stereo foley generation.

  13. SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

    eess.AS 2024-12 conditional novelty 6.0 of 10

    A language-model song generator adapted for multi-task editing, supporting segment-wise and track-wise modifications alongside full-song generation.

  14. Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Boomerang sampling, applied to a pretrained music diffusion model, creates audio variations that improve beat tracking when training data is scarce and can change instruments via text prompts.

  15. WAKE: Watermarking Audio with Key Enrichment

    cs.SD 2025-06 conditional novelty 5.0 of 10

    WAKE embeds and decodes multiple 32-bit audio watermarks with separate 8-bit keys using an invertible neural network, avoiding the overwriting problem in existing systems.

  16. VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

    cs.SD 2024-12 conditional novelty 5.0 of 10

    An adaptation of MusicGen that adds semantic conditioning from CLIP global features and rhythmic conditioning from CLIP local inter-frame similarity, trained in two stages, generates video-aligned background music.

  17. Compression of Higher Order Ambisonics with Multichannel RVQGAN

    cs.SD 2024-11 conditional novelty 5.0 of 10

    A multichannel RVQGAN with a covariance loss compresses 16-channel third-order Ambisonics to 16 kbps and outperforms Opus at 160 kbps in a MUSHRA listening test on ambient scenes.

  18. Universal Semantic Disentangled Privacy-preserving Speech Representation Learning

    eess.AS 2025-05 conditional novelty 4.0 of 10

    USC is a low-bitrate codec whose first token stream keeps content, prosody, and sentiment while removing speaker identity, enabling privacy-preserving speech LLM training.

  19. AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

    cs.SD 2026-08 conditional novelty 3.0 of 10

    A narrative review of 30 recent papers classifies AI sound-effect generators by input modality and summarizes reported progress and remaining limitations in temporal sync, evaluation, and controllability.

  20. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

  21. Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

    cs.SD 2024-11 conditional novelty 2.0 of 10

    A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.

Pith tools