Pith. sign in

REVIEW 4 cited by

MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01546 v1 pith:4EHUUS3L submitted 2023-08-03 cs.SD cs.AIcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.LGcs.MMeess.AS
keywords musicmixupdatabeat-synchronousmodelmusicldmstrategiestraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited availability of music data and sensitive issues related to copyright and plagiarism. In this paper, to tackle these challenges, we first construct a state-of-the-art text-to-music model, MusicLDM, that adapts Stable Diffusion and AudioLDM architectures to the music domain. We achieve this by retraining the contrastive language-audio pretraining model (CLAP) and the Hifi-GAN vocoder, as components of MusicLDM, on a collection of music data samples. Then, to address the limitations of training data and to avoid plagiarism, we leverage a beat tracking model and propose two different mixup strategies for data augmentation: beat-synchronous audio mixup and beat-synchronous latent mixup, which recombine training audio directly or via a latent embeddings space, respectively. Such mixup strategies encourage the model to interpolate between musical training samples and generate new music within the convex hull of the training data, making the generated music more diverse while still staying faithful to the corresponding style. In addition to popular evaluation metrics, we design several new evaluation metrics based on CLAP score to demonstrate that our proposed MusicLDM and beat-synchronous mixup strategies improve both the quality and novelty of generated music, as well as the correspondence between input text and generated music.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

    cs.SD 2025-07 conditional novelty 6.0 of 10

    JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.

  2. The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Word-based style descriptors generated by an LLM can shift MusicGen outputs toward a target artist's sound almost as much as using the artist's name, defining a name-free gap.

  3. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

    cs.SD 2026-07 reject novelty 4.0 of 10

    FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.

  4. WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

    cs.SD 2025-06 conditional novelty 4.0 of 10

    WhisQ uses Whisper and Qwen with co-attention and optimal transport to predict music quality and text-alignment scores, but its reported improvements do not match its own data.

Pith tools