Pith. sign in

REVIEW 2 cited by

Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09313 v3 pith:N7SZM6RP submitted 2024-04-14 eess.AS cs.AI

classification eess.AScs.AI
keywords synthesisgenerationmelodistmusicsingingsongtext-to-songvoice
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel task called text-to-song synthesis which incorporating both vocals and accompaniments generation. We develop Melodist, a two-stage text-to-song method that consists of singing voice synthesis (SVS) and vocal-to-accompaniment (V2A) synthesis. Melodist leverages tri-tower contrastive pretraining to learn more effective text representation for controllable V2A synthesis. A Chinese song dataset mined from a music website is built up to alleviate data scarcity for our research. The evaluation results on our dataset demonstrate that Melodist can synthesize songs with comparable quality and style consistency. Audio samples can be found in https://text2songMelodist.github.io/Sample/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

    cs.SD 2025-07 conditional novelty 6.0 of 10

    JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.

  2. DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

    eess.AS 2025-07 conditional novelty 6.0 of 10

    DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.

Pith tools