Pith. sign in

REVIEW 4 cited by

SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17645 v2 pith:RPIGQBBF submitted 2024-02-27 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords lyricssongmelodiescompositiongenerationmodelsongcomposermelody
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Creating lyrics and melodies for the vocal track in a symbolic format, known as song composition, demands expert musical knowledge of melody, an advanced understanding of lyrics, and precise alignment between them. Despite achievements in sub-tasks such as lyric generation, lyric-to-melody, and melody-to-lyric, etc, a unified model for song composition has not yet been achieved. In this paper, we introduce SongComposer, a pioneering step towards a unified song composition model that can readily create symbolic lyrics and melodies following instructions. SongComposer is a music-specialized large language model (LLM) that, for the first time, integrates the capability of simultaneously composing lyrics and melodies into LLMs by leveraging three key innovations: 1) a flexible tuple format for word-level alignment of lyrics and melodies, 2) an extended tokenizer vocabulary for song notes, with scalar initialization based on musical knowledge to capture rhythm, and 3) a multi-stage pipeline that captures musical structure, starting with motif-level melody patterns and progressing to phrase-level structure for improved coherence. Extensive experiments demonstrate that SongComposer outperforms advanced LLMs, including GPT-4, in tasks such as lyric-to-melody generation, melody-to-lyric generation, song continuation, and text-to-song creation. Moreover, we will release SongCompose, a large-scale dataset for training, containing paired lyrics and melodies in Chinese and English.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DanceChat: Large Language Model-Guided Music-to-Dance Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An LLM-generated text choreography, fused with music and beat features, guides a diffusion model to produce more diverse and physically plausible dance motion, with a multi-modal alignment loss intended to bridge musi...

  2. Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    A bar-level symbolic-score song generator (BACH) is claimed to beat published systems and commercial Suno on human-rated quality, duration, and efficiency, but the supporting full text is corrupted and unverifiable.

  3. SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling

    cs.SD 2025-06 reject novelty 3.0 of 10

    The paper announces a large-scale music metadata and link dataset from Genius, but the lack of access, code, and validation makes its claimed utility unverifiable.

  4. Content filtering methods for music recommendation: A review

    cs.IR 2025-07 conditional

    A survey of content-based music recommendation methods, including audio analysis, lyrics analysis, and context awareness, with no new experimental results.

Pith tools