REVIEW 4 cited by
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Creating lyrics and melodies for the vocal track in a symbolic format, known as song composition, demands expert musical knowledge of melody, an advanced understanding of lyrics, and precise alignment between them. Despite achievements in sub-tasks such as lyric generation, lyric-to-melody, and melody-to-lyric, etc, a unified model for song composition has not yet been achieved. In this paper, we introduce SongComposer, a pioneering step towards a unified song composition model that can readily create symbolic lyrics and melodies following instructions. SongComposer is a music-specialized large language model (LLM) that, for the first time, integrates the capability of simultaneously composing lyrics and melodies into LLMs by leveraging three key innovations: 1) a flexible tuple format for word-level alignment of lyrics and melodies, 2) an extended tokenizer vocabulary for song notes, with scalar initialization based on musical knowledge to capture rhythm, and 3) a multi-stage pipeline that captures musical structure, starting with motif-level melody patterns and progressing to phrase-level structure for improved coherence. Extensive experiments demonstrate that SongComposer outperforms advanced LLMs, including GPT-4, in tasks such as lyric-to-melody generation, melody-to-lyric generation, song continuation, and text-to-song creation. Moreover, we will release SongCompose, a large-scale dataset for training, containing paired lyrics and melodies in Chinese and English.
Forward citations
Cited by 4 Pith papers
-
DanceChat: Large Language Model-Guided Music-to-Dance Generation
An LLM-generated text choreography, fused with music and beat features, guides a diffusion model to produce more diverse and physically plausible dance motion, with a multi-modal alignment loss intended to bridge musi...
-
Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation
A bar-level symbolic-score song generator (BACH) is claimed to beat published systems and commercial Suno on human-rated quality, duration, and efficiency, but the supporting full text is corrupted and unverifiable.
-
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
The paper announces a large-scale music metadata and link dataset from Genius, but the lack of access, code, and validation makes its claimed utility unverifiable.
-
Content filtering methods for music recommendation: A review
A survey of content-based music recommendation methods, including audio analysis, lyrics analysis, and context awareness, with no new experimental results.
Discussion (0). Sign in to comment.