Pith. sign in

REVIEW 2 cited by

Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11780 v3 pith:AJ34KW2E submitted 2024-03-18 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords audioprompt-singercontrolcontrollingdataenableslanguagemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent singing-voice-synthesis (SVS) methods have achieved remarkable audio quality and naturalness, yet they lack the capability to control the style attributes of the synthesized singing explicitly. We propose Prompt-Singer, the first SVS method that enables attribute controlling on singer gender, vocal range and volume with natural language. We adopt a model architecture based on a decoder-only transformer with a multi-scale hierarchy, and design a range-melody decoupled pitch representation that enables text-conditioned vocal range control while keeping melodic accuracy. Furthermore, we explore various experiment settings, including different types of text representations, text encoder fine-tuning, and introducing speech data to alleviate data scarcity, aiming to facilitate further research. Experiments show that our model achieves favorable controlling ability and audio quality. Audio samples are available at http://prompt-singer.github.io .

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

    cs.SD 2025-07 conditional novelty 6.0 of 10

    STARS unifies lyric alignment, note transcription, vocal technique detection, and global style prediction into one multi-level neural model that matches or beats several single-task baselines.

  2. SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling

    cs.SD 2025-06 reject novelty 3.0 of 10

    The paper announces a large-scale music metadata and link dataset from Genius, but the lack of access, code, and validation makes its claimed utility unverifiable.

Pith tools