Pith. sign in

REVIEW 1 cited by

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.14910 v3 pith:IXJWZB2F submitted 2025-05-20 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords singingstyletcsingervoicezero-shotmultilingualpromptssynthesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting their robustness in zero-shot scenarios and producing poor transitions between phonemes and notes. Moreover, they also lack effective multi-level style control via diverse prompts. To overcome these challenges, we introduce TCSinger 2, a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts. TCSinger 2 mainly includes three key modules: 1) Blurred Boundary Content (BBC) Encoder, predicts duration, extends content embedding, and applies masking to the boundaries to enable smooth transitions. 2) Custom Audio Encoder, uses contrastive learning to extract aligned representations from singing, speech, and textual prompts. 3) Flow-based Custom Transformer, leverages Cus-MOE, with F0 supervision, enhancing both the synthesis quality and style modeling of the generated singing voice. Experimental results show that TCSinger 2 outperforms baseline models in both subjective and objective metrics across multiple related tasks. Singing voice samples are available at https://aaronz345.github.io/TCSinger2Demo/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Conan achieves chunkwise online zero-shot voice conversion, preserving source content while adopting the reference speaker's timbre and style, with a latency as low as 37 milliseconds.

Pith tools