Pith. sign in

REVIEW 3 cited by

TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12572 v2 pith:TJA2KF4F submitted 2025-02-18 cs.SD

classification cs.SD
keywords singingcontroltechniquetechsingervoicevoicesmodelsynthesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Singing voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy tones, thus limiting the expressive potential of synthetic voices. We introduce TechSinger, an advanced system for controllable singing voice synthesis that supports five languages and seven vocal techniques. TechSinger leverages a flow-matching-based generative model to produce singing voices with enhanced expressive control over various techniques. To enhance the diversity of training data, we develop a technique detection model that automatically annotates datasets with phoneme-level technique labels. Additionally, our prompt-based technique prediction model enables users to specify desired vocal attributes through natural language, offering fine-grained control over the synthesized singing. Experimental results demonstrate that TechSinger significantly enhances the expressiveness and realism of synthetic singing voices, outperforming existing methods in terms of audio quality and technique-specific control. Audio samples can be found at https://gwx314.github.io/tech-singer/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  2. STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

    cs.SD 2025-07 conditional novelty 6.0 of 10

    STARS unifies lyric alignment, note transcription, vocal technique detection, and global style prediction into one multi-level neural model that matches or beats several single-task baselines.

  3. TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

    eess.AS 2025-05 conditional novelty 6.0 of 10

    TCSinger 2 generates zero-shot singing voices in nine languages with style transfer from audio prompts and multi-level style control from natural language prompts, using blurred boundary encoders, contrastive prompt a...

Pith tools