Pith. sign in

REVIEW

Controllable and Interpretable Singing Voice Decomposition via Assem-VC

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.12676 v1 pith:7SFWSTDM submitted 2021-10-25 eess.AS cs.SD

classification eess.AScs.SD
keywords singingvoicespeakertargetassem-vcdecompositionconclusioncontent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose a singing decomposition system that encodes time-aligned linguistic content, pitch, and source speaker identity via Assem-VC. With decomposed speaker-independent information and the target speaker's embedding, we could synthesize the singing voice of the target speaker. In conclusion, we made a perfectly synced duet with the user's singing voice and the target singer's converted singing voice.

Discussion (0). Continue with ORCID to comment.

Pith tools