Pith. sign in

REVIEW 1 cited by

Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.00794 v2 pith:N53TSNVD submitted 2019-06-03 cs.LG cs.SDeess.ASstat.ML

classification cs.LGcs.SDeess.ASstat.ML
keywords blowconversiondatanon-parallelvoiceaudioend-to-endflow
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations. Voice conversion, in which a model has to impersonate a speaker in a recording, is one of those situations. In this paper, we propose Blow, a single-scale normalizing flow using hypernetwork conditioning to perform many-to-many voice conversion between raw audio. Blow is trained end-to-end, with non-parallel data, on a frame-by-frame basis using a single speaker identifier. We show that Blow compares favorably to existing flow-based architectures and other competitive baselines, obtaining equal or better performance in both objective and subjective evaluations. We further assess the impact of its main components with an ablation study, and quantify a number of properties such as the necessary amount of training data or the preference for source or target speakers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning

    cs.SD 2025-01 reject novelty 5.0 of 10

    Stepback trains a voice converter with two decoders and a self-destructive loss to separate speaker identity from linguistic content, but the preprint contains no reported evaluation results.

Pith tools