Pith. sign in

REVIEW 1 cited by

A Multi-Discriminator CycleGAN for Unsupervised Non-Parallel Speech Domain Adaptation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.00522 v4 pith:LNM4ZHZJ submitted 2018-03-27 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelspeechadaptationdomaincyclegandatadiscriminatorserror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Domain adaptation plays an important role for speech recognition models, in particular, for domains that have low resources. We propose a novel generative model based on cyclic-consistent generative adversarial network (CycleGAN) for unsupervised non-parallel speech domain adaptation. The proposed model employs multiple independent discriminators on the power spectrogram, each in charge of different frequency bands. As a result we have 1) better discriminators that focus on fine-grained details of the frequency features, and 2) a generator that is capable of generating more realistic domain-adapted spectrogram. We demonstrate the effectiveness of our method on speech recognition with gender adaptation, where the model only has access to supervised data from one gender during training, but is evaluated on the other at test time. Our model is able to achieve an average of $7.41\%$ on phoneme error rate, and $11.10\%$ word error rate relative performance improvement as compared to the baseline, on TIMIT and WSJ dataset, respectively. Qualitatively, our model also generates more natural sounding speech, when conditioned on data from the other domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Customizing Speech Recognition Model with Large Language Model Feedback

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLM log-probability scores combined with acoustic scores serve as RL rewards to adapt ASR models to new domains without labeled data.

Pith tools