Pith. sign in

REVIEW

Using Speech Synthesis to Train End-to-End Spoken Language Understanding Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.09463 v1 pith:K3KQLACE submitted 2019-10-21 eess.AS

classification eess.AS
keywords speechdataend-to-endlanguageunderstandingapproachmodelsspoken
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end models are an attractive new approach to spoken language understanding (SLU) in which the meaning of an utterance is inferred directly from the raw audio without employing the standard pipeline composed of a separately trained speech recognizer and natural language understanding module. The downside of end-to-end SLU is that in-domain speech data must be recorded to train the model. In this paper, we propose a strategy for overcoming this requirement in which speech synthesis is used to generate a large synthetic training dataset from several artificial speakers. Experiments on two open-source SLU datasets confirm the effectiveness of our approach, both as a sole source of training data and as a form of data augmentation.

Discussion (0). Sign in to comment.

Pith tools