REVIEW 1 cited by
Parsing Speech: A Neural Approach to Integrating Lexical and Acoustic-Prosodic Information
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In conversational speech, the acoustic signal provides cues that help listeners disambiguate difficult parses. For automatically parsing spoken utterances, we introduce a model that integrates transcribed text and acoustic-prosodic features using a convolutional neural network over energy and pitch trajectories coupled with an attention-based recurrent neural network that accepts text and prosodic features. We find that different types of acoustic-prosodic features are individually helpful, and together give statistically significant improvements in parse and disfluency detection F1 scores over a strong text-only baseline. For this study with known sentence boundaries, error analyses show that the main benefit of acoustic-prosodic features is in sentences with disfluencies, attachment decisions are most improved, and transcription errors obscure gains from prosody.
Forward citations
Cited by 1 Pith paper
-
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
Jointly training pitch accent detection with ASR on wav2vec2 reduces LibriSpeech WER from 6.0 to 4.3 in a one-hour fine-tuning setting.
Discussion (0). Sign in to comment.