Pith. sign in

REVIEW 1 cited by

CNN Encoding of Acoustic Parameters for Prominence Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.05488 v3 pith:SY2AR6BI submitted 2021-04-12 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords readingacrossacousticcontextdetectiondifferentfeaturefeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Expressive reading, considered the defining attribute of oral reading fluency, comprises the prosodic realization of phrasing and prominence. In the context of evaluating oral reading, it helps to establish the speaker's comprehension of the text. We consider a labeled dataset of children's reading recordings for the speaker-independent detection of prominent words using acoustic-prosodic and lexico-syntactic features. A previous well-tuned random forest ensemble predictor is replaced by an RNN sequence classifier to exploit potential context dependency across the longer utterance. Further, deep learning is applied to obtain word-level features from low-level acoustic contours of fundamental frequency, intensity and spectral shape in an end-to-end fashion. Performance comparisons are presented across the different feature types and across different feature learning architectures for prominent word prediction to draw insights wherever possible.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pitch Accent Detection improves Pretrained Automatic Speech Recognition

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Jointly training pitch accent detection with ASR on wav2vec2 reduces LibriSpeech WER from 6.0 to 4.3 in a one-hour fine-tuning setting.

Pith tools