Pith. sign in

End-to-End Speech Recognition and Disfluency Removal with Acoustic Language Model Pretraining

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The SOTA in transcription of disfluent and conversational speech has in recent years favored two-stage models, with separate transcription and cleaning stages. We believe that previous attempts at end-to-end disfluency removal have fallen short because of the representational advantage that large-scale language model pretraining has given to lexical models. Until recently, the high dimensionality and limited availability of large audio datasets inhibited the development of large-scale self-supervised pretraining objectives for learning effective audio representations, giving a relative advantage to the two-stage approach, which utilises pretrained representations for lexical tokens. In light of recent successes in large scale audio pretraining, we revisit the performance comparison between two-stage and end-to-end model and find that audio based language models pretrained using weak self-supervised objectives match or exceed the performance of similarly trained two-stage models, and further, that the choice of pretraining objective substantially effects a model's ability to be adapted to the disfluency removal task.

citation-role summary

background 1

citation-polarity summary

fields

cs.HC 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Rambler in the Wild: A Diary Study of LLM-Assisted Writing With Speech cs.HC · 2025-02-08 · conditional · none · ref 2 · internal anchor

    In a ten-day field study, twelve writers used an LLM-assisted dictation tool and reported that speaking their drafts helped productivity and emotional expression, with nine of twelve adapting to the new paradigm.