Pith. sign in

REVIEW 1 cited by

Towards End-to-End Spoken Grammatical Error Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05550 v2 pith:M3PUHA63 submitted 2023-11-09 cs.CL cs.LGeess.AS

classification cs.CLcs.LGeess.AS
keywords end-to-endspokenapproachescascadeddatadisfluencyfeedbackgrammatical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, with the associated concern of propagating errors between these individual modules. In this paper, we introduce an alternative "end-to-end" approach to spoken GEC, exploiting a speech recognition foundation model, Whisper. This foundation model can be used to replace the whole framework or part of it, e.g., ASR and disfluency removal. These end-to-end approaches are compared to more standard cascaded approaches on the data obtained from a free-speaking spoken language assessment test, Linguaskill. Results demonstrate that end-to-end spoken GEC is possible within this architecture, but the lack of available data limits current performance compared to a system using large quantities of text-based GEC data. Conversely, end-to-end disfluency detection and removal, which is easier for the attention-based Whisper to learn, does outperform cascaded approaches. Additionally, the paper discusses the challenges of providing feedback to candidates when using end-to-end systems for spoken GEC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction

    cs.CL 2024-12 reject novelty 4.0 of 10

    A two-level curriculum, batch ordering by loss and instance/token reweighting by Monte Carlo dropout confidence, yields about 0.5 to 1.2 F0.5 gains for BART, mT5, and SynGEC on NLPCC and MuCGEC.

Pith tools