Pith. sign in

REVIEW 1 cited by

Selective Attention Merging for low resource tasks: A case study of Child ASR

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.08468 v1 pith:LZV6D5FP submitted 2025-01-14 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechtasksattentiondatalow-resourcemergemergingmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While Speech Foundation Models (SFMs) excel in various speech tasks, their performance for low-resource tasks such as child Automatic Speech Recognition (ASR) is hampered by limited pretraining data. To address this, we explore different model merging techniques to leverage knowledge from models trained on larger, more diverse speech corpora. This paper also introduces Selective Attention (SA) Merge, a novel method that selectively merges task vectors from attention matrices to enhance SFM performance on low-resource tasks. Experiments on the MyST database show significant reductions in relative word error rate of up to 14%, outperforming existing model merging and data augmentation techniques. By combining data augmentation techniques with SA Merge, we achieve a new state-of-the-art WER of 8.69 on the MyST database for the Whisper-small model, highlighting the potential of SA Merge for improving low-resource ASR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust fine-tuning of speech recognition models via model merging: application to disordered speech

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Merging multiple fine-tuned Whisper models reduces word error rate on dysarthric speech by 12-16% relative to standard fine-tuning, with gains on long audio and low-data settings.

Pith tools