Pith. sign in

REVIEW 2 cited by

SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.10429 v2 pith:HA542PIA submitted 2024-09-16 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords speechlanguagessmilein-contextlearninglow-resourceadaptationautomatic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due to limited training data and insufficient cross-lingual generalization. Existing adaptation strategies, such as shallow fusion, data augmentation, and direct fine-tuning, either rely on external resources, suffer computational inefficiencies, or fail in test-time adaptation scenarios. To address these limitations, we introduce Speech Meta In-Context LEarning (SMILE), an innovative framework that combines meta-learning with speech in-context learning (SICL). SMILE leverages meta-training from high-resource languages to enable robust, few-shot generalization to low-resource languages without explicit fine-tuning on the target domain. Extensive experiments on the ML-SUPERB benchmark show that SMILE consistently outperforms baseline methods, significantly reducing character and word error rates in training-free few-shot multilingual ASR tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advancing STT for Low-Resource Real-World Speech

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Fine-tuning Whisper on a new 303-hour corpus of spontaneous Swiss German broadcast speech improves WER from 21% to 17% for large-v3 and outperforms models trained on sentence-level corpora.

  2. In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Providing 12 in-context audio-text examples reduces Phi-4-Multimodal's average word error rate by 19.7% relative across English varieties, with the largest gains for low-resource accents.

Pith tools