Providing 12 in-context audio-text examples reduces Phi-4-Multimodal's average word error rate by 19.7% relative across English varieties, with the largest gains for low-resource accents.
14 Table 3: Complete speaker-level Word Error Rates (WER) for 0-shot and 12-shot conditions
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
Providing 12 in-context audio-text examples reduces Phi-4-Multimodal's average word error rate by 19.7% relative across English varieties, with the largest gains for low-resource accents.