{"paper":{"title":"Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Adapting TTS to an accent with under ten utterances and LLM phoneme edits produces synthetic data that reduces ASR word error rates on real accented speech.","cross_cats":[],"primary_cat":"cs.SD","authors_text":"Dilek Hakkani-T\\\"ur, Mark Hasegawa-Johnson, Nimet Beyza Bozdag, Volodymyr Kindratenko, Yurii Halychanskyi","submitted_at":"2026-04-30T00:05:03Z","abstract_excerpt":"Synthetic accented speech is a promising way to improve automatic speech recognition (ASR) when real accented recordings are scarce. We ask what makes such data useful for ASR fine-tuning: target-accent phoneme edits that expose the recognizer to accent-specific pronunciations, or random phoneme perturbations that act as augmentation in phoneme space. In a few-shot TTS pipeline, we compare LLM-generated accent edits with matched-rate random substitutions and oracle controls using ground-truth accented phonemes and prosody. Random substitutions recover much of the ASR gain: LLM target-accent ed"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Experiments demonstrate consistent word error rate (WER) reductions on real accented speech, including cross-speaker evaluation and ultra-low data regimes. A matched-rate random phoneme baseline shows that phoneme-space perturbation itself is a strong form of augmentation, while LLM-guided edits provide additional gains through accent-conditioned structure.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That LLM-based phoneme editing, guided by fewer than ten reference utterances, reliably produces accent-conditioned pronunciations whose synthetic speech transfers to measurable improvements on real accented ASR test sets.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Few-shot TTS adaptation combined with LLM-guided phoneme editing produces synthetic accented speech that improves ASR word error rates on real accented audio even in cross-speaker and ultra-low-data settings.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Adapting TTS to an accent with under ten utterances and LLM phoneme edits produces synthetic data that reduces ASR word error rates on real accented speech.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"fb145624292c262c98dd17bd7574132c8726198bd0f10af269111174862a03a3"},"source":{"id":"2604.27273","kind":"arxiv","version":2},"verdict":{"id":"73593e90-a4cf-4fbb-aa30-732bdac13838","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-07T09:01:57.698631Z","strongest_claim":"Experiments demonstrate consistent word error rate (WER) reductions on real accented speech, including cross-speaker evaluation and ultra-low data regimes. A matched-rate random phoneme baseline shows that phoneme-space perturbation itself is a strong form of augmentation, while LLM-guided edits provide additional gains through accent-conditioned structure.","one_line_summary":"Few-shot TTS adaptation combined with LLM-guided phoneme editing produces synthetic accented speech that improves ASR word error rates on real accented audio even in cross-speaker and ultra-low-data settings.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That LLM-based phoneme editing, guided by fewer than ten reference utterances, reliably produces accent-conditioned pronunciations whose synthetic speech transfers to measurable improvements on real accented ASR test sets.","pith_extraction_headline":"Adapting TTS to an accent with under ten utterances and LLM phoneme edits produces synthetic data that reduces ASR word error rates on real accented speech."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.27273/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"ai_meta_artifact","ran_at":"2026-05-20T22:40:44.099732Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-19T19:25:11.615394Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"f792094dfca903e06ac8fb1355a90a6d22dbf84f6d0bcdffac1a18ca794813ce"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}