Pith. sign in

REVIEW 5 cited by

Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.11680 v1 pith:6ZAS3DOG submitted 2021-09-23 cs.CL cs.LGcs.SD

classification cs.CLcs.LGcs.SD
keywords languagescross-lingualdatalabeledlearningmodelpretrainedrecognition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data. However, in many cases there is labeled data available for related languages which is not utilized by these methods. This paper extends previous work on zero-shot cross-lingual transfer learning by fine-tuning a multilingually pretrained wav2vec 2.0 model to transcribe unseen languages. This is done by mapping phonemes of the training languages to the target language using articulatory features. Experiments show that this simple method significantly outperforms prior work which introduced task-specific architectures and used only part of a monolingually pretrained model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    PhonoQ-2.0 directly predicts structured phonological features from self-supervised models with a gating mechanism, outperforming phoneme baselines by 8+ macro-F1 points on average across in-domain, out-of-domain, and ...

  2. Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Training on within-listener comparison scores with a unified scoring scale, without listener embeddings, improves speech quality and continuous emotion prediction over listener-averaged baselines.

  3. FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A three-stage fusion of a sparse compression separator, an SSL-conditioned codec-token generator, and a fusion network achieves third place in URGENT 2025 and improves perceptual metrics with a modest fidelity trade-off.

  4. Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

    cs.LG 2025-05 conditional novelty 5.0 of 10

    On a single-speaker MRI speech corpus, adding vocal-tract video to audio does not improve phoneme recognition, but attention analysis shows articulatory cues can lead acoustic cues in time.

  5. Voice Adaptation for Swiss German

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...

Pith tools