REVIEW 5 cited by
Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data. However, in many cases there is labeled data available for related languages which is not utilized by these methods. This paper extends previous work on zero-shot cross-lingual transfer learning by fine-tuning a multilingually pretrained wav2vec 2.0 model to transcribe unseen languages. This is done by mapping phonemes of the training languages to the target language using articulatory features. Experiments show that this simple method significantly outperforms prior work which introduced task-specific architectures and used only part of a monolingually pretrained model.
Forward citations
Cited by 5 Pith papers
-
Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
PhonoQ-2.0 directly predicts structured phonological features from self-supervised models with a gating mechanism, outperforming phoneme baselines by 8+ macro-F1 points on average across in-domain, out-of-domain, and ...
-
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
Training on within-listener comparison scores with a unified scoring scale, without listener embeddings, improves speech quality and continuous emotion prediction over listener-averaged baselines.
-
FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge
A three-stage fusion of a sparse compression separator, an SSL-conditioned codec-token generator, and a fusion network achieves third place in URGENT 2025 and improves perceptual metrics with a modest fidelity trade-off.
-
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
On a single-speaker MRI speech corpus, adding vocal-tract video to audio does not improve phoneme recognition, but attention analysis shows articulatory cues can lead acoustic cues in time.
-
Voice Adaptation for Swiss German
Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...
Discussion (0). Sign in to comment.