REVIEW 3 major objections 1 minor 32 references
How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA
T0 review · 3 major / 1 minor · reviewed 2026-07-01 · grok-4.3
Pith's one-line read SSL speech models predict articulatory movements from bilingual Finnish-Russian speakers with Pearson r up to 0.68 using only about 5 minutes of training data.
desk verdict Bilingual EMA data lets them compare SSL models on articulatory prediction across Finnish and Russian with the same speakers, but the abstract leaves the controls for speaker and proficiency effects unclear. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Cross-lingual probing that correlates SSL latent representations with EMA-tracked articulatory movements from the same bilingual speakers.
What would settle it
Finding that prediction correlations fall to near zero when the same model layers are tested on held-out bilingual speakers or when language-specific biases in the EMA data are removed.
Extended reading notes
Core claim
Using EMA data from bilingual Finnish-Russian speakers, the study measures how well SSL model representations predict articulatory trajectories. Models reach Pearson correlations of 0.68 with roughly 5 minutes of training data per speaker. Multilingual pretraining yields higher accuracy than monolingual pretraining. Intermediate layers align best with the articulatory signals, tongue sensors correlate more strongly than lip sensors, read speech outperforms spontaneous speech, and results generalize across speaker proficiency levels.
Load-bearing premise
The articulatory movements recorded by EMA from bilingual speakers form a language-independent ground truth that can be compared directly between Finnish and Russian.
Editorial extensions
If this is right
- Multilingual SSL models capture articulatory features more effectively than monolingual ones across these two languages.
- Intermediate layers of the models are the most useful for recovering articulatory information.
- Tongue movements can be recovered from model representations more reliably than lip movements.
- Performance remains high for both read and spontaneous speech and across proficiency levels.
Reading between the lines
- If the same pattern holds for other language pairs, SSL models could support articulatory-based tools in low-resource languages without new EMA collections.
- The layer-wise pattern suggests that articulatory supervision could be added most efficiently at intermediate depths during fine-tuning.
- Higher accuracy on read speech implies that controlled elicitation tasks may be sufficient to benchmark articulatory encoding.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that self-supervised learning (SSL) speech models encode articulatory information in a cross-lingual manner, evaluated via EMA data from bilingual Finnish-Russian speakers. Multilingual models outperform monolingual ones with Pearson r up to 0.68 even using ~5 minutes of training data; intermediate layers are most effective, tongue movements are more predictable than lip movements, and performance is higher for read speech with generalization across proficiency levels.
Significance. If the cross-lingual comparisons are robust to speaker-specific effects, the results would strengthen the case that multilingual SSL models capture more language-independent articulatory features than monolingual counterparts. The small-data probing efficiency and layer-wise findings could aid interpretability efforts and guide applications in speech technology that rely on articulatory grounding.
major comments (3)
- [Abstract] Abstract: The abstract states performance numbers (Pearson r up to 0.68) and qualitative trends but supplies no information on data partitioning, statistical testing, baseline comparisons, or controls for speaker identity, so it is impossible to judge whether the reported correlations genuinely support the cross-lingual claim.
- [Cross-lingual evaluation setup] Cross-lingual evaluation setup: The central claim rests on treating EMA recordings from bilingual speakers as directly comparable, language-independent targets. Without explicit quantification of residual speaker variance or a monolingual-speaker control condition, speaker-idiosyncratic or L1-dominant signatures could inflate the reported multilingual advantage and layer-wise patterns.
- [Results on task and proficiency] Results on task and proficiency: The abstract states that proficiency and task effects were assessed, yet the absence of details on how residual variance was isolated or quantified leaves open whether these factors confound the multilingual vs. monolingual comparison.
minor comments (1)
- [Methods] The methods section would benefit from explicit description of the exact SSL models, probe architecture, and normalization procedures to support reproducibility.
Simulated Author's Rebuttal
Thank you for the opportunity to respond to the referee's report. We address each of the major comments point by point below, indicating where revisions will be made to the manuscript.
read point-by-point responses
-
Referee: [Abstract] Abstract: The abstract states performance numbers (Pearson r up to 0.68) and qualitative trends but supplies no information on data partitioning, statistical testing, baseline comparisons, or controls for speaker identity, so it is impossible to judge whether the reported correlations genuinely support the cross-lingual claim.
Authors: We agree with this observation. The revised abstract will incorporate details on the data partitioning (speaker-independent cross-validation), statistical significance testing of the Pearson correlations, baseline comparisons (e.g., against MFCC features), and the use of bilingual speakers as a control for speaker identity. These additions will strengthen the abstract's support for the cross-lingual claims. revision: yes
-
Referee: [Cross-lingual evaluation setup] Cross-lingual evaluation setup: The central claim rests on treating EMA recordings from bilingual speakers as directly comparable, language-independent targets. Without explicit quantification of residual speaker variance or a monolingual-speaker control condition, speaker-idiosyncratic or L1-dominant signatures could inflate the reported multilingual advantage and layer-wise patterns.
Authors: Our experimental design uses bilingual speakers to ensure that language comparisons are within the same speaker, thereby controlling for speaker-specific effects. We report per-speaker results and inter-speaker variance in the paper to quantify residual speaker variance. A monolingual control would necessitate new recordings and is beyond the scope of this work; however, we will expand the discussion to explicitly address potential L1 influences and why the bilingual setup mitigates speaker confounds. revision: partial
-
Referee: [Results on task and proficiency] Results on task and proficiency: The abstract states that proficiency and task effects were assessed, yet the absence of details on how residual variance was isolated or quantified leaves open whether these factors confound the multilingual vs. monolingual comparison.
Authors: We will add detailed explanations in the methods and results sections on the statistical approaches used to isolate task and proficiency effects, including mixed-effects modeling to quantify their contribution to variance while assessing the multilingual advantage. This will demonstrate that these factors were properly controlled and do not confound the main findings. revision: yes
Circularity Check
No circularity: purely empirical correlations on held-out data
full rationale
The paper reports measured Pearson r values (up to 0.68) between SSL latent features and EMA targets on held-out test data from bilingual speakers, with comparisons across model types, layers, and conditions. These are direct statistical measurements from independent data splits; no equations, fitted parameters renamed as predictions, self-citations, or ansatzes reduce the reported results to the inputs by construction. The analysis is self-contained against external benchmarks (held-out EMA recordings) and contains no load-bearing self-referential steps.
Assumptions & free parameters
assumptions (2)
- domain assumption Linear regression on SSL hidden states is sufficient to reveal whether articulatory information is encoded.
- domain assumption EMA sensor data collected from bilingual speakers provides comparable articulatory ground truth across Finnish and Russian.
Cite this review
Pith. "Pith review of How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA." pith.science (2026). https://pith.science/paper/A3F6GRDV
@misc{pith2026260631527,
author = {Pith},
title = {Pith review of: How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA},
year = {2026},
howpublished = {\url{https://pith.science/paper/A3F6GRDV}},
note = {Machine review of arXiv:2606.31527}
}
read the original abstract
SSL speech models capture rich phonetic, prosodic, and acoustic patterns from raw audio, yet how they encode articulatory information across diverse languages remains unclear. Using EMA data from bilingual Finnish-Russian speakers, we evaluate cross-lingual correlations between SSL latent representations and articulatory movements. Models achieve strong prediction performance (Pearson r up to 0.68) even with approximately 5 minutes of training data, with multilingual models outperforming monolingual ones. Intermediate layers encode articulatory features most effectively, and tongue movements are more predictable than lip movements. We also assess the impact of task type (read versus spontaneous speech) and language proficiency, finding higher accuracy for structured tasks and strong generalization across proficiency levels. These results enhance the interpretability of SSL models and show their potential for speech-technology applications.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Self-supervised learning (SSL) models such as wav2vec 2.0 [1], HuBERT [2], and WavLM [3] have become foundational in modern speech research. Trained on large-scale unlabeled au- dio, they learn structured representations that transfer effec- tively to diverse downstream tasks [4]. However, the precise nature of these representations remains o...
-
[2]
further showed that SSL models retain strong AAI perfor- mance in a zero-shot English-to-Dutch transfer setting, particu- larly when using intermediate encoder layers. These results support the hypothesis that SSL representa- tions capture partially language-agnostic articulatory features, although cross-lingual performance may drop when the tar- get lang...
-
[3]
FROST-EMA Dataset The articulatory dataset used is the recently collected, size- able FROST-EMA corpus [16]. It consists of parallel acous- tic and EMA recordings from 18 bilingual speakers (8 female, 10 male), covering Finnish and Russian as first and second lan- guages, along with imitated foreign-accented speech. The age range of the subjects is 21 to ...
work page Pith review arXiv 2026
-
[4]
Experimental design Speech recordings were passed through each pre-trained SSL speech model, which served as a feature extractor. For each model, we extracted hidden representations from all Trans- former layers. All encoders used in this study have 24 layers with 1024-dimensional hidden representations. This resulted in a set of layer-wise feature repres...
-
[5]
Unless otherwise noted, scores are Pearson correlations averaged over EMA channels
Results This section reports the results for the five experiments summa- rized in Table 1. Unless otherwise noted, scores are Pearson correlations averaged over EMA channels. E1: Cross-model comparison.Table 2 summarizes the av- erage articulatory scores over all speakers, sensors, and lay- ers. MMS-300m and the language fine-tuned XLS variants achieve th...
-
[6]
Across models, X/Z channels are consistently easier to pre- dict than Y , with TBZ among the strongest channels and UL Z among the weakest. E3: Training-size sensitivity.As revealed in Figure 3, artic- ulatory scores increase steeply as training duration grows from seconds to a few minutes. This indicates high sample efficiency of SSL representations for ...
-
[7]
Discussion Drawing upon the findings presented above, we have the fol- lowing four main findings. FINDING 1: Intermediate layers consistently yield the strongest articulatory alignment.Results support the view that SSL encoders capture physically meaningful production dynamics: linear probes reached moderate-to-strong correla- tions without task-specific ...
-
[8]
Conclusion Our case study on SSL model probing using EMA position data of bilingual speakers (Finnish and Russian) contributes to the growing body of evidence of the presence of articulatory dy- namics encoding by these models. Overall, the relatively strong correlations (up tor≈0.78in LOSO speaker generalization (MMS-300m), and meanr≈0.69in within-speake...
Show all 32 references
-
[9]
Acknowledgments This research has been initially conducted at the University of Eastern Finland and received funding from the BPI project THERADIA at a later stage
-
[10]
After using these tools, the authors reviewed and edited the content and take full responsibility for the publication
Generative AI Use Disclosure During the preparation of this work, the authors used DeepL and ChatGPT to assist with the editing of the English in some para- graphs. After using these tools, the authors reviewed and edited the content and take full responsibility for the publication
-
[11]
wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,”arXiv, 2020
2020
-
[12]
Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdi- nov, and A. Mohamed, “Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,”IEEE/ACM TASLP, vol. 29, pp. 3451–3460, 2021
2021
-
[13]
Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, D. Yu, S. Yu, Z. Yao, Y . Qian, X. Huang, and F. Wei, “Wavlm: Large-scale self- supervised pre-training for full stack speech processing,”IEEE JSTSP, vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[14]
SUPERB: Speech pro- cessing universal performance benchmark,
S.-w. Yang, P.-H. Chi, Y .-S. Chuanget al., “SUPERB: Speech pro- cessing universal performance benchmark,” inProc. Interspeech, 2021, pp. 1194–1198
2021
-
[15]
Predicting within and across language phoneme recognition performance of self- supervised learning speech pre-trained models,
H. Ji, T. Patel, and O. Scharenborg, “Predicting within and across language phoneme recognition performance of self- supervised learning speech pre-trained models,”arXiv preprint arXiv:2206.12489, 2022
2022
-
[16]
Prob- ing self-supervised speech models for phonetic and phone- mic information: A case study in aspiration,
K. Martin, J. Gauthier, C. Breiss, and R. Levy, “Prob- ing self-supervised speech models for phonetic and phone- mic information: A case study in aspiration,”arXiv preprint arXiv:2306.06232, 2023
2023
-
[17]
Probing for phonology in self-supervised speech representations: A case study on accent perception,
N. Venkateswaran, K. Tang, and R. Wayland, “Probing for phonology in self-supervised speech representations: A case study on accent perception,”arXiv preprint arXiv:2506.17542, 2025
2025
-
[18]
A large-scale probing analysis of speaker-specific attributes in self- supervised speech representations,
A. Y . F. Chiu, K. C. Fung, R. T. Y . Li, J. Li, and T. Lee, “A large-scale probing analysis of speaker-specific attributes in self- supervised speech representations,” 2025
2025
-
[19]
A layer-wise analysis of Man- darin and English suprasegmentals in SSL speech models,
A. de la Fuente and D. Jurafsky, “A layer-wise analysis of Man- darin and English suprasegmentals in SSL speech models,” inIn- terspeech 2024, 2024, pp. 1290–1294
2024
-
[20]
Evidence of vocal tract articulation in self-supervised learning of speech,
C. J. Cho, P.-c. Wu, A. Mohamed, and G. K. Anumanchipalli, “Evidence of vocal tract articulation in self-supervised learning of speech,” inICASSP 2023, 2023, pp. 1–5
2023
-
[21]
Improved acoustic-to- articulatory inversion using representations from pretrained self- supervised learning models,
S. Udupa, S. C, and P. K. Ghosh, “Improved acoustic-to- articulatory inversion using representations from pretrained self- supervised learning models,” inICASSP 2023, 2023, pp. 1–5
2023
-
[22]
Speaker-independent acoustic-to-articulatory speech inversion,
P.-c. Wu, L.-W. Chen, C. J. Cho, S. Watanabe, L. Goldstein, A. W. Black, and G. K. Anumanchipalli, “Speaker-independent acoustic-to-articulatory speech inversion,” inICASSP 2023, 2023, pp. 1–5
2023
-
[23]
Layer-wise analysis of a self-supervised speech representation model,
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” inASRU 2021, 2021, pp. 914–921
2021
-
[24]
Self-supervised models of speech infer universal articulatory kinematics,
C. J. Cho, A. Mohamed, A. W. Black, and G. K. Anumanchipalli, “Self-supervised models of speech infer universal articulatory kinematics,” inICASSP 2024, 2024, pp. 12 061–12 065
2024
-
[25]
Exploring self-supervised speech representa- tions for cross-lingual acoustic-to-articulatory inversion,
Y . Hao, R. Amooie, W. de Vries, T. Tienkamp, R. van Noord, and M. Wieling, “Exploring self-supervised speech representa- tions for cross-lingual acoustic-to-articulatory inversion,” inIn- terspeech 2024, 2024, pp. 4603–4607
2024
-
[26]
FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents,
S. Hopponen, T. Kinnunen, A. Nikolaev, R. Gonz´alez Hautam¨aki, L. Tavi, and E. Meister, “FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents,” inProc. Interspeech 2025, 2025, pp. 364–368
2025
-
[27]
MAIN: Multilingual Assessment Instrument for Narratives – Revised,
N. Gagarina, D. Klop, S. Kunnari, K. Tantele, T. V ¨alimaa, U. Bohnacker, and J. Walters, “MAIN: Multilingual Assessment Instrument for Narratives – Revised,”ZAS Papers in Linguistics, vol. 63, pp. 20–20, 2019
2019
-
[28]
Evaluation of speech-based digital biomark- ers: Review and recommendations,
J. Robin, J. E. Harrison, L. D. Kaufman, F. Rudzicz, W. Simpson, and M. Yancheva, “Evaluation of speech-based digital biomark- ers: Review and recommendations,”Digital Biomarkers, vol. 4, no. 3, pp. 99–108, 2020
2020
-
[29]
Audio self-supervised learning: A survey,
S. Liu, A. Mallol-Ragolta, E. Parada-Cabaleiro, K. Qian, X. Jing, A. Kathan, B. Hu, and B. W. Schuller, “Audio self-supervised learning: A survey,”Patterns, vol. 3, no. 12, p. 100616, 2022
2022
-
[30]
Understanding intermediate layers us- ing linear classifier probes,
G. Alain and Y . Bengio, “Understanding intermediate layers us- ing linear classifier probes,” in5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24- 26, 2017, Workshop Track Proceedings, 2017
2017
-
[31]
Hastie, R
T. Hastie, R. Tibshirani, and J. Friedman,The Elements of Statis- tical Learning. Springer, 2009
2009
-
[32]
Mms-1b: Massively multilingual speech mod- els,
V . Pratapet al., “Mms-1b: Massively multilingual speech mod- els,” 2024, large multilingual SSL model referenced as MMS- 300m/1B family
2024
Reviewed July 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.