Pith. sign in

REVIEW 3 major objections 1 minor 32 references

How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

T0 review · 3 major / 1 minor · reviewed 2026-07-01 · grok-4.3

Pith's one-line read SSL speech models predict articulatory movements from bilingual Finnish-Russian speakers with Pearson r up to 0.68 using only about 5 minutes of training data.

desk verdict Bilingual EMA data lets them compare SSL models on articulatory prediction across Finnish and Russian with the same speakers, but the abstract leaves the controls for speaker and proficiency effects unclear. read the letter →

arxiv 2606.31527 v1 pith:A3F6GRDV submitted 2026-06-30 eess.AS cs.SD

classification eess.AScs.SD
keywords SSLspeechmodelsarticulatoryencodingEMAcross-lingualprobingbilingualspeakersFinnishRussianintermediatelayers
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests whether self-supervised learning speech models encode articulatory information in a way that transfers across languages. It does this by probing model layers against electromagnetic articulography recordings from the same bilingual speakers producing Finnish and Russian. Multilingual models outperform monolingual ones, intermediate layers show the strongest correlations, and tongue movements prove easier to predict than lip movements. Performance holds across read and spontaneous speech tasks and across different proficiency levels. The results indicate that these models capture articulatory patterns that are not tied to a single language.

What carries the argument

Cross-lingual probing that correlates SSL latent representations with EMA-tracked articulatory movements from the same bilingual speakers.

What would settle it

Finding that prediction correlations fall to near zero when the same model layers are tested on held-out bilingual speakers or when language-specific biases in the EMA data are removed.

Watch

Extended reading notes

Core claim

Using EMA data from bilingual Finnish-Russian speakers, the study measures how well SSL model representations predict articulatory trajectories. Models reach Pearson correlations of 0.68 with roughly 5 minutes of training data per speaker. Multilingual pretraining yields higher accuracy than monolingual pretraining. Intermediate layers align best with the articulatory signals, tongue sensors correlate more strongly than lip sensors, read speech outperforms spontaneous speech, and results generalize across speaker proficiency levels.

Load-bearing premise

The articulatory movements recorded by EMA from bilingual speakers form a language-independent ground truth that can be compared directly between Finnish and Russian.

Editorial extensions

If this is right

  • Multilingual SSL models capture articulatory features more effectively than monolingual ones across these two languages.
  • Intermediate layers of the models are the most useful for recovering articulatory information.
  • Tongue movements can be recovered from model representations more reliably than lip movements.
  • Performance remains high for both read and spontaneous speech and across proficiency levels.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the same pattern holds for other language pairs, SSL models could support articulatory-based tools in low-resource languages without new EMA collections.
  • The layer-wise pattern suggests that articulatory supervision could be added most efficiently at intermediate depths during fine-tuning.
  • Higher accuracy on read speech implies that controlled elicitation tasks may be sufficient to benchmark articulatory encoding.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The paper claims that self-supervised learning (SSL) speech models encode articulatory information in a cross-lingual manner, evaluated via EMA data from bilingual Finnish-Russian speakers. Multilingual models outperform monolingual ones with Pearson r up to 0.68 even using ~5 minutes of training data; intermediate layers are most effective, tongue movements are more predictable than lip movements, and performance is higher for read speech with generalization across proficiency levels.

Significance. If the cross-lingual comparisons are robust to speaker-specific effects, the results would strengthen the case that multilingual SSL models capture more language-independent articulatory features than monolingual counterparts. The small-data probing efficiency and layer-wise findings could aid interpretability efforts and guide applications in speech technology that rely on articulatory grounding.

major comments (3)
  1. [Abstract] Abstract: The abstract states performance numbers (Pearson r up to 0.68) and qualitative trends but supplies no information on data partitioning, statistical testing, baseline comparisons, or controls for speaker identity, so it is impossible to judge whether the reported correlations genuinely support the cross-lingual claim.
  2. [Cross-lingual evaluation setup] Cross-lingual evaluation setup: The central claim rests on treating EMA recordings from bilingual speakers as directly comparable, language-independent targets. Without explicit quantification of residual speaker variance or a monolingual-speaker control condition, speaker-idiosyncratic or L1-dominant signatures could inflate the reported multilingual advantage and layer-wise patterns.
  3. [Results on task and proficiency] Results on task and proficiency: The abstract states that proficiency and task effects were assessed, yet the absence of details on how residual variance was isolated or quantified leaves open whether these factors confound the multilingual vs. monolingual comparison.
minor comments (1)
  1. [Methods] The methods section would benefit from explicit description of the exact SSL models, probe architecture, and normalization procedures to support reproducibility.

Simulated Author's Rebuttal

3 responses · 0 unresolved

Thank you for the opportunity to respond to the referee's report. We address each of the major comments point by point below, indicating where revisions will be made to the manuscript.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The abstract states performance numbers (Pearson r up to 0.68) and qualitative trends but supplies no information on data partitioning, statistical testing, baseline comparisons, or controls for speaker identity, so it is impossible to judge whether the reported correlations genuinely support the cross-lingual claim.

    Authors: We agree with this observation. The revised abstract will incorporate details on the data partitioning (speaker-independent cross-validation), statistical significance testing of the Pearson correlations, baseline comparisons (e.g., against MFCC features), and the use of bilingual speakers as a control for speaker identity. These additions will strengthen the abstract's support for the cross-lingual claims. revision: yes

  2. Referee: [Cross-lingual evaluation setup] Cross-lingual evaluation setup: The central claim rests on treating EMA recordings from bilingual speakers as directly comparable, language-independent targets. Without explicit quantification of residual speaker variance or a monolingual-speaker control condition, speaker-idiosyncratic or L1-dominant signatures could inflate the reported multilingual advantage and layer-wise patterns.

    Authors: Our experimental design uses bilingual speakers to ensure that language comparisons are within the same speaker, thereby controlling for speaker-specific effects. We report per-speaker results and inter-speaker variance in the paper to quantify residual speaker variance. A monolingual control would necessitate new recordings and is beyond the scope of this work; however, we will expand the discussion to explicitly address potential L1 influences and why the bilingual setup mitigates speaker confounds. revision: partial

  3. Referee: [Results on task and proficiency] Results on task and proficiency: The abstract states that proficiency and task effects were assessed, yet the absence of details on how residual variance was isolated or quantified leaves open whether these factors confound the multilingual vs. monolingual comparison.

    Authors: We will add detailed explanations in the methods and results sections on the statistical approaches used to isolate task and proficiency effects, including mixed-effects modeling to quantify their contribution to variance while assessing the multilingual advantage. This will demonstrate that these factors were properly controlled and do not confound the main findings. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: purely empirical correlations on held-out data

full rationale

The paper reports measured Pearson r values (up to 0.68) between SSL latent features and EMA targets on held-out test data from bilingual speakers, with comparisons across model types, layers, and conditions. These are direct statistical measurements from independent data splits; no equations, fitted parameters renamed as predictions, self-citations, or ansatzes reduce the reported results to the inputs by construction. The analysis is self-contained against external benchmarks (held-out EMA recordings) and contains no load-bearing self-referential steps.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on standard assumptions of linear probing and the validity of EMA as a proxy for articulation; no free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption Linear regression on SSL hidden states is sufficient to reveal whether articulatory information is encoded.
    The evaluation uses Pearson correlation after (implicit) linear mapping; this is a standard but unproven assumption in probing studies.
  • domain assumption EMA sensor data collected from bilingual speakers provides comparable articulatory ground truth across Finnish and Russian.
    Cross-lingual comparison assumes the measurement modality does not introduce language-specific artifacts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA." pith.science (2026). https://pith.science/paper/A3F6GRDV

@misc{pith2026260631527,
  author       = {Pith},
  title        = {Pith review of: How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A3F6GRDV}},
  note         = {Machine review of arXiv:2606.31527}
}
read the original abstract

SSL speech models capture rich phonetic, prosodic, and acoustic patterns from raw audio, yet how they encode articulatory information across diverse languages remains unclear. Using EMA data from bilingual Finnish-Russian speakers, we evaluate cross-lingual correlations between SSL latent representations and articulatory movements. Models achieve strong prediction performance (Pearson r up to 0.68) even with approximately 5 minutes of training data, with multilingual models outperforming monolingual ones. Intermediate layers encode articulatory features most effectively, and tongue movements are more predictable than lip movements. We also assess the impact of task type (read versus spontaneous speech) and language proficiency, finding higher accuracy for structured tasks and strong generalization across proficiency levels. These results enhance the interpretability of SSL models and show their potential for speech-technology applications.

Figures

Figures reproduced from arXiv: 2606.31527 by the authors.

Figure 1
Figure 1. Layer-wise articulatory prediction performance (average Pearson correlation r) across encoder depth for (a) XLSR-53, (b) wav2vec 2.0, and (c) MMS-300m. Curves are vertically offset by sensor for visual clarity. While distinct correlation trajectories are observed across spatial axes, the overall layer-wise profiles remain highly similar across articulators and coordinate axes, suggesting consistent probing dynamics … view at source ↗
Figure 2
Figure 2. Layer-wise articulatory prediction performance (Pearson correlation r) for MMS-300m under a LOSO evaluation. (a) Task comparison across Finnish and Russian speakers, showing higher correlations for controlled reading tasks (T1-T2) than for spontaneous picture description (T3). (b) Language-condition comparison (L1, L2, and simulated accent conditions.) 20 30 60 300 600 1200 Training Size (s) 0.1 0.2 0.3 0.4 0.5 0.6 … view at source ↗
Figure 3
Figure 3. Performance increases with training size and stabi￾lizes around 300 seconds. The solid green lines correspond to Finnish speakers, while the dashed purple lines indicate Rus￾sian speakers. pronounced for vertical upper-lip movement (UL Z), which re￾mained hardest to predict in our experiments. FINDING 3: Tongue articulators are more reliably encoded than lips. Tongue trajectories (e.g., TB X, TT Z) were con￾sistentl… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 32 canonical work pages

  1. [1]

    Trained on large-scale unlabeled au- dio, they learn structured representations that transfer effec- tively to diverse downstream tasks [4]

    Introduction Self-supervised learning (SSL) models such as wav2vec 2.0 [1], HuBERT [2], and WavLM [3] have become foundational in modern speech research. Trained on large-scale unlabeled au- dio, they learn structured representations that transfer effec- tively to diverse downstream tasks [4]. However, the precise nature of these representations remains o...

  2. [2]

    further showed that SSL models retain strong AAI perfor- mance in a zero-shot English-to-Dutch transfer setting, particu- larly when using intermediate encoder layers. These results support the hypothesis that SSL representa- tions capture partially language-agnostic articulatory features, although cross-lingual performance may drop when the tar- get lang...

  3. [3]

    How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

    FROST-EMA Dataset The articulatory dataset used is the recently collected, size- able FROST-EMA corpus [16]. It consists of parallel acous- tic and EMA recordings from 18 bilingual speakers (8 female, 10 male), covering Finnish and Russian as first and second lan- guages, along with imitated foreign-accented speech. The age range of the subjects is 21 to ...

  4. [4]

    familiarity

    Experimental design Speech recordings were passed through each pre-trained SSL speech model, which served as a feature extractor. For each model, we extracted hidden representations from all Trans- former layers. All encoders used in this study have 24 layers with 1024-dimensional hidden representations. This resulted in a set of layer-wise feature repres...

  5. [5]

    Unless otherwise noted, scores are Pearson correlations averaged over EMA channels

    Results This section reports the results for the five experiments summa- rized in Table 1. Unless otherwise noted, scores are Pearson correlations averaged over EMA channels. E1: Cross-model comparison.Table 2 summarizes the av- erage articulatory scores over all speakers, sensors, and lay- ers. MMS-300m and the language fine-tuned XLS variants achieve th...

  6. [6]

    E3: Training-size sensitivity.As revealed in Figure 3, artic- ulatory scores increase steeply as training duration grows from seconds to a few minutes

    Across models, X/Z channels are consistently easier to pre- dict than Y , with TBZ among the strongest channels and UL Z among the weakest. E3: Training-size sensitivity.As revealed in Figure 3, artic- ulatory scores increase steeply as training duration grows from seconds to a few minutes. This indicates high sample efficiency of SSL representations for ...

  7. [7]

    Discussion Drawing upon the findings presented above, we have the fol- lowing four main findings. FINDING 1: Intermediate layers consistently yield the strongest articulatory alignment.Results support the view that SSL encoders capture physically meaningful production dynamics: linear probes reached moderate-to-strong correla- tions without task-specific ...

  8. [8]

    Conclusion Our case study on SSL model probing using EMA position data of bilingual speakers (Finnish and Russian) contributes to the growing body of evidence of the presence of articulatory dy- namics encoding by these models. Overall, the relatively strong correlations (up tor≈0.78in LOSO speaker generalization (MMS-300m), and meanr≈0.69in within-speake...

Show all 32 references
  1. [9]

    Acknowledgments This research has been initially conducted at the University of Eastern Finland and received funding from the BPI project THERADIA at a later stage

  2. [10]

    After using these tools, the authors reviewed and edited the content and take full responsibility for the publication

    Generative AI Use Disclosure During the preparation of this work, the authors used DeepL and ChatGPT to assist with the editing of the English in some para- graphs. After using these tools, the authors reviewed and edited the content and take full responsibility for the publication

  3. [11]

    wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,”arXiv, 2020

  4. [12]

    Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdi- nov, and A. Mohamed, “Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,”IEEE/ACM TASLP, vol. 29, pp. 3451–3460, 2021

  5. [13]

    Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, D. Yu, S. Yu, Z. Yao, Y . Qian, X. Huang, and F. Wei, “Wavlm: Large-scale self- supervised pre-training for full stack speech processing,”IEEE JSTSP, vol. 16, no. 6, pp. 1505–1518, 2022

  6. [14]

    SUPERB: Speech pro- cessing universal performance benchmark,

    S.-w. Yang, P.-H. Chi, Y .-S. Chuanget al., “SUPERB: Speech pro- cessing universal performance benchmark,” inProc. Interspeech, 2021, pp. 1194–1198

  7. [15]

    Predicting within and across language phoneme recognition performance of self- supervised learning speech pre-trained models,

    H. Ji, T. Patel, and O. Scharenborg, “Predicting within and across language phoneme recognition performance of self- supervised learning speech pre-trained models,”arXiv preprint arXiv:2206.12489, 2022

  8. [16]

    Prob- ing self-supervised speech models for phonetic and phone- mic information: A case study in aspiration,

    K. Martin, J. Gauthier, C. Breiss, and R. Levy, “Prob- ing self-supervised speech models for phonetic and phone- mic information: A case study in aspiration,”arXiv preprint arXiv:2306.06232, 2023

  9. [17]

    Probing for phonology in self-supervised speech representations: A case study on accent perception,

    N. Venkateswaran, K. Tang, and R. Wayland, “Probing for phonology in self-supervised speech representations: A case study on accent perception,”arXiv preprint arXiv:2506.17542, 2025

  10. [18]

    A large-scale probing analysis of speaker-specific attributes in self- supervised speech representations,

    A. Y . F. Chiu, K. C. Fung, R. T. Y . Li, J. Li, and T. Lee, “A large-scale probing analysis of speaker-specific attributes in self- supervised speech representations,” 2025

  11. [19]

    A layer-wise analysis of Man- darin and English suprasegmentals in SSL speech models,

    A. de la Fuente and D. Jurafsky, “A layer-wise analysis of Man- darin and English suprasegmentals in SSL speech models,” inIn- terspeech 2024, 2024, pp. 1290–1294

  12. [20]

    Evidence of vocal tract articulation in self-supervised learning of speech,

    C. J. Cho, P.-c. Wu, A. Mohamed, and G. K. Anumanchipalli, “Evidence of vocal tract articulation in self-supervised learning of speech,” inICASSP 2023, 2023, pp. 1–5

  13. [21]

    Improved acoustic-to- articulatory inversion using representations from pretrained self- supervised learning models,

    S. Udupa, S. C, and P. K. Ghosh, “Improved acoustic-to- articulatory inversion using representations from pretrained self- supervised learning models,” inICASSP 2023, 2023, pp. 1–5

  14. [22]

    Speaker-independent acoustic-to-articulatory speech inversion,

    P.-c. Wu, L.-W. Chen, C. J. Cho, S. Watanabe, L. Goldstein, A. W. Black, and G. K. Anumanchipalli, “Speaker-independent acoustic-to-articulatory speech inversion,” inICASSP 2023, 2023, pp. 1–5

  15. [23]

    Layer-wise analysis of a self-supervised speech representation model,

    A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” inASRU 2021, 2021, pp. 914–921

  16. [24]

    Self-supervised models of speech infer universal articulatory kinematics,

    C. J. Cho, A. Mohamed, A. W. Black, and G. K. Anumanchipalli, “Self-supervised models of speech infer universal articulatory kinematics,” inICASSP 2024, 2024, pp. 12 061–12 065

  17. [25]

    Exploring self-supervised speech representa- tions for cross-lingual acoustic-to-articulatory inversion,

    Y . Hao, R. Amooie, W. de Vries, T. Tienkamp, R. van Noord, and M. Wieling, “Exploring self-supervised speech representa- tions for cross-lingual acoustic-to-articulatory inversion,” inIn- terspeech 2024, 2024, pp. 4603–4607

  18. [26]

    FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents,

    S. Hopponen, T. Kinnunen, A. Nikolaev, R. Gonz´alez Hautam¨aki, L. Tavi, and E. Meister, “FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents,” inProc. Interspeech 2025, 2025, pp. 364–368

  19. [27]

    MAIN: Multilingual Assessment Instrument for Narratives – Revised,

    N. Gagarina, D. Klop, S. Kunnari, K. Tantele, T. V ¨alimaa, U. Bohnacker, and J. Walters, “MAIN: Multilingual Assessment Instrument for Narratives – Revised,”ZAS Papers in Linguistics, vol. 63, pp. 20–20, 2019

  20. [28]

    Evaluation of speech-based digital biomark- ers: Review and recommendations,

    J. Robin, J. E. Harrison, L. D. Kaufman, F. Rudzicz, W. Simpson, and M. Yancheva, “Evaluation of speech-based digital biomark- ers: Review and recommendations,”Digital Biomarkers, vol. 4, no. 3, pp. 99–108, 2020

  21. [29]

    Audio self-supervised learning: A survey,

    S. Liu, A. Mallol-Ragolta, E. Parada-Cabaleiro, K. Qian, X. Jing, A. Kathan, B. Hu, and B. W. Schuller, “Audio self-supervised learning: A survey,”Patterns, vol. 3, no. 12, p. 100616, 2022

  22. [30]

    Understanding intermediate layers us- ing linear classifier probes,

    G. Alain and Y . Bengio, “Understanding intermediate layers us- ing linear classifier probes,” in5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24- 26, 2017, Workshop Track Proceedings, 2017

  23. [31]

    Hastie, R

    T. Hastie, R. Tibshirani, and J. Friedman,The Elements of Statis- tical Learning. Springer, 2009

  24. [32]

    Mms-1b: Massively multilingual speech mod- els,

    V . Pratapet al., “Mms-1b: Massively multilingual speech mod- els,” 2024, large multilingual SSL model referenced as MMS- 300m/1B family

Pith tools

Reviewed July 1, 2026 · model on record in the stance chip above.