REVIEW 2 major objections 6 minor 33 references
FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents
T0 review · 2 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper introduces the FROST-EMA corpus, the first electromagnetic articulography dataset covering Finnish and Russian, with 18 bilingual speakers recorded in their native language, second language, and an imitated foreign accent.
desk verdict FROST-EMA fills a real gap with a sensible bilingual L1/L2/imitated-accent design, but the paper needs a factual correction and a harder look at audio denoising before release. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the corpus's balanced within-speaker design. Each of the 18 bilingual speakers performs three speech tasks under three conditions: native language (A), native language with an imitated foreign accent (B), and second language (C), yielding nine recording types per speaker. Articulatory data come from seven sensors on the lips and tongue tracked at 1250 Hz by an electromagnetic articulograph, with head movement removed using reference sensors on the mastoids; audio is recorded at 48 kHz. This design is what lets accent effects be studied while holding speaker and task content constant.
What would settle it
Take a random sample of released sessions and recompute the head-movement correction from the raw mastoid reference sensors; if the corrected tongue trajectories retain significant correlation with the reference sensors, the normalization has failed. Alternatively, inspect spectrograms of the released audio in the 7–13.5 kHz band to see whether the denoising removed frication noise from consonants like /s/ and /ʃ/ rather than only the interference.
Extended reading notes
Core claim
The central claim is that FROST-EMA is the first publicly planned EMA corpus to include Finnish and Russian, and the first to record the same bilingual speakers in L1, L2, and imitated L2, allowing both natural and fake foreign accents to be compared within speakers. The corpus contains 18 speakers, each with up to 227 or 223 recordings depending on L1, covering three tasks (reading a passage, a word list, and picture description) under three conditions. The recordings pair 48 kHz audio with seven-sensor EMA trajectories at 1250 Hz, after head-movement normalization. The two case studies illustrate the corpus's value: the speaker verification experiment shows that imitated accents and L2 speech shift equal error rates relative to L1 in gender-dependent ways, and the articulatory analysis of one speaker reveals a diphthong-like tongue movement in L2 that is absent in L1 and imitation.
Load-bearing premise
The corpus is only as good as its recordings: if the head-movement subtraction leaves sensor artifacts, or the audio denoising removes linguistically relevant high-frequency sounds, the data will not support the phonetic and technological uses claimed.
Editorial extensions
If this is right
- Researchers can use the corpus to quantify articulatory differences between authentic L2 production and stereotyped imitations, across two typologically distinct languages.
- The paired acoustic and articulatory data can train and evaluate articulatory inversion models that estimate tongue and lip movements from audio.
- The L2 and fake-accent conditions give speaker verification developers a natural test set for measuring how voice disguise and accent variability degrade recognition accuracy.
- If released as planned, the 18-speaker dataset will support statistical, not just illustrative, analysis of co-articulation and accent effects.
Reading between the lines
- The gender-dependent ASV results (females improved in L2 while males worsened) are based on a small sample; once the corpus is public, the effect could be tested with matched proficiency and corpus balance controls.
- The imitated-accent condition may reveal whether speakers imitate a foreign accent by copying salient segmental features or by adopting global articulatory strategies; the paper does not test this directly.
- Because one speaker was kept despite partially missing positional data and another dropped for missing audio, the released corpus will need per-session quality flags; verifying the actual usable data volume would clarify how much of the advertised 227/223 recordings are complete.
- The single-speaker finding that L2 [æ] shows a backward-forward tongue movement, suggesting diphthongization or a preceding [j], could be tested across all speakers to see if it is a genuine L2 effect or a personal articulation quirk.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FROST-EMA, a new electromagnetic articulography corpus with parallel audio and articulatory data from 18 Finnish and Russian bilingual speakers, each recorded in L1, L2, and imitated-L2 conditions across three tasks (paragraph reading, word list, and picture description). The paper also reports two preliminary case studies: one using a pretrained ECAPA-TDNN speaker verification system to compare EER across the three conditions, and one illustrating articulatory and formant trajectories for a single speaker's production of [æ] across conditions. The corpus is planned for public release in the first half of 2026.
Significance. If released as described, FROST-EMA would fill a clear gap: no existing public EMA corpus covers Finnish or Russian, and the balanced L1/L2/imitated-accent design is novel and potentially valuable for both phonetic research and speech-technology applications such as articulatory inversion and speaker verification under accent variation. The corpus has more speakers than most existing EMA datasets (18 vs. 1–9 in other listed corpora), though not more than EMA-MAE's 20. The ASV case study is methodologically transparent because it uses a standard pretrained SpeechBrain model, and the articulatory case study, while single-speaker and illustrative, demonstrates the intended parallel acoustic-articulatory analysis. The main contribution is the corpus itself, and its value depends on the quality and accessibility of the released data.
major comments (2)
- [Section 2.5] The assertion that the iZotope RX 10 spectral denoising (15.3 dB reduction in the 7–13.5 kHz band) is 'not a problem for the acoustic analysis' is unsupported, and this band overlaps substantially with the spectral energy of sibilants and fricatives that are phonemically contrastive in Finnish and Russian (e.g., Finnish /s/, Russian /s, ʂ, ʐ/). Because the paper does not state that raw audio will be released alongside the processed version, the irreversible processing could compromise the acoustic half of the corpus for the very phonetic research it is intended to enable. Please provide a validation of the processing (e.g., spectral comparisons, perceptual judgments, or automatic-recognition tests), commit to releasing raw audio as well, or at minimum document the exact processing parameters and explicitly flag the affected frequency range as a limitation in the corpus description.
- [Section 1, bullet (3), and Table 1] The novelty claim of 'an overall larger number of subjects compared to existing corpora' is contradicted by Table 1, which lists EMA-MAE as having 20 speakers (10 L1 English + 10 L1 Mandarin) versus the 18 speakers in FROST-EMA. This is a factual error in a headline claim; please revise the wording to something like 'larger than most existing EMA corpora' or 'one of the largest', or adjust the comparison in Table 1.
minor comments (6)
- [Section 1] There is a typo in the sentence 'makes the compilation a corpus a tedious task'; it should read 'makes the compilation of such a corpus a tedious task.'
- [Section 3.2] The phrase 'the use of a imitated accent' should be 'the use of an imitated accent.'
- [Section 4, Figures 3, 5, and 7] The text describes Figures 3, 5, and 7 as showing 'articulatory trajectories' and 'overlapping gestures,' but these figures display F1 and F2 formant traces over time; please clarify the terminology or relabel the figures to avoid confusion.
- [Table 1] The EMA-MAE row is visually confusing because it lists two speaker groups (L1 English and L1 Mandarin) in a single row; consider splitting it into two rows or adding a clearer subheader.
- [Section 2.3] The paper does not specify the EMA data format (e.g., file extension, coordinate system, units) or the synchronization procedure between the audio and EMA streams; these details are important for prospective users and should be added or referenced.
- [References] The reference for Rebernik et al. contains a typo: 'electromagnetic articulogragphy' should be 'electromagnetic articulography.'
Circularity Check
No significant circularity: FROST-EMA is a corpus introduction with self-contained case studies, not a derivation that reduces to its inputs.
full rationale
This paper introduces a new EMA corpus and presents two illustrative case studies. There is no derivation chain, no fitted parameter renamed as a prediction, and no equation whose output is an input by construction. The central claim is that the corpus exists and is planned for release; this is supported by the recording procedures, equipment descriptions, speaker table, and data counts. The speaker verification case study uses an independently pretrained ECAPA-TDNN model from SpeechBrain trained on VoxCeleb, which is external to this paper and not fitted to the FROST-EMA data. The articulatory case study is a direct measurement of one speaker's EMA trajectories and formants, with no model-based inference that could be circular. The only self-citation is the prior pilot study [7], used for novelty context and to state that no other EMA database provides a two-directional accent design; that citation is not load-bearing for the validity of the new corpus or its results. The iZotope spectral denoising and possible high-frequency energy loss described in Section 2.5 is a data-quality risk and an external validity concern, not a circularity: the audio is not being used to justify its own denoising, and the risk does not make the corpus claim equivalent to its inputs. No circular step meets the evidentiary standard of quoting a specific reduction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Head-movement normalization in the AG501 leaves only articulatory movements.
- domain assumption Spectral denoising with iZotope RX 10 does not alter acoustically relevant phonetic information.
- domain assumption Self-reported L2 proficiency levels in Table 2 are accurate reflections of speaker ability.
- domain assumption ECAPA-TDNN pretrained on VoxCeleb transfers to Finnish and Russian data.
Cite this review
Pith. "Pith review of FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents." pith.science (2026). https://pith.science/paper/MOR44XVF
@misc{pith2026250608981,
author = {Pith},
title = {Pith review of: FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents},
year = {2026},
howpublished = {\url{https://pith.science/paper/MOR44XVF}},
note = {Machine review of arXiv:2506.08981}
}
read the original abstract
We introduce a new FROST-EMA (Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography) corpus. It consists of 18 bilingual speakers, who produced speech in their native language (L1), second language (L2), and imitated L2 (fake foreign accent). The new corpus enables research into language variability from phonetic and technological points of view. Accordingly, we include two preliminary case studies to demonstrate both perspectives. The first case study explores the impact of L2 and imitated L2 on the performance of an automatic speaker verification system, while the second illustrates the articulatory patterns of one speaker in L1, L2, and a fake accent.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Introduction Humans produce speech by coordinating the movements of var- iousarticulators, such as the tongue and the lips. Humans con- tinuously reshape their vocal tract to modulate the airstream produced by the voice source, the larynx. The resulting acous- tic waveform, commonly captured with a microphone, contains latent (i.e. not directly observable...
-
[2]
The collection of FROST-EMA dataset 2.1. Phonetic basis for the collected data Finnish and Russian, though geographically adjacent, belong to different language families—Finno-Ugric and Indo-European [11], respectively—and exhibit distinct phonetic and phonolog- ical characteristics (for Finnish, see [12] and for Russian, see [13]). Finnish features a lar...
work page 2025
-
[3]
Case study I: Speaker verification 3.1. Setup We carried out an experiment to investigate the effects of con- ditions A, B, and C on automatic speaker verification (ASV). We used 1A as speaker enrollment. For testing, we used paired speech samples from stimuli section 2 with all variants (A, B Table 2:Participants in the data collection Participant ID Sex...
-
[4]
We extracted formant measurements with [26]
Case Study II: Articulation of [æ] kh5 (a female participant with L1 Finnish and L2 Russian) pro- vides an illustrative example of the insights that EMA data can offer. We extracted formant measurements with [26]. In the figures below, co-articulation effects caused by the surrounding consonants can be seen in the movements of the three tongue sensors (Fi...
-
[5]
Discussion Our new FROST-EMA corpus with an innovative design intro- duced in this work is designed to serve purposes of both speech science and speech technology for bilingual analysis. Follow- ing metadata clean-up and further auditing of the individual sen- tences, we anticipate the release of the database in the first half of 2026. -8 -4 0 -15 -10 -5 ...
work page 2026
-
[6]
Acknowledgements The authors wish to thank the School of Computing of the Uni- versity of Eastern Finland and the Logolab research environ- ment located at the School of Humanities, University of Eastern Finland, for the use of its facilities. In addition, the authors wish to thank senior laboratory technician Petri Pulli and research assistants Martta Si...
-
[7]
Articulation during voice disguise: a pilot study,
L. Tavi, T. Kinnunen, E. Meister, R. Gonz ´alez-Hautam¨aki, and A. Malmi, “Articulation during voice disguise: a pilot study,” Proc. Speech and Computer (SPECOM´21), Springer LNAI 12997, pp. 680–691, 2021-09
work page 2021
-
[8]
Fricative consonants: acoustic and x-ray measure- ments,
P. Badin, “Fricative consonants: acoustic and x-ray measure- ments,”Journal of Phonetics, vol. 19, no. 3, pp. 397–408
Show all 33 references
-
[9]
Another interesting study [10] used linear probingto study the association of EMA coordinate val- ues with the latent features extracted from various foundational speech models
utilized EMA data as part in a brain-to-speech synthesis to address data limitations. Another interesting study [10] used linear probingto study the association of EMA coordinate val- ues with the latent features extracted from various foundational speech models. Their results...
2025 arXiv
-
[10]
Articulatory phonetics in the market: combining public engagement with ultrasound data collection,
C. Nance, M. Dewhurst, L. Fairclough, P. Forster, S. Kirkham, J. J. H. Lo, J. McMonagle, T. Nagamine, S. Parkman, H. Rabani, A. Siem, D. Turton, and D. Wang, “Articulatory phonetics in the market: combining public engagement with ultrasound data collection,”Linguistics Vanguar...
2024 doi
-
[11]
Real-time magnetic resonance imaging and electromagnetic articulography database for speech production research (tc),
S. S. Narayanan, A. Toutios, V . Ramanarayanan, A. C. Lammert, J. Kim, S. Lee, K. S. Nayak, Y .-C. Kim, Y . Zhu, L. M. Goldstein, D. Byrd, E. Bresch, P. K. Ghosh, A. Katsamanis, and M. I. Proctor, “Real-time magnetic resonance imaging and electromagnetic articulography databas...
2014 doi
-
[12]
The impact of l1 negative phonological transfer on l2 word identification and production,
J. Fan and L. Yongbing, “The impact of l1 negative phonological transfer on l2 word identification and production,”International Journal of Linguistics, vol. 6, no. 5, pp. 37–50, 2014
2014
-
[13]
Acoustic-to-articulatory mapping with joint optimization of deep speech enhancement and articulatory inversion models,
A. S. Shahrebabaki, G. Salvi, T. Svendsen, and S. M. Sinis- calchi, “Acoustic-to-articulatory mapping with joint optimization of deep speech enhancement and articulatory inversion models,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, vol. 30, pp. 135–147, 2022
2022
-
[14]
Modeling consonant-vowel coarticulation for articulatory speech synthesis,
P. Birkholz, “Modeling consonant-vowel coarticulation for articulatory speech synthesis,”PLOS ONE, vol. 8, no. 4, pp. 1–17, 2013-04, publisher: Public Library of Science. [Online]. Available: https://doi.org/10.1371/journal.pone.0060603
2013 doi
-
[15]
V oice disguise using a foreign accent: phonetic and linguistic variation
S. Neuhauser, “V oice disguise using a foreign accent: phonetic and linguistic variation.”International Journal of Speech, Lan- guage & the Law, vol. 15, no. 2, pp. 131–159, 2008
2008
-
[16]
Speech synthesis from neural decoding of spoken sentences,
G. Anumanchipalli, J. Chartier, and E. Chang, “Speech synthesis from neural decoding of spoken sentences,”Nature, no. 568, p. 493–498, 2019
2019
-
[17]
Ev- idence of vocal tract articulation in self-supervised learning of speech,
C. J. Cho, P. Wu, A. Mohamed, and G. K. Anumanchipalli, “Ev- idence of vocal tract articulation in self-supervised learning of speech,” inICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
2023
-
[18]
Connections between uralic and other language fami- lies,
S. Georg, “Connections between uralic and other language fami- lies,” inThe Uralic languages, 2nd ed., ser. Routledge Language Family Series, D. Abondolo and R.-L. Valij¨arvi, Eds. Routledge, 2023, pp. 176–209
2023
-
[19]
The finnic languages,
R. Gr ¨unthal, “The finnic languages,” inThe Uralic languages, 2nd ed., ser. Routledge Language Family Series, D. Abondolo and R.-L. Valij¨arvi, Eds. Routledge, 2023, pp. 433–480
2023
-
[20]
Sussex and P
R. Sussex and P. V . Cubberley,The Slavic languages, ser. Cam- bridge language surveys. Cambridge, UK: Cambridge University Press, 2006
2006
-
[21]
The MOCHA-TIMIT articulatory database,
A. Wrench, “The MOCHA-TIMIT articulatory database,” https: //www.cstr.ed.ac.uk/research/projects/artic/mocha.html, 1999
1999
-
[22]
Announcing the electro- magnetic articulography (day 1) subset of the mngu0 articulatory corpus,
K. Richmond, P. Hoole, and S. King, “Announcing the electro- magnetic articulography (day 1) subset of the mngu0 articulatory corpus,” inInterspeech 2011, 2011, pp. 1505–1508
2011
-
[23]
The electromagnetic ar- ticulography mandarin accented english (EMA-MAE) corpus of acoustic and 3d articulatory kinematic data,
A. Ji, J. J. Berry, and M. T. Johnson, “The electromagnetic ar- ticulography mandarin accented english (EMA-MAE) corpus of acoustic and 3d articulatory kinematic data,” in2014 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 7719–7723
2014
-
[24]
A new italian dataset of parallel acoustic and articulatory data,
C. Canevari, L. Badino, and L. Fadiga, “A new italian dataset of parallel acoustic and articulatory data,” inInterspeech 2015, 2015, pp. 2152–2156, ISSN: 2958-1796
2015
-
[25]
The DKU- JNU-EMA electromagnetic articulography database on mandarin and chinese dialects with tandem feature based acoustic-to- articulatory inversion,
Z. Cai, X. Qin, D. Cai, M. Li, X. Liu, and H. Zhong, “The DKU- JNU-EMA electromagnetic articulography database on mandarin and chinese dialects with tandem feature based acoustic-to- articulatory inversion,” in2018 11th International Symposium on Chinese Spoken Language Proces...
2018
-
[26]
Main: Multilingual assessment instrument for narratives–revised,
N. Gagarina, D. Klop, S. Kunnari, K. Tantele, T. V ¨alimaa, U. Bohnacker, and J. Walters, “Main: Multilingual assessment instrument for narratives–revised,”ZAS Papers in Linguistics, vol. 63, pp. 20–20, 2019
2019
-
[27]
VisArtico: a visual- ization tool for articulatory data,
S. Ouni, L. Mangeonjean, and I. Steiner, “VisArtico: a visual- ization tool for articulatory data,”13th Annual Conference of the International Speech Communication Association - InterSpeech 2012, 2012-09, portland, OR, USA
2012
-
[28]
An optimal set of flesh points on tongue and lips for speech-movement classi- fication,
J. Wang, A. Samal, P. Rong, and J. R. Green, “An optimal set of flesh points on tongue and lips for speech-movement classi- fication,”Journal of Speech, Language, and Hearing Research, vol. 59, pp. 15–26, 2016-02
2016
-
[29]
A review of data collection practices using electromagnetic articu- logragphy,
T. Rebernik, J. Jacobi, R. Jonkers, A. Noiray, and M. Wieling, “A review of data collection practices using electromagnetic articu- logragphy,”Laboratory phonology, vol. 12(1), no. 6, pp. 1–42, 2021
2021
-
[30]
Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” inInterspeech 2020, 2020, pp. 3830–3834
2020
-
[31]
SpeechBrain: A general-purpose speech toolkit,
M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, J.-C. Chou, S.-L. Yeh, S.-W. Fu, C.-F. Liao, E. Rastorgueva, F. Grondin, W. Aris, H. Na, Y . Gao, R. D. Mori, and Y . Ben- gio, “SpeechBrain: A general-pu...
2021 arXiv
-
[32]
V oxCeleb: Large-scale speaker verification in the wild,
A. Nagraniet al., “V oxCeleb: Large-scale speaker verification in the wild,”Computer Speech & Language, vol. 60, p. 101027, 2020
2020
-
[33]
Praat: doing phonetics by computer [computer program],
P. Boersma and D. Weenink, “Praat: doing phonetics by computer [computer program],” 2024-10-27. [Online]. Available: http://www.praat.org
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.