{"id":"66975ec9-31e5-4cc9-85a6-15510cdfa662","arxiv_id":"2607.05060","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An SI system trained only on American English estimates oral tract variables and source features on unseen French and Russian with PPMC 0.83 and 0.74, and also tracks velopharyngeal opening via nasalance.","lead":"An English-trained speech inversion model recovers tongue/lip tract variables and source features from French and Russian audio with average correlations of 0.83 and 0.74. This matters because articulatory data are hard to collect, so a language-agnostic acoustic-to-articulatory mapper would open multi-lingual and clinical analysis without new sensors.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Cross-lingual PPMC rests on tiny held-out N and English-derived TV geometry without variance or speaker-level reporting.","rationale":"The reader already isolates the exact soft spot: English-defined TV geometry + tiny non-English N. That is the load-bearing assumption; everything else (WavLM English bias, modest gain over [8], private data) is secondary. The paper’s own numbers still show useful transfer on French oral TVs, so the verdict stays CONDITIONAL rather than REJECT—the French result is credible partial evidence—but the stronger “language-agnostic” title and the Russian/VP claims remain under-powered until speaker-level variance is shown. No independent formal verification or public artifacts exist to offset the sample-size risk. Agreement with the reader is therefore full; no verdict change is required.","tokens_in":8843,"tokens_out":534,"duration_ms":6700,"concrete_test":"Recompute Table 3 and Table 4 as leave-one-speaker-out means with 95% bootstrap CIs (or report the full per-speaker PPMC matrix) on the existing YU French/Russian files; if any speaker’s oral-TV average falls below ~0.65 or the Russian CI includes values <0.60, the cross-lingual claim is not supported by the current data.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Table 3 averages 0.83 French / 0.74 Russian; Table 4 VP 0.89/0.82) treats those PPMC numbers as evidence of language-agnostic recovery of oral TVs + SFs + VP. That inference requires that (i) the geometric flesh-point-to-TV transform defined on English XRMB/YU (Sec. 2.1.1–2.1.2, citing [21]) remains a faithful, language-independent representation for French and Russian articulatory targets, and (ii) the 4 French + 3 Russian speakers (and only 1 Russian for nasalance) are representative enough that a single pooled correlation is stable. Neither is secured: no per-speaker or bootstrap intervals appear, Russian is already flagged as degraded by one noisy recording (Sec. 3.1), and VP training for XRMB uses pseudo-labels from the authors’ own prior nasal SI system [8] (Sec. 3.3). If either the TV geometry or the small cohorts are non-representative, the language-agnostic framing does not follow from the reported averages.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper presents a multi-task speech inversion (SI) system that maps audio to six oral tract variables (LA, LP, TBCL, TBCD, TTCL, TTCD) plus three source features (Per, Aper, F0), and optionally a velopharyngeal (VP) TV, using WavLM-Large embeddings, Conformer layers, and a combined Pearson+RMSE loss. The system is trained only on American English (XRMB + YU English splits) and evaluated on held-out English as well as previously unseen French and Russian speakers from a newly collected multi-lingual EMA/nasalance corpus. Table 3 reports average PPMC of 0.86 (XRMB English), 0.85 (YU English), 0.83 (French), and 0.74 (Russian) across the nine oral+source parameters; Table 4 reports VP-TV vs. nasalance PPMC of 0.92 (English), 0.89 (French), and 0.82 (Russian). Qualitative trajectory plots and a comparison to a prior SI model on XRMB are also provided.","tokens_in":9091,"tokens_out":1223,"duration_ms":9127,"significance":"If the cross-lingual numbers hold under more rigorous scrutiny, the work supplies a practical English-trained SI pipeline that recovers both oral constriction timing and source/VP information for French and Russian without language-specific articulatory training data. That would be useful for multi-lingual articulatory research, clinical assessment of nasality, and low-resource settings where EMA is unavailable. Strengths include speaker-independent splits, simultaneous multi-parameter estimation, a modest improvement over the prior SI baseline on XRMB (0.86 vs 0.85), and the collection of a multi-lingual EMA+nasalance corpus. The language-agnostic framing is therefore of genuine interest to the speech-inversion and articulatory-phonetics communities, provided the small held-out cohorts and English-derived TV geometry are shown to be representative.","major_comments":[{"comment":"Table 3 and Sec. 3.1: the central cross-lingual claim rests on pooled PPMC averages over only 4 French and 3 Russian speakers (and, for VP in Table 4 / Sec. 3.3, only 3 French and 1 Russian). No per-speaker scores, standard errors, or bootstrap intervals are reported. The text itself attributes the Russian drop (especially Per/F0) to one noisy recording. Without speaker-level or uncertainty statistics, it is not possible to judge whether 0.83 / 0.74 (and 0.89 / 0.82 for VP) are stable estimates of language-agnostic performance or are driven by individual speakers/recording conditions.","section":null},{"comment":"Sec. 2.1.1–2.1.2 and the geometric transform of [21]: oral TVs are obtained by an English-derived flesh-point-to-TV mapping that is assumed language-independent. The manuscript never tests whether that mapping remains faithful for French or Russian articulatory targets (e.g., by comparing alternative normalizations or reporting residual anatomical variance). If the transform itself is English-centric, the reported PPMC on French/Russian does not cleanly support a language-agnostic claim.","section":null},{"comment":"Sec. 3.3: VP training for XRMB uses nasalance pseudo-labels generated by the authors’ own prior nasal SI system [8]. While the final evaluation on YU French/Russian uses measured nasalance, the training target for a large fraction of the data is model-derived. The manuscript should quantify how sensitive the VP results are to these pseudo-labels (e.g., train without XRMB nasalance or report agreement between the prior nasal SI and true nasalance on YU English) so that the cross-lingual VP numbers are not partly circular.","section":null}],"minor_comments":[{"comment":"Table 2 / Sec. 2.1.2: the YU French and Russian test sets are small (59 min / 35 min). Explicitly state total hours and number of utterances per language in the abstract or introduction so readers can gauge statistical power immediately.","section":null},{"comment":"Eq. (1) and Sec. 2.2: the mixing weight α = 0.2 is described as “empirically chosen.” A short ablation (or at least the range explored) would strengthen reproducibility.","section":null},{"comment":"Figure 2 and Figure 4: axis labels and units for the TV / nasalance trajectories are hard to read in the manuscript rendering; ensure they remain legible in the final version.","section":null},{"comment":"Sec. 3.1: the claim that WavLM-Large “outperforms” other SSL models is supported only by citations; a one-sentence confirmation that the same ranking holds on the present multi-lingual test sets would be useful.","section":null},{"comment":"Typographical consistency: “product–moment” vs “product-moment,” and occasional spacing around PPMC values, should be standardized.","section":null}],"recommendation":"major_revision","confidential_remarks":"The core English SI pipeline and the multi-lingual data collection are solid contributions. The language-agnostic framing is currently overstated relative to N and the untested TV geometry; a revised version that reports speaker-level statistics, acknowledges the English-centric mapping as a limitation, and clarifies the pseudo-label dependence for VP would be suitable for the journal. I do not see evidence of circularity on the main oral-TV/SF results, only on the XRMB nasalance retrofit."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: they trained an English SI system (WavLM-Large + Conformers, multi-task oral TVs + Per/Aper/F0) and report average PPMC 0.83 on French and 0.74 on Russian against held-out EMA, plus VP-nasalance correlations of 0.89/0.82. That cross-lingual table is new; almost everything before this was English-only or accent work.\n\nWhat is actually new is the YU multi-lingual collection (EMA + nasalance + audio for English/French/Russian) and the first reported oral-TV + source-feature + VP numbers on French and Russian. The pipeline itself is a clean extension of their prior line, not a redesign. They do the basics right: speaker-independent splits, a direct XRMB comparison that edges their earlier model (0.86 vs 0.85), and trajectory plots that look plausible across languages. The French oral-TV result is the most credible piece of evidence for partial transfer. They also flag the noisy Russian recording themselves.\n\nSoft spots are real and proportional, not fatal. Non-English N is tiny (4 French, 3 Russian; only one Russian for nasalance), with no speaker-level scores or intervals. Russian is already pulled down by that noisy speaker. The flesh-point-to-TV geometry and WavLM features are English-derived; treating pooled PPMC as proof of language-agnostic recovery overclaims what those cohorts can support. VP training on XRMB uses pseudo-labels from their own prior nasal SI system, so that path is less independent than the oral-TV results. No code or data release makes the numbers hard to re-run.\n\nThis is for people who need audio-only proxies for articulatory timing in multi-lingual or clinical settings. It is solid subfield progress, not a methods leap. I would send it to peer review: the French numbers and the new corpus deserve referee time, with the expectation that reviewers will ask for more speakers, variance, and a quieter title. Worth engaging if you work on SI or cross-lingual phonetics; otherwise the tables are enough.","headline":"Useful first French/Russian SI numbers on a new multi-lingual EMA set, but the language-agnostic claim is ahead of the sample size and the English-rooted TV geometry.","tokens_in":9776,"tokens_out":544,"would_cite":true,"duration_ms":13853,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"An English-trained speech inversion system recovers oral tract variables, source features, and velopharyngeal opening on French and Russian speech.","keywords":["Speech Inversion","Articulatory Modeling","Tract Variables","Cross-Lingual Speech Analysis","Velopharyngeal Port","Source Features","WavLM"],"falsifier":"Collect EMA, nasalance, and audio from a larger set of French and Russian speakers (or additional languages) and recompute PPMC; a sharp drop below the reported correlations would falsify the cross-lingual claim.","tokens_in":9679,"feed_emoji":"🗣️","tokens_out":877,"duration_ms":6618,"temperature":0.7,"pith_summary":"This paper shows that a speech inversion system trained only on American English audio paired with articulatory measurements can recover the same articulatory timing patterns from French and Russian speech it never saw in training. The system jointly estimates six oral tract variables (lip and tongue constrictions), three source features (periodicity, aperiodicity, and fundamental frequency), and a velopharyngeal opening measure that tracks nasalance. On held-out French and Russian speakers the estimated trajectories match ground-truth articulatory and nasalance data with average Pearson correlations of 0.83 and 0.74 for the oral and source parameters, and 0.89 and 0.82 for velopharyngeal opening. Because collecting articulatory data is expensive and language-specific, a system that works across languages from English training alone would let researchers and clinicians study articulatory timing in many languages from ordinary audio recordings.","feed_headline":"English-trained speech inversion works on French and Russian","feed_subtitle":"Oral tract variables, source features, and nasal opening recovered from audio never seen in training","key_machinery":"Multi-task speech inversion that maps WavLM-Large embeddings through Conformer layers to six oral tract variables (LA, LP, TBCL, TBCD, TTCL, TTCD), three source features (Per, Aper, F0), and a velopharyngeal tract variable, trained with a combined Pearson-correlation and RMSE loss on English XRMB and YU data.","core_discovery":"An SI system trained exclusively on co-recorded American English speech and articulatory kinematics simultaneously estimates six oral tract variables plus three source features on previously unseen French and Russian with average PPMC scores of 0.83 and 0.74 against ground-truth measurements, and the same English-trained system also estimates a velopharyngeal tract variable that correlates with nasalance at 0.89 (French) and 0.82 (Russian).","pith_inferences":["If the transfer holds for more language families, speech inversion could serve as a low-cost proxy for field articulatory phonetics where EMA or X-ray microbeam is impractical.","The same joint oral-plus-source-plus-VP architecture may improve robustness for disordered speech or child speech once more diverse training data are added.","Performance gaps between French and Russian (and the noted noisy Russian speaker) suggest recording conditions and phonetic inventory distance remain practical limits on claimed language-agnosticism."],"forward_implications":["Articulatory timing patterns for oral constrictions and source control can be recovered from audio alone in French and Russian without language-specific articulatory training data.","Velopharyngeal opening and closing can be estimated across languages that differ in nasalization patterns (English vs. French nasal vowels) from an English-trained model.","Multi-lingual speech processing and clinical articulatory assessment become feasible from ordinary audio recordings rather than specialized articulatory hardware.","Future collection of more speakers and under-resourced languages can test and extend the same English-trained inversion pipeline."],"fun_headline_variants":["English-trained speech inversion generalizes to French and Russian","Speech inversion recovers articulatory features across unseen languages","English-only SI estimates French/Russian tracts at 0.83/0.74 PPMC","Cross-lingual speech inversion from English-trained articulatory model","SI recovers oral tracts source features and nasalance in French Russian"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The geometric mapping from flesh-point sensors to tract variables, and the speech features used, transfer well enough from English to French and Russian that results on a small number of held-out speakers measure true language-agnostic performance.","fun_headline_variants_meta":{"raw":{"variants":["English-trained speech inversion generalizes to French and Russian","Speech inversion recovers articulatory features across unseen languages","English-only SI estimates French/Russian tracts at 0.83/0.74 PPMC","Cross-lingual speech inversion from English-trained articulatory model","SI recovers oral tracts source features and nasalance in French Russian"]},"model":"grok-4.5","effort":"low","cost_usd":0.003988,"raw_usage":{"total_tokens":1152,"prompt_tokens":683,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":39880000,"prompt_tokens_details":{"text_tokens":683,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":377,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":683,"tokens_out":92,"duration_ms":3662,"temperature":1.0,"reasoning_tokens":377,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T09:24:40.916135+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Collect EMA, nasalance, and audio from a larger set of French and Russian speakers (or additional languages) and recompute PPMC; a sharp drop below the reported correlations would falsify the cross-lingual claim.","supporting_citations":[],"review_version":1}