{"id":"c77eacbf-deec-4335-a79b-1266cdb74185","arxiv_id":"2507.01888","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Inferred articulatory vocal tract variables from adult-trained speech inversion align with clinician perceptual ratings for /r/ errors in children with SSD, and less strongly for /s/.","lead":"This paper tests whether articulatory movements inferred from audio by a neural network match expert clinicians' perceptual ratings of /r/ and /s/ errors in children with speech sound disorders. It finds strong alignment for /r/ and partial alignment for /s/, suggesting audio-only kinematics could support clinical assessment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Adult-trained speech inversion is applied to child SSD audio without child ground truth; positive-control contrasts do not establish that inferred tract variables are valid for children, so the reported articulatory-perceptual alignment could reflect acoustic artifacts.","rationale":"The reader identified the same weakest assumption: adult-trained speech inversion applied to child speech without child ground-truth validation. My stress-test agrees and sharpens it: the positive-control phonemes validate phoneme discriminability in child voices, not metric validity of the inferred tract variables for children. The missing specification of how the speaker-level min-max normalization (Eq. 1) is applied at inference creates a concrete, testable technical gap. If adult min/max ranges are reused for child audio, the inferred values are not directly comparable across age groups. The RQ3 result—PERCEPT scores predicting mean squared articulatory difference—is consistent with the network learning an acoustic mapping that correlates with perceptual ratings; that is useful evidence for a clinical tool, but it is not evidence that the inferred articulatory kinematics are true. None of this invalidates the paper; it does mean the conclusion should remain conditional on child articulography validation or on clearly tempering the causal/validity language. The proposed concrete test—direct child articulography comparison, or failing that, a vocal-tract-length scaling simulation on the adult test set—would settle whether the concern lands. The paper already acknowledges the limitation, so the reader's CONDITIONAL verdict is appropriate and I recommend no change to it.","tokens_in":26363,"tokens_out":2866,"duration_ms":38335,"concrete_test":"Collect concurrent ultrasound or EMA articulography from 10–15 children with SSD (ages roughly 7–12) producing the study's wordlist, compute gold-standard Tract Variables using the same geometric transformations as Attia et al. (2024), and compare per-child, per-variable inversion accuracy (Pearson r and RMSE) for correct and error tokens. If the average r for TTCL/TTCD in children falls markedly below the adult test-set values (Table 3: 0.82–0.95), or if error-token accuracy is substantially worse than correct-token accuracy, the central claim is not supported as stated. If no child dataset is feasible, a second-best check: take adult X-ray Microbeam test utterances, apply formant and pitch scaling to child-like ranges, and verify that inversion accuracy and the /w/-vs-/ɹ/-vs-/ʌ/ ordering are preserved under scaling.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that inferred vocal tract variables track category and severity of /r/ and /s/ errors—rests on the unvalidated transfer of an adult-trained speech inversion network to child SSD speech. The paper's internal evidence (17/18 /r/ hypotheses, positive-control contrasts for /w/, /ʌ/, /θ/, /ʃ/, /l/) shows that the network can discriminate adult phoneme categories in child voices, but it does not establish that the inferred constriction locations and degrees are accurate for child vocal tracts or for atypical productions. Critical details are unspecified: Eq. 1 applies speaker-level min-max normalization, yet the paper never states what min/max values are used at inference on child audio. If adult training ranges are reused, child tract variables—being absolutely smaller—would be systematically biased, and group differences could be driven by that bias plus acoustic covariation with perceptual ratings (e.g., F3 for /r/) rather than by true articulatory kinematics. The RQ3 association is also compatible with the network simply encoding the same acoustic cues that raters use, rather than a real articulatory distance. Therefore the load-bearing assumption is not merely 'no child ground truth'; it is that the inversion output has metric validity for children, which the positive-control contrasts cannot certify.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether articulatory kinematics inferred by an adult-trained acoustic-to-articulatory speech inversion network align with expert perceptual ratings of /ɹ/ and /s/ errors in children with speech sound disorders (SSD). The authors apply a WavLM-based speech inversion system trained on the Wisconsin X-ray Microbeam adult dataset to 5,961 utterances from 118 children and 3 adults, extract six Articulatory Phonology vocal tract variables over forced-aligned phone intervals, and test 33 a priori hypotheses about tract-variable differences among correct targets, error subtypes, and paired positive-control phonemes using linear mixed models with Benjamini-Hochberg FDR correction. They also test whether PERCEPT Rating Scale scores predict mean squared articulatory distance from correct productions. Results: 17 of 18 /ɹ/ hypotheses and 7 of 15 /s/ hypotheses were supported, and the gradient analysis showed a significant negative association (β = -0.019, p < 0.001). The authors conclude that speech inversion tract variables are clinically interpretable, particularly for /ɹ/.","tokens_in":26765,"tokens_out":6022,"duration_ms":71850,"significance":"If the central inference is valid, this would be a substantial methodological advance: it would provide a non-instrumented, scalable way to obtain articulatory kinematic descriptions of speech sound errors in children with SSD, with potential clinical applications in assessment and biofeedback. The study is notable for its large pediatric corpus, the explicit pre-registration-style design with 33 a priori hypotheses, the use of positive-control phonemes, FDR-corrected contrasts, and the provision of analysis code. The authors are appropriately cautious in interpreting the /s/ results, which have lower support. The main risk is the unvalidated transfer of an adult-trained inversion model to child speech; the paper's internal evidence (positive-control contrasts, observed bunched/retroflex-like subclusters for /ɹ/) is encouraging but does not by itself certify metric validity of the inferred tract variables for children or for atypical productions.","major_comments":[{"comment":"The central claim that inferred vocal tract variables track both category and severity of articulatory errors rests on the assumption that the adult-trained speech inversion network produces metrically valid tract-variable estimates for child SSD speech. The positive-control contrasts in RQ1/RQ2 show that the network can discriminate some phonemic categories in child voices, but they do not establish that the inferred constriction locations and degrees are accurate for children's smaller vocal tracts or for atypical /ɹ/ and /s/ productions. The RQ3 association (β = -0.019, p < 0.001) could in principle be driven by the network encoding the same acoustic cues (e.g., F3 for /ɹ/, spectral moments for /s/) that the raters use, rather than by true articulatory kinematics. I request an explicit treatment of this alternative explanation, preferably with a control analysis: for instance, add acoustic distance covariates (F2/F3 for /ɹ/; spectral mean/peak for /s/) to the RQ3 model, or compare the inversion-derived articulatory distance against a purely acoustic distance in nested models. Without such a control, the 'articulatory' interpretation of the RQ3 result remains underdetermined.","section":"Methods, Speech Inversion Neural Network; Results, RQ3"},{"comment":"Equation (1) describes speaker-level min-max normalization of the vocal tract variables, applied to the adult Wisconsin X-ray Microbeam ground truth during training. The manuscript never states what normalization is used at inference when the trained network processes child SSD audio. If the network directly outputs normalized tract variables (as is implied by the training targets), the description should say so explicitly, because no child-specific min/max would then be needed. If, instead, the child audio or the network output is scaled using adult training-set minima/maxima, the systematically smaller child vocal tracts would produce biased tract-variable values, and the group differences reported in Tables 4 and 5 could be artifacts of that bias. This point is load-bearing because the magnitude and even direction of several contrast estimates are central to the conclusions.","section":"Methods, Equation (1) and Speech Inversion Neural Network"},{"comment":"The positive-control phonemes /w/, /ʌ/, /θ/, /ʃ/, and /l/ are described as 'assumed to be correctly articulated' without any perceptual verification. In children with SSD, target sounds other than /ɹ/ and /s/ can also be misarticulated; if some positive-control phones are themselves errored, the intended interpretative function of the positive controls (as articulatorily well-defined endpoints) is weakened. The authors should report whether any screening was performed on these phones, or, failing that, discuss the likely impact of this assumption on the RQ1/RQ2 validity checks. This is not a fatal flaw, but it bears directly on the internal-validation argument.","section":"Methods, Dataset and Independent Variables"}],"minor_comments":[{"comment":"The acronym 'PERCEPT' is used throughout but only expanded informally in the Methods; consider providing a definition or a brief note on its status as a proprietary research instrument.","section":"Methods, PERCEPT Rating Scale"},{"comment":"The 'Interpretation of Hypothesis' column largely restates the 'Vocal Tract Variable Interpretation' column in plain language; consider merging or removing to reduce redundancy.","section":"Table 1"},{"comment":"The captions state that higher values on the x-axis correspond to more anterior locations, but for constriction-degree variables the y-axis 'higher' values correspond to narrower constrictions; a brief clarification would help the reader.","section":"Figure captions (Figures 4-8)"},{"comment":"The custom 'PERCEPT-TX' child speech acoustic models used for forced alignment are not described or referenced; please provide details or a citation.","section":"Methods, Phone segmentation boundary alignment"},{"comment":"The authors acknowledge that averaging tract variables over the entire phone interval may attenuate differences; this is an appropriate limitation, but it could be expanded to note that dynamic measures (e.g., gestural coordination) may be more sensitive for /s/.","section":"Discussion, Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-executed in its design and reporting, and the a priori hypothesis framework is a strength. The main concern is the external validity of the adult-trained inversion model for child speech, which the authors acknowledge but do not fully resolve; the requested control analyses (acoustic-covariate models, explicit normalization details, screening of positive controls) are feasible within the existing scope. I do not see grounds for rejection, but the manuscript would benefit from a round of revision addressing these validity issues. The self-citation pattern is frequent but appears to reflect the authors' prior work on the same speech inversion systems, which is appropriate here."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe two things to know: this paper is the first to test a broad set of 33 a priori articulatory hypotheses about /r/ and /s/ errors in child SSD using adult-trained speech inversion tract variables, and the results are stronger than I expected for /r/. Second, the load-bearing assumption—that adult-trained inversion yields metric-valid vocal tract variables for child speech—is explicitly acknowledged but not validated, and the paper's internal checks don't fully close that gap.\n\nWhat's genuinely good: the study design is careful. The 33 hypotheses were set before analysis, grounded in ultrasound, MRI, and Articulatory Phonology literature. The positive-control phonemes (/w/, /ʌ/, /θ/, /ʃ/, /l/) are a smart addition; they let the authors ask whether the inversion can at least recover known phonemic contrasts in child voices. The stats are appropriate: linear mixed models with Satterthwaite df and FDR correction. 17 of 18 /r/ hypotheses are supported with large effects; the /s/ results are more mixed (7/15), which the authors interpret honestly. The RQ3 gradient result (beta = -0.019, p < .001) is real but small. They also ship the analysis code.\n\nThe soft spots, in order. The biggest is the adult-to-child transfer. The network was trained on Wisconsin X-ray Microbeam adult articulography and applied directly to child SSD audio. The paper says the vocal tract variable normalization facilitates this, but it never states what min/max values are used at inference. If the adult training ranges are reused, child tract variables are systematically biased toward adult ranges; if per-speaker child min/max are used, the network was never trained on such inputs. Either way, the positive-control contrasts show the network can discriminate adult phoneme categories in child voices, but they don't certify that the inferred constriction locations and degrees are accurate for child vocal tracts or for atypical productions. The RQ3 association is also compatible with the network encoding the same acoustic cues the raters heard (e.g., F3 for /r/), rather than a genuine articulatory distance. The authors mention the child ground-truth gap in Limitations, but the conclusion still leans on 'clinical interpretability' a bit harder than the evidence supports.\n\nThe /s/ tongue body failures are a useful pattern, not a flaw; they suggest the midsagittal tract variables are less interpretable for secondary articulators. The lateralized /s/ results are weak, and the authors openly discuss the /l/ control confound.\n\nBottom line: this deserves a serious referee. The core design is transparent and the /r/ results are impressive, conditional on the transfer assumption. A revision should (a) spell out the normalization at inference, (b) temper the clinical interpretability language, and (c) ideally add a small validation experiment (e.g., adult articulatory ground truth held out, or synthetic child vocal tract simulation). I'd bring this to reading group.\n\nRecommendation: send to review.","headline":"Systematic a priori test of speech inversion for /r/ in child SSD, with real strengths and an unvalidated adult-to-child transfer that should temper the claims.","tokens_in":27146,"tokens_out":2826,"would_cite":true,"duration_ms":32106,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Computer-inferred vocal tract variables match clinician-perceived subtypes and severity of misarticulated /r/ and /s/ in children with speech sound disorders.","keywords":["speech sound disorder","speech inversion","articulatory phonology","vocal tract variables","perceptual rating","rhotics","sibilants","child speech"],"falsifier":"Collect articulography (for instance electromagnetic articulography or ultrasound) from children with speech sound disorders producing the same /r/ and /s/ error subtypes, and compare each child's measured tongue tip and tongue body constrictions against the model's inferred vocal tract variables for the same utterances; the central claim fails if the inferred values do not reproduce the measured articulatory directions and distances within children.","tokens_in":26208,"feed_emoji":"🗣️","tokens_out":9009,"duration_ms":88659,"temperature":0.7,"pith_summary":"This paper asks whether articulatory kinematics inferred from ordinary audio recordings, through speech inversion neural networks trained on adult articulatory data, can recover clinically meaningful details of misarticulated /r/ and /s/ in children with speech sound disorders. Using inferred Articulatory Phonology vocal tract variables, the authors test 33 a priori hypotheses that link perceived error subtypes to articulatory patterns and report that 17 of 18 /r/ hypotheses and 7 of 15 /s/ hypotheses are supported. They also show that scores on a 5-point perceptual rating scale predict the mean squared articulatory distance of an errored phone from correct productions: sounds rated more severely incorrect sit further from the correct articulatory configuration. If correct, this means clinicians and researchers could perform articulatory kinematic analyses of /r/ errors from audio alone, without specialized articulography equipment.","feed_headline":"17 of 18 predicted /r/ error patterns show up in inferred articulation","feed_subtitle":"Computer-inferred tongue movements track clinician ratings of error severity in child /r/ and /s/.","key_machinery":"The central object is the Articulatory Phonology vocal tract variable set, six normalized articulatory parameters (lip aperture, lip protrusion, tongue tip constriction location, tongue tip constriction degree, tongue body constriction location, tongue body constriction degree) plus glottal source variables, inferred at 100 Hz by a speech inversion network from self-supervised audio embeddings. The network was trained on adult articulography recordings and applied unchanged to child SSD audio; the inferred variables are averaged over each phone interval and compared across phones with linear mixed models, so the vocal tract variables are the medium through which perceptual error categories and severity are mapped to articulation.","core_discovery":"The central claim is that an adult-trained speech inversion neural network, applied directly to children's SSD audio, produces Articulatory Phonology vocal tract variable estimates that track both the categorical subtype and the gradient severity of articulatory errors. For /r/, estimated marginal means from a linear mixed model support 17 of 18 hypotheses: derhotic /r/ → [w] phones show lower tongue tip constriction degree and more posterior tongue body constriction location than correct /r/, derhotic /r/ → [+vocalic] phones show less lip protrusion, lower tongue tip, and looser tongue body constriction, and positive controls /w/ and /ʌ/ fall at the expected endpoints. For /s/, 7 of 15 hypotheses are supported, chiefly those involving tongue tip constriction location for dentalized and palatalized errors, while tongue body and lateralized /s/ hypotheses fail. A third model shows a significant negative association between PERCEPT Rating Scale scores and mean squared articulatory difference from correct productions, indicating that more perceptually incorrect phones are articulatorily further from the target.","pith_inferences":["A natural next test is to collect child articulatory ground truth for the same utterances; if measured and inferred vocal tract variables agree within children, adult-trained inversion could replace articulography for monitoring /r/ treatment remotely from audio.","The two tongue tip subclusters visible in correct /r/ estimates invite the non-invasive hypothesis that speech inversion separates bunched from retroflexed /r/ variants, which would allow treatment to target the child's actual tongue shape without imaging.","The /s/ tongue body failures may be partly an artifact of midsagittal measurement; extending inversion with lateral or tongue-root variables, or with child-trained models, is the testable path toward capturing lateralized /s/ errors.","If the perceptual-to-articulatory distance association replicates across raters and sessions, the PERCEPT scale could be calibrated as a continuous measure of articulatory change in clinical trials, without requiring kinematic instrumentation."],"forward_implications":["Articulatory kinematic analysis of /r/ in childhood speech sound disorders can be conducted from ordinary recordings, bypassing articulography instrumentation for this class of questions.","PERCEPT Rating Scale scores can serve as a perceptual proxy for articulatory proximity to a correct target, supporting their use as a clinical and research outcome measure.","For /r/ error subtypes, tongue tip and tongue body variables carry most of the discriminative information, while lip protrusion contributes the least, suggesting clinical cueing should emphasize tongue configuration.","For /s/, tongue tip constriction location distinguishes fronted and backed errors, but midsagittal speech inversion does not yet reliably capture lateralized /s/ or tongue body contributions; conclusions about /s/ subtypes should be limited to tongue tip place."],"supporting_citations":[{"why":"Supplies the geometric transformations from articulography pellets to vocal tract variables and the baseline speech inversion system this study extends.","marker":"Attia et al. (2024)"},{"why":"Provides the adult X-ray microbeam articulography dataset with coregistered audio used as training ground truth.","marker":"Westbury et al. (1994)"},{"why":"Shows that glottal source features improve acoustic-to-articulatory speech inversion, motivating the network's source variables.","marker":"Siriwardena & Espy-Wilson (2023)"},{"why":"Establishes neural-network speech inversion from x-ray microbeam audio as the long-studied approach this paper applies to child SSD.","marker":"Papcun et al. (1992)"},{"why":"Prior demonstration that adult-trained speech inversion inferred tongue body constriction location differences between fully rhotic and derhotic child speech.","marker":"Benway et al. (2023)"},{"why":"Defines Articulatory Phonology, the framework from which the six vocal tract variables are drawn.","marker":"Browman & Goldstein (1992)"},{"why":"Provides evidence that subtle /s/ articulation differences shape the acoustic spectrum, grounding the /s/ hypotheses.","marker":"Munson (2004)"},{"why":"Ultrasound-based description of derhotic /r/ tongue configurations that grounds the a priori /r/ hypotheses.","marker":"Preston, Benway, et al. (2020)"}],"fun_headline_variants":["AI tongue tracking matches perceptual ratings for child /r/ errors","Inferred articulation supports 17 of 18 /r/ error hypotheses in SSD","Perceptual ratings predict articulatory proximity for /r/ and /s/","Speech inversion reveals articulatory patterns behind perceived /r/ errors","Computer-estimated tongue movements align with clinician ratings in kids"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a speech inversion network trained solely on adult articulatory recordings produces valid vocal tract variable estimates for children's smaller vocal tracts and for atypical /r/ and /s/ productions, with no child or error ground truth used to confirm or correct those estimates.","fun_headline_variants_meta":{"raw":{"variants":["AI tongue tracking matches perceptual ratings for child /r/ errors","Inferred articulation supports 17 of 18 /r/ error hypotheses in SSD","Perceptual ratings predict articulatory proximity for /r/ and /s/","Speech inversion reveals articulatory patterns behind perceived /r/ errors","Computer-estimated tongue movements align with clinician ratings in kids"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1627,"prompt_tokens":1057,"completion_tokens":570,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":673,"completion_tokens_details":{"reasoning_tokens":478}},"tokens_in":673,"tokens_out":570,"duration_ms":5960,"temperature":1.0,"reasoning_tokens":478,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:39:48.058825+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect articulography (for instance electromagnetic articulography or ultrasound) from children with speech sound disorders producing the same /r/ and /s/ error subtypes, and compare each child's measured tongue tip and tongue body constrictions against the model's inferred vocal tract variables for the same utterances; the central claim fails if the inferred values do not reproduce the measured articulatory directions and distances within children.","supporting_citations":[],"review_version":1}