{"id":"c434f489-2228-448f-95c4-15258b5cd129","arxiv_id":"2506.11072","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"OpenSMILE and Praat produce significantly different speech features from identical audio, and the choice of tool changes observed group differences and autism classification results.","lead":"An automated comparison of two open-source speech analysis tools, OpenSMILE and Praat, finds large differences in the pitch, loudness, and speech-rate features they extract from the same recordings of adolescents. The paper argues these tool differences can change which demographic or diagnostic differences appear and how well autism classification models perform.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The most load-bearing risk is not diarization noise but a possible unit mismatch: OpenSMILE eGeMAPS F0 is typically semitone-scaled, while Praat pitch is in Hz, which would make the reported pitch differences an artifact of incompatible scales.","rationale":"I read the paper in good faith. The practical recommendation—report tool and parameters—is valuable and likely robust. However, the central quantitative evidence for that recommendation is the assertion that all five features differ between tools. The diarization issue flagged by the reader is real but limited to speech rate and is admitted. A more serious, unexamined risk is that the OpenSMILE pitch features are in semitones while Praat's are in Hz. The numerical pattern in Table 1 is consistent with that explanation. This would make the pitch differences an artifact of the authors' comparison, not of the tools. It is testable from the missing code/config. If the unit mismatch is confirmed, the strongest claim should be narrowed to loudness and speech-rate operationalizations, and the paper's evidence for 'uncontrolled source of variance' would be weakened, though the recommendation to report settings remains sensible. Because the concern is testable and the paper is otherwise transparent, I retain the reader's CONDITIONAL verdict rather than moving to reject or unverdict.","tokens_in":9087,"tokens_out":9685,"duration_ms":90506,"concrete_test":"Request the exact OpenSMILE configuration and extraction script for the pitch features. Re-run the pipeline on a random subset of utterance audio, extracting both 'F0semitoneFrom27.5Hz' and the Hz-valued F0 (or convert semitone to Hz via f = 27.5 * 2^(S/12)). Compare Praat pitch to the Hz-valued OpenSMILE feature. If mean pitch difference drops from ~157 Hz to a few Hz and Pitch Std/Range differences shrink correspondingly, Section 4.1's claim for pitch features is a unit artifact. If the large differences persist after the conversion, the unit-mismatch explanation is refuted and the original claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in §4.1 depends on the unstated definition of the OpenSMILE pitch feature. The eGeMAPS configuration cited for OpenSMILE provides F0 as 'F0semitoneFrom27.5Hz', a log-scale semitone value, alongside an optional Hz-valued contour. If the authors took the eGeMAPS F0 field and labeled it 'Mean Pitch', comparing it to Praat's Hz-scale pitch would produce exactly the pattern in Table 1: a mean difference of 156.94 Hz is roughly the offset between a semitone value near 36 and a Hz value near 193 Hz for typical adolescent speech; similar scaling explains Pitch Std (31.84) and Pitch Range (125.62). The paper does not specify which OpenSMILE F0 output was used, and no code is provided. If this is true, three of the five features in the headline claim are not measuring the same quantity, so the 'significantly different' result is an artifact of unit mismatch rather than evidence that OpenSMILE and Praat measure pitch differently. This is more load-bearing than the acknowledged diarization confound, which affects only the speech-rate feature and is already disclosed in §4.1.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares five speech features (mean F0, F0 standard deviation, F0 range, loudness, and speech rate) extracted with OpenSMILE (eGeMAPSv02) and Praat (Parselmouth) from automatically diarized ADOS-2 recordings of 29 adolescents (14 ASD, 15 TD; 21 male, 8 female) across 14 tasks. Using task-level aggregates, it reports large statistically significant differences between tools for all five features (Table 1), ANOVA-based demographic/tool effects (Table 2), and leave-one-user-out random forest classification of ASD versus TD for each task (Table 3). The paper concludes that open-source speech tools are not interchangeable, that feature-extraction choices affect downstream group-level conclusions and classification performance, and that researchers should report their exact parameter choices.","tokens_in":9324,"tokens_out":11817,"duration_ms":119373,"significance":"If the central comparison were clean, this would be a useful and timely finding: feature extraction is an under-audited link in the clinical speech-ML pipeline, and the paper shows that task, diagnostic group, and tool choice can all shift surface conclusions. The comparison is empirical and external; no fitted parameter enters the headline pairwise differences; and the authors explicitly disclose the diarization/segmentation confound in Section 4.1 and Section 5.1 rather than hiding it. The practical recommendation to report tool versions, configurations, and preprocessing choices is well motivated. However, the value of the paper hinges on whether the headline differences reflect genuine measurement divergence. As written, that is not established for the pitch and loudness features because the OpenSMILE output definitions are not stated, and the speech-rate row is explicitly contaminated by segmentation artifacts. The study is also small (n=29), with multiple-testing concerns in the ANOVA section and no uncertainty quantification in the classification section. The topic deserves attention, but the strong claims need correction and re-analysis before they can stand.","major_comments":[{"comment":"The manuscript says only that 'for pitch, we used fundamental frequency (F0)', but it does not state which OpenSMILE eGeMAPS output field was captured. The eGeMAPSv02 configuration exposes pitch as F0semitoneFrom27.5Hz, a semitone-scale value, whereas Praat pitch is in Hz; the mean differences in Table 1 (Mean Pitch: 156.94, Pitch Std: 31.84, Pitch Range: 125.62) are exactly the pattern expected from combining values in these two different units. The same issue likely affects the loudness comparison: eGeMAPS Loudness is a perceptual loudness measure in sone, while the paper's Praat 'loudness' is based on intensity in dB. A paired t-test on raw values in different units is not a test of whether the tools measure the same construct. The authors must specify the exact OpenSMILE features used; if semitone/sone values were used, the comparisons must be redone on a common scale (for example, converting semitone values back to Hz) or re-scoped to the features that genuinely share units. At minimum, a correlation or rank-correlation analysis after unit conversion would show whether the tools agree up to a monotone transformation.","section":"Section 3.1 / Section 4.1 / Table 1"},{"comment":"The speech-rate row is not a clean comparison of tool behavior. The paper reports an OpenSMILE mean of 5.78 syllables/second versus a Praat mean of 54.12 syllables/second, which the authors themselves describe as behaviorally impossible and attribute to diarization noise, silence, clipped audio, and misattributed turns. Because this artifact drives the 48.37 mean difference in Table 1, the speech-rate result cannot be cited as evidence of how the tools' speech-rate algorithms differ on valid audio. The authors should redo this part of the analysis after filtering implausible syllable counts or using manually validated utterance boundaries, and should report medians or trimmed means and the number of utterances excluded. The acknowledgment in Section 4.1 is helpful, but it does not cure the fact that the headline Table 1 includes an artifact-dominated value.","section":"Section 4.1 / Figure 1 / Table 1"},{"comment":"The demographic comparison runs a large number of tests without a stated global correction: five features, fourteen tasks, and two group factors. Table 2 then reports individual p<0.05 effects, several with partial eta-squared values around 0.07 to 0.17. Many of these would not survive a family-wise or false-discovery-rate correction. In addition, the text says 'Bonferroni adjustments' but cites Reference [28], which is the Benjamini-Hochberg false-discovery-rate paper, and the post-hoc columns report significance for one tool without exact p-values for the other. The authors should state the total number of tests, prespecify a correction procedure, apply it consistently, and fix the citation.","section":"Section 3.2 / Section 4.2 / Table 2"},{"comment":"The statement that 'OpenSMILE generally performed worse than Praat' on F1 is not supported by Table 3: OpenSMILE has equal or higher F1 in many task/group cells (for example, Creating, Cartoons, Demonstration, Telling, Description, Conversation for ASD, Social for TD, Construction for TD, and Joint for both groups). The table reports point estimates only, with no confidence intervals, error bars, or significance tests, and the leave-one-user-out procedure on 29 participants produces noisy per-task estimates. The qualitative conclusion that performance varies by tool, task, and group may be true, but the specific 'generally worse' claim needs summary statistics across tasks and uncertainty quantification before it can be accepted.","section":"Section 4.3 / Table 3"}],"minor_comments":[{"comment":"The caption contains a typo: 'extracted form Praat and OpenSMILE' should be 'extracted from Praat and OpenSMILE'.","section":"Table 1 caption"},{"comment":"The conversion factor for OpenSMILE speech rate, 'assuming 1.5 words per segment', is stated without a source or a sensitivity check; it is a constant and cannot by itself create cross-tool differences, but it does affect the reported OpenSMILE speech-rate values and should be justified.","section":"Section 3.1"},{"comment":"There is no code or data availability statement. For a paper whose central message is that parameter choices matter, releasing the exact eGeMAPS configuration, Praat/Parselmouth version, pitch floor/ceiling settings, and diarization parameters would substantially increase reproducibility.","section":"Reproducibility"},{"comment":"The wording 'only present when using Praat (p=0.2801)' is ambiguous; the sentence likely means that the OpenSMILE comparison gave p=0.2801, but both p-values should be written explicitly.","section":"Section 4.2"},{"comment":"The column headers 'Precision (Praat — OpenSMILE)' do not align clearly with the ASD/TD subcolumns, and the doubling of numbers per cell is confusing; clarify the layout and define which number belongs to which tool.","section":"Table 3"},{"comment":"The paired-sample t-tests are reported only as p<0.001; please report the t-values and degrees of freedom, and clarify whether the tests are computed per task or on the aggregated means used in Table 1.","section":"Section 4.1 statistical reporting"}],"recommendation":"major_revision","confidential_remarks":"The unit-mismatch concern is the most important issue to resolve. I believe it is probable, not merely possible, that the OpenSMILE pitch values were read from the semitone-scale F0 field and the loudness values from a sone-scale loudness field. If so, four of the five headline feature comparisons are comparing different quantities, and the central claim of Section 4.1 would need substantial reframing. I would not reject the paper outright because a focused revision can resolve the issue: specify the features, rerun or convert to a common scale, report the speech-rate results after artifact filtering, and add multiple-testing corrections. The paper is otherwise a well-scoped empirical contribution that fits the journal's interests, but the editor should require a clear confirmation of the exact OpenSMILE output definitions and a reproducibility statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, but with one serious caveat that the reader's report missed. The paper does something the field needs: it takes real autism-assessment audio, runs two standard open-source speech feature extractors on the same utterances, and shows that downstream classifier results differ by task and group. The recommendation—report not just the tool but the exact parameter configuration—is solid and well motivated.\n\nThe problem is that the headline claim, that all five features are significantly different between tools, may be an artifact of comparing unlike quantities. The eGeMAPS configuration cited for OpenSMILE outputs pitch as F0semitoneFrom27.5Hz, a semitone scale, while Praat returns Hz. The paper never states which OpenSMILE pitch field was used. The reported mean pitch difference of ~157 Hz is suspiciously close to the offset you would get comparing a semitone value near 36 against a Hz value near 193 for typical adolescent speech. Pitch Std and Pitch Range show the same pattern. If that is what happened, then three of the five features in Table 1 are not measuring the same thing at all, and the paired t-test results for those features are meaningless. The paper also doesn't provide code or configuration files, so the reader can't resolve this from the text.\n\nThe other soft spots are more minor and mostly disclosed: the speech-rate outlier is acknowledged as a diarization artifact, the sample is small and gender-imbalanced, and multiple ANOVAs are run without correction. But the unit issue is load-bearing. It doesn't sink the paper's general message—tool choice matters—but it does mean the specific quantitative results in Section 4.1 should not be quoted until the authors clarify their feature definitions and reproduce the comparison with matched units.\n\nIs it worth peer review? Yes, conditionally. The question is important, the paper is honest about its limitations, and the authors are not overclaiming clinical significance. But a serious referee needs to require exact feature names, unit conversion, and ideally a release of the extraction scripts. If the unit mismatch is confirmed, the paper becomes an even more useful cautionary tale about how easy it is to make this mistake. I would not cite the specific numbers in my own work yet, but I might bring it to a reading group as a case study in tool pitfalls.","headline":"A well-intentioned cautionary study whose central pitch comparisons may be invalid due to an apparent unit mismatch between OpenSMILE and Praat; it needs major clarification before the numbers can be trusted.","tokens_in":9834,"tokens_out":2199,"would_cite":false,"duration_ms":23203,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"OpenSMILE and Praat return significantly different values for the same speech audio, and the choice of tool changes machine-learning predictions about autism diagnosis.","keywords":["speech feature extraction","OpenSMILE","Praat","autism spectrum disorder","machine learning reliability","reproducibility","speech rate","prosodic features"],"falsifier":"Take the same ADOS-2 recordings, produce utterance-level audio via manual diarization and forced alignment, and re-extract the five features with both tools. If Praat's speech-rate estimates fall to plausible values and the cross-tool differences largely disappear, the paper's reliability conclusion is an artifact of preprocessing rather than a property of the tools; if the differences persist on clean audio, the claim of tool-level unreliability is confirmed.","tokens_in":8880,"feed_emoji":"🎙️","tokens_out":6816,"duration_ms":62888,"temperature":0.7,"pith_summary":"The paper tries to establish that feature-extraction tool choice is an uncontrolled source of variance in behavioral speech modeling. Using the same utterance-level audio from autism diagnostic sessions, the authors show that OpenSMILE and Praat return statistically different values for every feature examined—mean pitch, pitch variability, pitch range, loudness, and speech rate—with p < 0.001. These differences are not neutral: two-way ANOVAs find group differences (by gender and by diagnosis) that appear only with one tool, and random-forest classifiers trained on the two tools' features perform differently across tasks and diagnostic groups. The authors do not claim one tool is correct; they argue that default open-source feature extraction lacks validation, so published clinical machine-learning results may not be comparable, and they call for reporting tool parameters and domain-relevant verification.","feed_headline":"Speech tools disagree on the same audio, study finds","feed_subtitle":"OpenSMILE and Praat produce statistically different pitch, loudness, and speech-rate values from identical recordings.","key_machinery":"The engine is a paired comparison design on a shared audio corpus. Each utterance-level audio file from ADOS-2 sessions is processed by both tools under default settings; the five features are aggregated to task level, then compared with paired t-tests and two-way ANOVAs with tool-by-group interaction, followed by Bonferroni post-hoc tests. The same features feed separate leave-one-user-out random forests to measure downstream classification consequences. The sanity check that exposes the mechanism is the speech-rate comparison: Praat's intensity-based syllable counting yields behaviorally impossible rates on noisy automated diarization output, whereas OpenSMILE's heuristic voiced-segment mapping stays near plausible values.","core_discovery":"On the paper's own terms, the central discovery is a mismatch: five speech features extracted from identical audio by OpenSMILE's eGeMAPS set and by Praat via Parselmouth are all significantly different (paired t-test, p < 0.001), with mean differences such as 156.9 Hz for mean pitch and 48.4 syllables per second for speech rate. The speech-rate gap is behaviorally implausible—Praat sometimes reports a mean of 54.12 syllables per second when typical adult speech is 4–6—and the authors attribute Praat's instability to its intensity-based syllable detection being sensitive to diarization noise, silence, and clipped audio. Group comparisons reinforce the point: Praat alone detected gender differences in several tasks and diagnostic differences in speech rate during the Cartoon task, while OpenSMILE did not. When the features are fed to leave-one-user-out random-forest classifiers, performance varies by tool, task, and diagnostic group (e.g., Praat's TD recall in Cartoons is 0.18 vs. ASD 0.77), so no single metric or tool gives a stable picture. The conclusion is that unvalidated default feature extraction can change both the measured behavior and the model's apparent accuracy.","pith_inferences":["If the cross-tool gap is driven largely by diarization artifacts, then re-running the comparison on manually segmented or forced-aligned audio should shrink the speech-rate gap; this is a testable implication the paper does not pursue.","The same tool-variance problem likely affects other clinical speech domains such as dementia and depression, and other tools such as librosa and ProsodyPro, so a shared benchmark corpus with reference acoustic values would help separate tool noise from behavioral signal.","A pragmatic design rule follows: report both tools or a calibrated third measure when the construct is speech rate, since intensity-based syllable counting is artifact-prone on automatically diarized clinical recordings.","The paper's per-task and per-group performance tables suggest that model selection should optimize for clinically meaningful error balance, such as ASD recall, rather than average F1, because tool choice flips which group is disadvantaged."],"forward_implications":["Studies that use different speech tools on comparable clinical audio cannot be assumed to measure the same construct; cross-study effect sizes may be inflated or obscured by tool differences.","Demographic or diagnostic differences found with a single tool should be treated as provisional until replicated with another tool or a validated reference.","Classification results should be reported per diagnostic group and per task, since aggregate accuracy hides large recall disparities, such as the TD recall of 0.18 versus ASD recall of 0.77 in the Cartoon task with Praat.","Method sections in clinical speech machine learning should state tool version, parameter set, and preprocessing choices as a reproducibility requirement."],"supporting_citations":[{"why":"Provides the OpenSMILE extractor whose default features are being evaluated.","marker":"[9]"},{"why":"Supplies the Praat tool whose default algorithms form the comparison feature set.","marker":"[16]"},{"why":"Defines the eGeMAPS parameter set used to extract OpenSMILE features.","marker":"[26]"},{"why":"Introduces Parselmouth, the Python interface used to run Praat in the pipeline.","marker":"[27]"},{"why":"Performs the speaker diarization that creates the utterance-level audio segments.","marker":"[24]"},{"why":"Describes the ADOS-2 diagnostic instrument that provides the recording tasks and diagnostic labels.","marker":"[23]"},{"why":"Supplies the syllables-per-second norm used to justify the speech-rate measure.","marker":"[22]"},{"why":"Motivates the study by showing that commonly used clinical speech features lack repeatability validation.","marker":"[7]"}],"fun_headline_variants":["OpenSMILE vs Praat: Same audio, wildly divergent features","Top speech tools give conflicting measurements on identical recordings","Autism study: Speech tools produce behaviorally implausible rates","Two speech tools disagree so much they change ML predictions","Praat reports 54 syllables/sec; OpenSMILE disagrees on same clip"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main load-bearing assumption is that the utterance-level audio produced by automatic speaker diarization is clean enough that the large cross-tool differences reflect genuine tool behavior rather than segmentation artifacts such as silence, clipping, or misattributed turns.","fun_headline_variants_meta":{"raw":{"variants":["OpenSMILE vs Praat: Same audio, wildly divergent features","Top speech tools give conflicting measurements on identical recordings","Autism study: Speech tools produce behaviorally implausible rates","Two speech tools disagree so much they change ML predictions","Praat reports 54 syllables/sec; OpenSMILE disagrees on same clip"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000731,"raw_usage":{"total_tokens":3266,"prompt_tokens":935,"completion_tokens":2331,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":2242}},"tokens_in":551,"tokens_out":2331,"duration_ms":17374,"temperature":1.0,"reasoning_tokens":2242,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:28:41.098093+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same ADOS-2 recordings, produce utterance-level audio via manual diarization and forced alignment, and re-extract the five features with both tools. If Praat's speech-rate estimates fall to plausible values and the cross-tool differences largely disappear, the paper's reliability conclusion is an artifact of preprocessing rather than a property of the tools; if the differences persist on clean audio, the claim of tool-level unreliability is confirmed.","supporting_citations":[{"cited_title":"Analysis of en- gagement behavior in children during dyadic interactions using prosodic cues,","cited_arxiv_id":null,"evidence_quote":"Defines the eGeMAPS parameter set used to extract OpenSMILE features."},{"cited_title":"Classifying autism from crowd- sourced semistructured speech recordings: machine learning model comparison study,","cited_arxiv_id":null,"evidence_quote":"Provides the OpenSMILE extractor whose default features are being evaluated."},{"cited_title":"Analyzing short term dynamic speech fea- tures for understanding behavioral traits of children with autism spectrum disorder","cited_arxiv_id":null,"evidence_quote":"Supplies the Praat tool whose default algorithms form the comparison feature set."},{"cited_title":"Abnormal speech spectrum and increased pitch vari- ability in young autistic children,","cited_arxiv_id":null,"evidence_quote":"Introduces Parselmouth, the Python interface used to run Praat in the pipeline."},{"cited_title":"Quantitative analysis of pitch in speech of children with neu- rodevelopmental disorders,","cited_arxiv_id":null,"evidence_quote":"Performs the speaker diarization that creates the utterance-level audio segments."},{"cited_title":"Quantification of speech and syn- chrony in the conversation of adults with autism spectrum disor- der,","cited_arxiv_id":null,"evidence_quote":"Describes the ADOS-2 diagnostic instrument that provides the recording tasks and diagnostic labels."},{"cited_title":"Praat: doing phonetics by computer,","cited_arxiv_id":null,"evidence_quote":"Supplies the syllables-per-second norm used to justify the speech-rate measure."},{"cited_title":"V ocal markers of autism: Assessing the generalizability of machine learning models,","cited_arxiv_id":null,"evidence_quote":"Motivates the study by showing that commonly used clinical speech features lack repeatability validation."}],"review_version":1}