{"id":"62c1dddb-ef4b-4f2e-8703-e127ea62e6e3","arxiv_id":"2501.04789","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"On a sustained-vowel dataset of disordered voices, FCN-F0 achieved the highest per-frame fundamental frequency accuracy (96%) and the lowest subharmonic error rate among five estimators tested.","lead":"This study compared five fundamental frequency estimators on 15,941 frames of pathological sustained vowels, using manual annotations and a subharmonic-to-harmonics ratio to classify errors. It found that the deep-learning model FCN-F0 was most accurate overall and best at handling subharmonic voice signals, with CREPE and Harvest close behind.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'best' ranking is not yet supported without cluster-aware inference: ~15.9k intervals come from 703 recordings, so a few subharmonic-heavy recordings could drive the 96-vs-95% gap and the 34% subharmonic-error reduction.","rationale":"I read the paper as a careful, clinically motivated empirical comparison, and the experimental design has genuine strengths: it uses a substantial disordered-voice corpus, applies publicly available estimators under their default settings, and analyzes per-frame outcomes rather than only recording-level averages. The manual annotation protocol is described in detail, and the authors explicitly acknowledge the hardest case: strong subharmonics can make a female voice resemble a male voice. That acknowledged limitation is real and could affect absolute accuracy values. However, I do not think it is the single most load-bearing weakness for the central ranking claim, because the authors state that such annotation errors are likely few and would 'slightly change' the numbers. The more directly testable threat is the absence of any inferential procedure that accounts for the clustered structure of the data. Fifteen thousand intervals from 703 recordings are not 15,000 independent observations; a few long, subharmonic-heavy recordings could dominate the error counts and produce the observed 1-point accuracy differences and the 34% subharmonic-error reduction. This is not an accusation of bias—it is a missing analysis that the paper's own data could supply. The reader's conditional verdict already captures the need for additional support, and my concern reinforces that condition rather than replacing it. I therefore leave the verdict unchanged, but I would make the concrete cluster-aware test a stated requirement before the numeric ranking is treated as definitive. The paper should also be credited for publishing the classification methodology and for comparing against a primitive ACF baseline, which makes the subharmonic-error patterns easier to interpret. Overall, no fatal flaw is apparent, but the precision of the central claim currently exceeds what the statistical analysis supports.","tokens_in":11066,"tokens_out":4674,"duration_ms":50223,"concrete_test":"Perform recording-level inference: for each recording, compute per-estimator correct and subharmonic-error rates, then run a cluster bootstrap by recording and/or a mixed-effects logistic regression with a random intercept for recording to obtain 95% confidence intervals for (a) FCN-F0 minus CREPE accuracy, (b) FCN-F0 minus Harvest accuracy, and (c) the subharmonic-error rate difference. Also run McNemar's paired test stratified by recording. If any confidence interval includes zero or the ordering reverses under resampling, the 'best' claim should be softened to 'not statistically distinguishable.' A conservative secondary check is to repeat the full analysis using one randomly selected interval per recording.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central performance claim rests on pooled per-interval rates in Fig. 4: 15,941 intervals from only 703 recordings, yet all differences are reported without confidence intervals, paired tests, or any adjustment for within-recording correlation. Intervals from the same sustained vowel are strongly non-independent, and pathological voices can switch modes within and across recordings, so a small number of recordings with sustained strong subharmonics can contribute large blocks of correlated errors. The FCN-F0-vs-CREPE accuracy gap is roughly 160 intervals, and the claimed 34% subharmonic-error reduction corresponds to about a 1-percentage-point absolute difference; both are within the range that a handful of recordings could plausibly account for. The Limitations paragraph acknowledges annotation uncertainty, which is appropriately flagged, but it does not address this inferential gap. Consequently, the headline conclusion that FCN-F0 'performed the best' is not established with the precision the numbers imply, even if every manual annotation is correct.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares six fundamental-frequency estimators—an autocorrelation (ACF) baseline, Praat, YAAPT, Harvest, CREPE, and FCN-F0—on 50-ms intervals of sustained /a/ vowels from the KayPENTAX Disordered Voice Database. The authors manually annotate a speaking fundamental frequency f_o^* for each interval, classify each estimate as correct, subharmonic error, or other error using intervals of the harmonic power profile P(f_o), and compute a subharmonics-to-harmonics ratio (SHR) from ACF-based subharmonic candidates. The reported headline result is that FCN-F0 has the highest per-interval accuracy (96% vs. 95% for CREPE and Harvest, 88% for Praat and YAAPT, 62% for ACF), reduces subharmonic errors by 34% relative to CREPE, and no estimator reliably handles strong subharmonics (SHR > -3 dB).","tokens_in":11289,"tokens_out":5290,"duration_ms":54057,"significance":"If the accuracy ranking were established, the paper would be a useful clinical benchmark: it evaluates per-frame behavior rather than per-recording averages, uses publicly available estimator implementations, and stratifies by subharmonic strength. The SHR-conditional analyses and the contingency tables against the ACF baseline are informative and go beyond simple overall accuracy. The central ranking, however, is currently supported only by pooled per-interval rates without uncertainty quantification or cluster-aware inference, and the manual ground truth is acknowledged to be uncertain precisely in the subharmonic cases that matter. These issues are fixable with additional analysis, and the clinical application is potentially valuable for voice assessment.","major_comments":[{"comment":"The central claim that FCN-F0 'performed the best' rests on pooled per-interval percentages computed over 15,941 intervals from only 703 recordings, with no confidence intervals, paired tests, or adjustment for within-recording correlation. Sustained-vowel intervals from the same recording are strongly non-independent; a small number of pathological recordings with sustained strong subharmonics can contribute large blocks of correlated errors. The 96%-vs-95% accuracy gap between FCN-F0 and CREPE/Harvest corresponds to roughly 160 intervals, and the claimed 34% subharmonic-error reduction is about a one-percentage-point absolute difference—both are within the range that a few recordings could plausibly account for. I request cluster-robust inference (e.g., bootstrap by recording or a mixed-effects logistic regression with a recording random effect), with 95% confidence intervals reported for all accuracy and subharmonic-error differences. Without this, the headline ranking is not established at the precision implied by the text.","section":"Section III, Figs. 4-6"},{"comment":"The accuracy numbers depend entirely on the manually annotated f_o^*, which the authors acknowledge is uncertain for sustained strong subharmonics (e.g., a female voice with strong subharmonics could be mistaken for a normal male voice). Because the top three estimators differ by only about one percentage point, even a small number of annotation errors could reorder CREPE, Harvest, and FCN-F0. Please add a sensitivity analysis: exclude or relabel the intervals flagged as uncertain, quantify annotator agreement (e.g., a second annotation pass or test-retest reliability), or use an objective reference on a subset. The current statement that errors 'could slightly change' the accuracies is not sufficient to support the precision of the claimed ranking.","section":"Section II, 'fo Annotation' and the Limitations paragraph"},{"comment":"The rule that labels an estimate correct or as a subharmonic error based on whether it falls in an interval bounded by local minima of P(f_o) is novel and central to all error rates in Figs. 4-6, but no validation of this classification is provided. If the interval boundaries are sensitive to the periodogram's 0.5-Hz resolution, the Hamming window, or the choice of local-minimum search, the error counts could shift. I ask for at least a robustness check on a subset of intervals (e.g., comparison with a perceptual or manual classification) so that the reader can see that the quality labels are not driving the ranking.","section":"Section II, Quality-of-Estimate Classification"}],"minor_comments":[{"comment":"The percentages in the bar chart would be easier to interpret alongside a table reporting exact counts, percentages, and confidence intervals for every estimator and error category.","section":"Section III, Fig. 4"},{"comment":"The sentence 'reduces the subharmonic errors (34% less than CREPE)' is ambiguous without the absolute rates; please state both rates explicitly (e.g., CREPE 3.0% vs. FCN-F0 2.0%).","section":"Section III, paragraph beginning 'The types of the estimation errors...'"},{"comment":"The decision to resample all signals to 8 kHz is stated, but the rationale is not; since CREPE is then run at 16 kHz, please explain whether the 8-kHz lowpass filtering has any effect on the upper range of f_o or on the harmonic/subharmonic analysis.","section":"Section II, Acoustic Data"},{"comment":"The comparison with Vaysse et al. is interesting but would be more useful with a direct statement of which recording types and annotation conventions differ, since the reader cannot evaluate the claim from the cited abstract alone.","section":"Section III, 'There is also a notable discrepancy...'"},{"comment":"The paper would benefit from a data/code availability statement, including access to the custom spectrogram annotation program and the scripts used for the estimators, to support reproducibility of the 15,941-interval annotations and the contingency-table analyses.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the empirical design is appropriate for a comparison study. The requested revisions concern statistical inference and sensitivity analysis rather than a fundamental redesign; the central claim is plausible but not yet established. No novelty disclosure issue is apparent. If the authors can provide cluster-robust uncertainty estimates and a sensitivity analysis for the manual annotations, the paper could be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the per-frame quality-of-estimate classification that separates subharmonic errors from other errors, plus the use of SHR to characterize subharmonic strength. Applied to 15,941 intervals from 703 KayPENTAX sustained /a/ recordings, that is a useful contribution for clinical voice analysis. The manual annotation was careful—spectrogram review, sex-appropriate checks, and a stated limitation about strong subharmonics—and the paper is honest about that uncertainty. The scatter plots, contingency tables, and SHR histograms are informative and make the behavior of each estimator easy to grasp.\n\nThe soft spot is the one the stress-test note flags: the central claim that FCN-F0 is best rests on a 96% vs 95% correct-rate gap and a 34% relative reduction in subharmonic errors, which is roughly a one-percentage-point absolute difference. No confidence intervals, paired tests, or any correction for within-recording correlation are reported. With 15,941 intervals coming from only 703 recordings, intervals from the same sustained vowel are strongly non-independent, and a handful of subharmonic-heavy recordings could plausibly account for the gap. So the ranking is plausible, and the direction of the finding may well be right, but the precision the numbers suggest is not supported. A per-recording bootstrap or mixed-effects model would settle this.\n\nThe annotation uncertainty is a second soft spot, but the paper already acknowledges it, and it does not clearly favor one estimator over another. The custom error classification is also not externally validated, and the dataset and annotations are not released, which limits reproducibility—though the estimators themselves are public and the method is described well enough to re-implement.\n\nI disagree with the harsher version of the circularity concern: using Praat to seed the annotation could bias toward Praat, but the top estimators are deep-learning models, so the central ranking is not forced. Nor do I see invented entities or fitted parameters that reduce to the result; it is an empirical study.\n\nThis paper is for voice scientists, clinical acoustic analysts, and pitch-estimator developers. It deserves serious peer review. The revision should add cluster-aware inference and, ideally, release the annotations. I would engage with it.","headline":"A solid empirical comparison of pitch estimators on subharmonic voices, but the headline accuracy gaps lack cluster-aware statistics and should be read as trends, not a settled ranking.","tokens_in":11741,"tokens_out":1273,"would_cite":true,"duration_ms":14583,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FCN-F0, a deep-learning pitch tracker, is most accurate on subharmonic voice signals.","keywords":["fundamental frequency estimation","subharmonics","voice disorders","pitch detection","deep learning","CREPE","FCN-F0","subharmonics-to-harmonics ratio"],"falsifier":"Re-run the comparison on the same recordings with ground-truth f0 established by glottal inverse filtering or simultaneous high-speed videoendoscopy or electroglottography, and check whether FCN-F0's 96% accuracy and its 34% subharmonic-error reduction over CREPE persist; a second check would be testing the ranking on running speech with subharmonics, since the paper attributes its disagreement with [27] to the difference between sustained vowels and connected speech.","tokens_in":10896,"feed_emoji":"🎤","tokens_out":5148,"duration_ms":45096,"temperature":0.7,"pith_summary":"This paper asks which fundamental-frequency estimators can correctly identify the speaking pitch of pathological voices that contain subharmonic vibrations, a condition that makes the true pitch ambiguous because the signal repeats over several glottal cycles. Using 15,941 manually annotated 50-ms intervals from 703 recordings of sustained /a/ vowels, the authors compared five estimators plus a baseline against a quality-of-estimate classification that separates subharmonic errors (estimates at $f_o^*/M$) from other errors. The paper claims that FCN-F0, a fully convolutional deep network, is the most accurate overall (96% correct estimates) and reduces subharmonic errors by a third relative to CREPE, with CREPE and Harvest close behind at 95%. It also claims that no estimator reliably handles strong subharmonics (SHR above $-3$ dB), a reliability limit that matters because clinical acoustic parameters such as jitter and shimmer depend on the speaking fundamental frequency and mishandling subharmonics can cause false negatives.","feed_headline":"FCN-F0 wins pitch-estimator comparison on subharmonic voice","feed_subtitle":"96% correct estimates on disordered sustained vowels; CREPE and Harvest follow at 95%.","key_machinery":"The study's key machinery is a quality-of-estimate classification built on the harmonic power profile $P(f_o)=\\sum_{k=1}^{K(f_o)}|S_{xx}(k f_o)|^2$, where $S_{xx}$ is the periodogram of the 50-ms Hamming-windowed signal and $K(f_o)=\\lfloor f_s/(2f_o)\\rfloor$. For each annotated truth $f_o^*$, the profile's local minima define intervals around $f_o^*/M$; an estimate landing in the $M=1$ interval is correct, in an $M>1$ interval is a subharmonic error, and elsewhere is another error. Subharmonic strength is measured by the subharmonics-to-harmonics ratio (SHR), computed from the autocorrelation baseline's estimates, which quantifies the power of subharmonic tones relative to harmonic tones. These tools separate incidental autocorrelation errors (the main SHR peak near $-25$ dB) from true subharmonic intervals (secondary peak near $-10$ dB) and attribute each estimator's failures.","core_discovery":"The central discovery is that FCN-F0, a deep-learning model, performs the best both in overall accuracy and in correctly resolving subharmonic signals: it achieved 96% correct estimates versus 95% for CREPE and Harvest, 88% for Praat and YAAPT, and 62% for the autocorrelation baseline. FCN-F0 reduced subharmonic errors by 34% compared with CREPE, and its error profile was the most balanced between subharmonic and other mistakes. The paper further establishes a reliability boundary: among the 755 intervals with subharmonics-to-harmonics ratio above $-10$ dB, FCN-F0 correctly estimated only 63.7%, and none of the estimators could reliably handle strong subharmonics with SHR above $-3$ dB.","pith_inferences":["The paper leaves implicit that the deep-learning models' advantage may come from using features beyond raw periodicity (such as harmonic amplitudes and phases), which would predict that they generalize less well to languages or recording conditions absent from their training corpora.","The 8-kHz resampling used in the study could limit estimates for very high-pitched voices; a testable extension is whether a higher sampling rate changes the ranking among the top three estimators.","The SHR-based classification could be turned into a practical clinical flag: intervals with ACF subharmonic errors and high SHR are likely true subharmonics, so a confidence mask accompanying F0 estimates could improve downstream diagnostics.","Because FCN-F0's training data are non-pathological English and French speech, a broader pathological corpus spanning more voice types and languages is a natural next test of whether its 96% accuracy is robust."],"forward_implications":["Clinical acoustic analysis of sustained vowels should prefer FCN-F0, with CREPE and Harvest as strong alternatives, to reduce false negatives caused by subharmonic voicing.","For intervals with SHR above $-3$ dB, no current estimator should be trusted; acoustic parameters computed there should be flagged as unreliable.","Praat's Viterbi postprocessing fixes incidental autocorrelation errors but not true subharmonic intervals, so its subharmonic error rate of about 9.3% is a lower bound for the fraction of subharmonic intervals in the dataset.","Retraining deep-learning models with subharmonic voice samples, as the paper suggests, may improve accuracy for the high-SHR cases that currently defeat all estimators."],"supporting_citations":[{"why":"Supplies the CREPE deep-learning model architecture and pretrained weights used as one of the compared estimators.","marker":"[35]"},{"why":"Supplies the FCN-F0 fully convolutional network architecture and pretrained weights, the estimator that performed best.","marker":"[36]"},{"why":"Describes the Harvest estimator, the leading non-data-driven estimator compared and the closest competitor to the deep-learning models.","marker":"[33]"},{"why":"Describes the Praat f0 detector, one of the compared estimators and the basis of the autocorrelation baseline.","marker":"[32]"},{"why":"Describes the YAAPT estimator, compared for its reported ability to handle subharmonics.","marker":"[34]"},{"why":"Prior comparison reporting YAAPT and FCN-F0 as subharmonic-capable; this study confirms and partially contradicts those findings.","marker":"[27]"},{"why":"Provides the time-varying harmonic model with gradient-based optimization used to refine the manual f0 annotations.","marker":"[30]"},{"why":"Provides the KayPENTAX Disordered Voice Database recordings used for all experiments.","marker":"[29]"},{"why":"Relates SHR values to amplitude modulation extent, supporting the interpretation of the $-10$ dB threshold.","marker":"[41]"}],"fun_headline_variants":["FCN-F0 dominates pitch estimation on subharmonic voice","Deep learning wins subharmonic pitch estimation showdown","FCN-F0 beats CREPE and Harvest on subharmonic voice","AI pitch estimator FCN-F0 tops subharmonic voice test","FCN-F0 hits 96% on subharmonic pitch, beats all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The manual annotation of the true speaking fundamental frequency is itself uncertain for strong subharmonics, where a female voice with strong subharmonics could be mistaken for a normal male voice; all accuracy numbers rest on this ground truth.","fun_headline_variants_meta":{"raw":{"variants":["FCN-F0 dominates pitch estimation on subharmonic voice","Deep learning wins subharmonic pitch estimation showdown","FCN-F0 beats CREPE and Harvest on subharmonic voice","AI pitch estimator FCN-F0 tops subharmonic voice test","FCN-F0 hits 96% on subharmonic pitch, beats all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0007,"raw_usage":{"total_tokens":3098,"prompt_tokens":817,"completion_tokens":2281,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":2196}},"tokens_in":433,"tokens_out":2281,"duration_ms":15537,"temperature":1.0,"reasoning_tokens":2196,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:25:06.570179+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison on the same recordings with ground-truth f0 established by glottal inverse filtering or simultaneous high-speed videoendoscopy or electroglottography, and check whether FCN-F0's 96% accuracy and its 34% subharmonic-error reduction over CREPE persist; a second check would be testing the ranking on running speech with subharmonics, since the paper attributes its disagreement with [27] to the difference between sustained vowels and connected speech.","supporting_citations":[{"cited_title":"Fully-convolutional net- work for pitch estimation of speech signals,","cited_arxiv_id":null,"evidence_quote":"Supplies the FCN-F0 fully convolutional network architecture and pretrained weights, the estimator that performed best."},{"cited_title":"Harvest: A high-performance fundamental frequency estimator from speech signals,","cited_arxiv_id":null,"evidence_quote":"Describes the Harvest estimator, the leading non-data-driven estimator compared and the closest competitor to the deep-learning models."},{"cited_title":"Accurate short-term analysis of the funda- mental frequency and the harmonics-to-noise ratio of a sampled sound,","cited_arxiv_id":null,"evidence_quote":"Describes the Praat f0 detector, one of the compared estimators and the basis of the autocorrelation baseline."},{"cited_title":"A spectral/temporal method for robust fundamental frequency tracking,","cited_arxiv_id":null,"evidence_quote":"Describes the YAAPT estimator, compared for its reported ability to handle subharmonics."},{"cited_title":"Performance analysis of various fundamental frequency estimation algorithms in the context of pathological speech,","cited_arxiv_id":null,"evidence_quote":"Prior comparison reporting YAAPT and FCN-F0 as subharmonic-capable; this study confirms and partially contradicts those findings."},{"cited_title":"Harmonics-to-noise ratio estimation with deterministically time-varying harmonic model for patho- logical voice signals,","cited_arxiv_id":null,"evidence_quote":"Provides the time-varying harmonic model with gradient-based optimization used to refine the manual f0 annotations."}],"review_version":1}