REVIEW 3 major objections 5 minor 1 cited by
Comparison of fundamental frequency estimators with subharmonic voice signals
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read FCN-F0, a deep-learning pitch tracker, is most accurate on subharmonic voice signals.
desk verdict A solid empirical comparison of pitch estimators on subharmonic voices, but the headline accuracy gaps lack cluster-aware statistics and should be read as trends, not a settled ranking. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The study's key machinery is a quality-of-estimate classification built on the harmonic power profile $P(f_o)=\sum_{k=1}^{K(f_o)}|S_{xx}(k f_o)|^2$, where $S_{xx}$ is the periodogram of the 50-ms Hamming-windowed signal and $K(f_o)=\lfloor f_s/(2f_o)\rfloor$. For each annotated truth $f_o^*$, the profile's local minima define intervals around $f_o^*/M$; an estimate landing in the $M=1$ interval is correct, in an $M>1$ interval is a subharmonic error, and elsewhere is another error. Subharmonic strength is measured by the subharmonics-to-harmonics ratio (SHR), computed from the autocorrelation baseline's estimates, which quantifies the power of subharmonic tones relative to harmonic tones. These tools separate incidental autocorrelation errors (the main SHR peak near $-25$ dB) from true subharmonic intervals (secondary peak near $-10$ dB) and attribute each estimator's failures.
What would settle it
Re-run the comparison on the same recordings with ground-truth f0 established by glottal inverse filtering or simultaneous high-speed videoendoscopy or electroglottography, and check whether FCN-F0's 96% accuracy and its 34% subharmonic-error reduction over CREPE persist; a second check would be testing the ranking on running speech with subharmonics, since the paper attributes its disagreement with [27] to the difference between sustained vowels and connected speech.
Extended reading notes
Core claim
The central discovery is that FCN-F0, a deep-learning model, performs the best both in overall accuracy and in correctly resolving subharmonic signals: it achieved 96% correct estimates versus 95% for CREPE and Harvest, 88% for Praat and YAAPT, and 62% for the autocorrelation baseline. FCN-F0 reduced subharmonic errors by 34% compared with CREPE, and its error profile was the most balanced between subharmonic and other mistakes. The paper further establishes a reliability boundary: among the 755 intervals with subharmonics-to-harmonics ratio above $-10$ dB, FCN-F0 correctly estimated only 63.7%, and none of the estimators could reliably handle strong subharmonics with SHR above $-3$ dB.
Load-bearing premise
The manual annotation of the true speaking fundamental frequency is itself uncertain for strong subharmonics, where a female voice with strong subharmonics could be mistaken for a normal male voice; all accuracy numbers rest on this ground truth.
Editorial extensions
If this is right
- Clinical acoustic analysis of sustained vowels should prefer FCN-F0, with CREPE and Harvest as strong alternatives, to reduce false negatives caused by subharmonic voicing.
- For intervals with SHR above $-3$ dB, no current estimator should be trusted; acoustic parameters computed there should be flagged as unreliable.
- Praat's Viterbi postprocessing fixes incidental autocorrelation errors but not true subharmonic intervals, so its subharmonic error rate of about 9.3% is a lower bound for the fraction of subharmonic intervals in the dataset.
- Retraining deep-learning models with subharmonic voice samples, as the paper suggests, may improve accuracy for the high-SHR cases that currently defeat all estimators.
Reading between the lines
- The paper leaves implicit that the deep-learning models' advantage may come from using features beyond raw periodicity (such as harmonic amplitudes and phases), which would predict that they generalize less well to languages or recording conditions absent from their training corpora.
- The 8-kHz resampling used in the study could limit estimates for very high-pitched voices; a testable extension is whether a higher sampling rate changes the ranking among the top three estimators.
- The SHR-based classification could be turned into a practical clinical flag: intervals with ACF subharmonic errors and high SHR are likely true subharmonics, so a confidence mask accompanying F0 estimates could improve downstream diagnostics.
- Because FCN-F0's training data are non-pathological English and French speech, a broader pathological corpus spanning more voice types and languages is a natural next test of whether its 96% accuracy is robust.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper compares six fundamental-frequency estimators—an autocorrelation (ACF) baseline, Praat, YAAPT, Harvest, CREPE, and FCN-F0—on 50-ms intervals of sustained /a/ vowels from the KayPENTAX Disordered Voice Database. The authors manually annotate a speaking fundamental frequency f_o^* for each interval, classify each estimate as correct, subharmonic error, or other error using intervals of the harmonic power profile P(f_o), and compute a subharmonics-to-harmonics ratio (SHR) from ACF-based subharmonic candidates. The reported headline result is that FCN-F0 has the highest per-interval accuracy (96% vs. 95% for CREPE and Harvest, 88% for Praat and YAAPT, 62% for ACF), reduces subharmonic errors by 34% relative to CREPE, and no estimator reliably handles strong subharmonics (SHR > -3 dB).
Significance. If the accuracy ranking were established, the paper would be a useful clinical benchmark: it evaluates per-frame behavior rather than per-recording averages, uses publicly available estimator implementations, and stratifies by subharmonic strength. The SHR-conditional analyses and the contingency tables against the ACF baseline are informative and go beyond simple overall accuracy. The central ranking, however, is currently supported only by pooled per-interval rates without uncertainty quantification or cluster-aware inference, and the manual ground truth is acknowledged to be uncertain precisely in the subharmonic cases that matter. These issues are fixable with additional analysis, and the clinical application is potentially valuable for voice assessment.
major comments (3)
- [Section III, Figs. 4-6] The central claim that FCN-F0 'performed the best' rests on pooled per-interval percentages computed over 15,941 intervals from only 703 recordings, with no confidence intervals, paired tests, or adjustment for within-recording correlation. Sustained-vowel intervals from the same recording are strongly non-independent; a small number of pathological recordings with sustained strong subharmonics can contribute large blocks of correlated errors. The 96%-vs-95% accuracy gap between FCN-F0 and CREPE/Harvest corresponds to roughly 160 intervals, and the claimed 34% subharmonic-error reduction is about a one-percentage-point absolute difference—both are within the range that a few recordings could plausibly account for. I request cluster-robust inference (e.g., bootstrap by recording or a mixed-effects logistic regression with a recording random effect), with 95% confidence intervals reported for all accuracy and subharmonic-error differences. Without this, the headline ranking is not established at the precision implied by the text.
- [Section II, 'fo Annotation' and the Limitations paragraph] The accuracy numbers depend entirely on the manually annotated f_o^*, which the authors acknowledge is uncertain for sustained strong subharmonics (e.g., a female voice with strong subharmonics could be mistaken for a normal male voice). Because the top three estimators differ by only about one percentage point, even a small number of annotation errors could reorder CREPE, Harvest, and FCN-F0. Please add a sensitivity analysis: exclude or relabel the intervals flagged as uncertain, quantify annotator agreement (e.g., a second annotation pass or test-retest reliability), or use an objective reference on a subset. The current statement that errors 'could slightly change' the accuracies is not sufficient to support the precision of the claimed ranking.
- [Section II, Quality-of-Estimate Classification] The rule that labels an estimate correct or as a subharmonic error based on whether it falls in an interval bounded by local minima of P(f_o) is novel and central to all error rates in Figs. 4-6, but no validation of this classification is provided. If the interval boundaries are sensitive to the periodogram's 0.5-Hz resolution, the Hamming window, or the choice of local-minimum search, the error counts could shift. I ask for at least a robustness check on a subset of intervals (e.g., comparison with a perceptual or manual classification) so that the reader can see that the quality labels are not driving the ranking.
minor comments (5)
- [Section III, Fig. 4] The percentages in the bar chart would be easier to interpret alongside a table reporting exact counts, percentages, and confidence intervals for every estimator and error category.
- [Section III, paragraph beginning 'The types of the estimation errors...'] The sentence 'reduces the subharmonic errors (34% less than CREPE)' is ambiguous without the absolute rates; please state both rates explicitly (e.g., CREPE 3.0% vs. FCN-F0 2.0%).
- [Section II, Acoustic Data] The decision to resample all signals to 8 kHz is stated, but the rationale is not; since CREPE is then run at 16 kHz, please explain whether the 8-kHz lowpass filtering has any effect on the upper range of f_o or on the harmonic/subharmonic analysis.
- [Section III, 'There is also a notable discrepancy...'] The comparison with Vaysse et al. is interesting but would be more useful with a direct statement of which recording types and annotation conventions differ, since the reader cannot evaluate the claim from the cited abstract alone.
- [General] The paper would benefit from a data/code availability statement, including access to the custom spectrogram annotation program and the scripts used for the estimators, to support reproducibility of the 15,941-interval annotations and the contingency-table analyses.
Circularity Check
No circularity found: the study is an empirical comparison with manually annotated ground truth, and the central ranking is not forced by construction or by the authors' prior work.
full rationale
This paper performs an empirical benchmark of five fundamental-frequency estimators (plus an ACF baseline) on sustained vowels. The central claim is that FCN-F0 achieves the highest accuracy and the best subharmonic-error handling. I walked the derivation chain and found no step that reduces to its own input. The ground-truth annotation uses Praat only as an initial rough estimate that is then manually reviewed on a spectrogram with playback and refined by a time-varying harmonic model; the final truth is not defined as the output of any evaluated estimator. This is further confirmed empirically: Praat itself scores only 88% correct while the two deep-learning models score 95-96%, so the annotation procedure did not simply re-import Praat's outputs. The subharmonic-error classification is based on the annotated truth f_o* and a harmonic power profile P(fo), not on any estimator's self-derived quantity. SHR is computed from the ACF baseline and used only as an explanatory variable to stratify subharmonic strength; it is not used to 'predict' the compared estimators' outcomes in a way that is equivalent to fitting them. The authors cite several of their own previous papers, but these citations support background claims (vocal bifurcation phenomena, harmonic-model refinement, kymographic analysis) rather than the ranking result. No uniqueness theorem or ansatz is imported from those citations to force the choice of FCN-F0. The paper also candidly discusses annotation uncertainty in the Limitations paragraph, which affects precision but is not a circularity. Therefore no circular step can be quoted, and the score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The manually annotated fundamental frequency f_o^* correctly represents the speaking fundamental frequency for every interval, including cases with strong subharmonics.
- ad hoc to paper The quality-of-estimate classification, using intervals bounded by local minima of the harmonic power profile P(f_o), correctly labels estimates as correct or subharmonic errors.
Cite this review
Pith. "Pith review of Comparison of fundamental frequency estimators with subharmonic voice signals." pith.science (2026). https://pith.science/paper/DNDT56XR
@misc{pith2026250104789,
author = {Pith},
title = {Pith review of: Comparison of fundamental frequency estimators with subharmonic voice signals},
year = {2026},
howpublished = {\url{https://pith.science/paper/DNDT56XR}},
note = {Machine review of arXiv:2501.04789}
}
read the original abstract
In clinical voice signal analysis, mishandling of subharmonic voicing may cause an acoustic parameter to signal false negatives. As such, the ability of a fundamental frequency estimator to identify speaking fundamental frequency is critical. This paper presents a sustained-vowel study, which used a quality-of-estimate classification to identify subharmonic errors and subharmonics-to-harmonics ratio (SHR) to measure the strength of subharmonic voicing. Five estimators were studied with a sustained vowel dataset: Praat, YAAPT, Harvest, CREPE, and FCN-F0. FCN-F0, a deep-learning model, performed the best both in overall accuracy and in correctly resolving subharmonic signals. CREPE and Harvest are also highly capable estimators for sustained vowel analysis.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
Fully convolutional networks trained on synthesized subharmonic phonation classify subharmonic period M=1-4 with >98% synthetic accuracy, with qualitative only evidence on real voices.
Reference graph
Works this paper leans on
-
[1]
Microphone and electroglottographic data from dyspho- nic patients: Type 1, 2 and 3 signals,
A. Behrman, C. J. Agresti, E. Blumstein, and N. Lee, “Microphone and electroglottographic data from dyspho- nic patients: Type 1, 2 and 3 signals,” J. V oice, vol. 12, no. 2, pp. 249–260, Jan. 1998
1998
-
[2]
Diplophonia reappraised,
L. Cavalli and A. Hirson, “Diplophonia reappraised,” J. V oice, vol. 13, no. 4, pp. 542–556, 1999
1999
-
[3]
A study of subharmonics in connected speech material,
E. Kramer, R. Linder, and R. Schönweiler, “A study of subharmonics in connected speech material,” Journal of V oice, vol. 27, no. 1, pp. 29–38, Jan. 2013
work page 2013
-
[4]
T. Ikuma, A. J. McWhorter, L. Adkins, and M. Kunduk, “Investigation of vocal bifurcations and voice patterns induced by asymmetry of pathological vocal folds,” J. Speech. Lang. Hear . Res., vol. 66, no. 1, pp. 48–60, Jan. 2023
work page 2023
-
[5]
Acoustic characteristics of rough voice: Subharmonics,
K. Omori, H. Kojima, R. Kakani, D. H. Slavit, and S. M. Blaugrund, “Acoustic characteristics of rough voice: Subharmonics,” J. V oice, vol. 11, no. 1, pp. 40–47, Mar. 1997
1997
-
[6]
I. R. Titze, Workshop on Acoustic V oice Analysis: Sum- mary Statement. Denver, CO, USA: National Center for V oice and Speech, 1994
work page 1994
-
[7]
S. Kiritani, H. Hirose, and H. Imagawa, “High-speed digital image analysis of vocal cord vibration in diplo- IKUMA ET AL. - JAN. 2025 8 phonia,” Speech Commun. , vol. 13, no. 1-2, pp. 23–32, 1993
work page 2025
-
[8]
The mechanisms of subharmonic tone generation in a synthetic larynx model,
S. Kniesburges, A. Lodermeyer, S. Becker, M. Traxdorf, and M. Döllinger, “The mechanisms of subharmonic tone generation in a synthetic larynx model,” J. Acoust. Soc. Am., vol. 139, no. 6, pp. 3182–3192, Jun. 2016
2016
Show all 43 references
-
[9]
Synthetic multi-line kymographic analysis: A spatiotem- poral data reduction technique for high-speed videoen- doscopy,
T. Ikuma, M. Kunduk, D. Fink, and A. J. McWhorter, “Synthetic multi-line kymographic analysis: A spatiotem- poral data reduction technique for high-speed videoen- doscopy,” J. Acoust. Soc. Am. , vol. 140, no. 4, pp. 2703– 2713, Oct. 2016
2016
-
[10]
Irregular vocal- fold vibration—High-speed observation and modeling,
P. Mergell, H. Herzel, and I. R. Titze, “Irregular vocal- fold vibration—High-speed observation and modeling,” J. Acoust. Soc. Am., vol. 108, no. 6, pp. 2996–3002, 2000
2000
-
[11]
Spatio-temporal analysis of irregular vocal fold oscil- lations: Biphonation due to desynchronization of spatial modes,
J. Neubauer, P. Mergell, U. Eysholdt, and H. Herzel, “Spatio-temporal analysis of irregular vocal fold oscil- lations: Biphonation due to desynchronization of spatial modes,” J. Acoust. Soc. Am. , vol. 110, no. 6, pp. 3179– 3192, 2001
2001
-
[12]
Simulation of multiple source vocalization in the larynx: How true folds, false folds, and aryepiglottic folds may interact,
I. R. Titze, “Simulation of multiple source vocalization in the larynx: How true folds, false folds, and aryepiglottic folds may interact,” J Speech Lang Hear Res , vol. 67, no. 3, pp. 802–810, Mar. 2024
2024
-
[13]
Multi-Dimensional V oice Program (MDVP) Model 5105 Software Instruction Manual,
KayPENTAX, “Multi-Dimensional V oice Program (MDVP) Model 5105 Software Instruction Manual,” Lincoln Park, NJ, Jun. 2008
2008
-
[14]
On the nature of vocal fry,
H. Hollien, P. Moore, R. W. Wendahl, and J. F. Michel, “On the nature of vocal fry,” Journal of Speech and Hearing Research, vol. 9, no. 2, pp. 245–247, Jun. 1966
1966
-
[15]
Freddie Mercury—acoustic analysis of speaking fundamental frequency, vibrato, and subhar- monics,
C. T. Herbst, S. Hertegard, D. Zangger-Borch, and P.- Å. Lindestad, “Freddie Mercury—acoustic analysis of speaking fundamental frequency, vibrato, and subhar- monics,” Logoped. Phoniatr . V ocol., vol. 42, no. 1, pp. 29–38, Jan. 2017
2017
-
[16]
Pitch period determination of aperiodic speech signals,
P. Hedelin and D. Huber, “Pitch period determination of aperiodic speech signals,” in Int. Conf. Acoust. Speech Signal Process., Apr. 1990, pp. 361–364 vol.1
1990
-
[17]
A pitch determination algorithm based on subharmonic-to-harmonic ratio,
X. Sun, “A pitch determination algorithm based on subharmonic-to-harmonic ratio,” in Proc. 6th ICSLP , vol. 4, Beijing, China, 2000, pp. 676–679
2000
-
[18]
Acoustic tracking of pitch, modal, and subhar- monic vibrations of vocal folds in Parkinson’s Disease and Parkinsonism,
J. Hlavni ˇcka, R. ˇCmejla, J. Klempí ˇr, E. R˚ užiˇcka, and J. Rusz, “Acoustic tracking of pitch, modal, and subhar- monic vibrations of vocal folds in Parkinson’s Disease and Parkinsonism,” IEEE Access , vol. 7, pp. 150 339– 150 354, 2019
2019
-
[19]
Poincaré pitch marks,
M. Hagmüller and G. Kubin, “Poincaré pitch marks,” Speech Communication, vol. 48, no. 12, pp. 1650–1665, Dec. 2006
2006
-
[20]
Measurement of fundamental frequencies in diplophonic voices,
P. Aichinger, M. Hagmüller, I. Roesner, W. Bigenzahn, B. Schneider-Stickler, J. Schoentgen, and F. Pernkopf, “Measurement of fundamental frequencies in diplophonic voices,” in Proc. 13th MA VEBA. Florence, Italy: Firenze University Press, 2015, Sept. 2-4, pp. 21–24
2015
-
[21]
Fundamental frequency tracking in diplophonic voices,
P. Aichinger, M. Hagmüller, I. Roesner, B. Schneider- Stickler, J. Schoentgen, and F. Pernkopf, “Fundamental frequency tracking in diplophonic voices,” Biomed. Sig- nal Process. Control , vol. 37, pp. 69–81, Aug. 2017
2017
-
[22]
Tracking of mul- tiple fundamental frequencies in diplophonic voices,
P. Aichinger, M. Hagmüller, B. Schneider-Stickler, J. Schoentgen, and F. Pernkopf, “Tracking of mul- tiple fundamental frequencies in diplophonic voices,” IEEEACM Trans. Audio Speech Lang. Process. , vol. 26, no. 2, pp. 330–341, Feb. 2018
2018
-
[23]
Comparative per- formance of pitch detection algorithms on dysphonic voices,
J. Laver, S. Hiller, and R. Hanson, “Comparative per- formance of pitch detection algorithms on dysphonic voices,” in ICASSP 82 IEEE Int. Conf. Acoust. Speech Signal Process., vol. 7, May 1982, pp. 192–195
1982
-
[24]
A comparison of high precision FO extraction algorithms for sustained vowels,
V . Parsa and D. G. Jamieson, “A comparison of high precision FO extraction algorithms for sustained vowels,” J. Speech. Lang. Hear . Res., vol. 42, no. 1, pp. 112–126, Feb. 1999
1999
-
[25]
Evaluation of performance of several established pitch detection algorithms in pathological voices,
S.-J. Jang, S.-H. Choi, H.-M. Kim, H.-S. Choi, and Y .-R. Yoon, “Evaluation of performance of several established pitch detection algorithms in pathological voices,” in 2007 29th Annu. Int. Conf. IEEE Eng. Med. Biol. Soc. , Lyon, France, Aug. 2007, pp. 620–623
2007
-
[26]
Robust fundamental frequency es- timation in sustained vowels: Detailed algorithmic com- parisons and information fusion with adaptive Kalman filtering,
A. Tsanas, M. Zañartu, M. A. Little, C. Fox, L. O. Ramig, and G. D. Clifford, “Robust fundamental frequency es- timation in sustained vowels: Detailed algorithmic com- parisons and information fusion with adaptive Kalman filtering,” J. Acoust. Soc. Am. , vol. 135, no. 5, pp. 2...
2014
-
[27]
Performance analysis of various fundamental frequency estimation algorithms in the context of pathological speech,
R. Vaysse, C. Astésano, and J. Farinas, “Performance analysis of various fundamental frequency estimation algorithms in the context of pathological speech,” J. Acoust. Soc. Am. , vol. 152, no. 5, pp. 3091–3101, 2022
2022
-
[28]
An analysis of the diplo- phonia phenomenon,
P. Dejonckere and J. Lebacq, “An analysis of the diplo- phonia phenomenon,” Speech Commun., vol. 2, no. 1, pp. 47–56, May 1983
1983
-
[29]
Disordered V oice Database and Program [Model 4337],
KayPENTAX and Massachusetts Eye and Ear Infirmary, “Disordered V oice Database and Program [Model 4337],” 2006
2006
-
[30]
Harmonics-to-noise ratio estimation with deterministically time-varying harmonic model for patho- logical voice signals,
T. Ikuma, B. Story, A. J. McWhorter, L. Adkins, and M. Kunduk, “Harmonics-to-noise ratio estimation with deterministically time-varying harmonic model for patho- logical voice signals,” J. Acoust. Soc. Am., vol. 152, no. 3, pp. 1783–1794, Sep. 2022
2022
-
[31]
R. J. Baken and R. F. Orlikoff, Clinical Measurement of Speech and V oice , 2nd ed. San Diego, CA, USA: Singular, 2000
2000
-
[32]
Accurate short-term analysis of the funda- mental frequency and the harmonics-to-noise ratio of a sampled sound,
P. Boersma, “Accurate short-term analysis of the funda- mental frequency and the harmonics-to-noise ratio of a sampled sound,” Proc. Inst. Phonet. Sci. , vol. 17, pp. 97– 110, 1993
1993
-
[33]
Harvest: A high-performance fundamental frequency estimator from speech signals,
M. Morise, “Harvest: A high-performance fundamental frequency estimator from speech signals,” in Interspeech
-
[34]
A spectral/temporal method for robust fundamental frequency tracking,
S. A. Zahorian and H. Hu, “A spectral/temporal method for robust fundamental frequency tracking,” J. Acoust. Soc. Am. , vol. 123, no. 6, pp. 4559–4571, Jun. 2008
2008
-
[35]
Crepe: A convolutional representation for pitch estimation,
J. W. Kim, J. Salamon, P. Li, and J. P. Bello, “Crepe: A convolutional representation for pitch estimation,” in IEEE ICASSP 2018 , Calgary, AB, Apr. 2018, pp. 161– 165
2018
-
[36]
Fully-convolutional net- work for pitch estimation of speech signals,
L. Ardaillon and A. Roebel, “Fully-convolutional net- work for pitch estimation of speech signals,” in Inter- speech 2019 , Sep. 2019, pp. 2005–2009. IKUMA ET AL. - JAN. 2025 9
2019
-
[37]
Nearly defect-free F0 trajectory extraction for expressive speech modifications based on STRAIGHT,
H. Kawahara, A. D. Cheveigné, H. Banno, T. Taka- hashi, and T. Irino, “Nearly defect-free F0 trajectory extraction for expressive speech modifications based on STRAIGHT,” in Interspeech 2005 . ISCA, Sep. 2005, pp. 537–540
2005
-
[38]
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm,
A. Viterbi, “Error bounds for convolutional codes and an asymptotically optimum decoding algorithm,” IEEE Trans. Inf. Theory, vol. 13, no. 2, pp. 260–269, Apr. 1967
1967
-
[39]
Introducing Parselmouth: A Python interface to Praat,
Y . Jadoul, B. Thompson, and B. de Boer, “Introducing Parselmouth: A Python interface to Praat,” Journal of Phonetics, vol. 71, pp. 1–15, Nov. 2018
2018
-
[40]
WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Trans. Inf. & Syst. , vol. E99.D, no. 7, pp. 1877–1884, 2016
2016
-
[41]
Performance evaluation of subharmonic- to-harmonic ratio (SHR) computation,
C. T. Herbst, “Performance evaluation of subharmonic- to-harmonic ratio (SHR) computation,” J. V oice, vol. 35, no. 3, pp. 365–375, May 2021
2021
-
[42]
Perception of pitch and roughness in vocal signals with subharmonics,
C. C. Bergan and I. R. Titze, “Perception of pitch and roughness in vocal signals with subharmonics,” J. V oice, vol. 15, no. 2, pp. 165–175, Jun. 2001
2001
-
[2017]
2017, pp
ISCA, Aug. 2017, pp. 2321–2325
2017
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.