REVIEW 4 major objections 6 minor 64 references
Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Co-articulated vowel spectra progress by interpolation between sustained-vowel templates, and a glottal-pulse synthesis method matches frequency-domain quality in perceptual tests.
desk verdict A careful, honest engineering paper with a genuinely new phase-based pitch-pulse analysis and solid listening tests, but the template-interpolation idea that motivates the whisper-restoration use case is not actually tested by the synthesis experiments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key object is the phase-oriented pitch-pulse segmentation, which locates each pitch period by searching for the time shift at which the DFT phase of the fundamental-frequency bin equals $-\pi/2$ (the onset of $\phi_0$), then derives per-period harmonic phases $\phi_\ell=\angle X_\ell[1+\ell]+\pi/2$ and the Normalized Relative Delay vector $\mathrm{NRD}_\ell=(\phi_\ell-(\ell+1)\phi_0)/2\pi$, taken modulo 1 and unwrapped. This yields a per-pulse view of harmonic magnitude and phase that frame-based analysis cannot provide. The comparison of voicing methods rests on three synthesis architectures; the one carrying the new claim is GLO, which generates each glottal excitation pulse individually with a Liljencrants-Fant glottal pulse model (a standard parametric shape for the airflow through the glottis), filters it with a dedicated all-pole vocal-tract filter, and overlap-adds the filtered pulses, keeping the glottal source explicitly separate from the tract response.
What would settle it
Record co-articulated vowel transitions from many speakers and languages, segment them with the phase-based pitch-pulse method, and compare each intermediate pulse's NRD vector and magnitude spectrum against the interpolation of that speaker's sustained-vowel templates. If the deviation between measured and interpolated spectra exceeds the natural period-to-period scatter seen within sustained vowels, the interpolation model is falsified.
Extended reading notes
Core claim
The paper's central claim is that the harmonic magnitude/phase structure across successive pitch pulses in a co-articulated vowel region can be modeled as a progressive interpolation between the natural, idiosyncratic spectral profiles of the same vowels produced in sustained mode, rather than as an arbitrary sequence of frames. The evidence comes from phase-based segmentation of the /i/–/a/ transition in the Portuguese word 'Tiago' for two speakers: the Normalized Relative Delay vectors and magnitude spectra move from the /i/ template toward the /a/ template over about 24 pitch periods. The perceptual claim is that, for co-articulated vowels in word context, GLO—which synthesizes individual glottal excitation pulses and filters each through a vocal-tract model—is rated at least as natural as frequency-domain synthesis and better than combined frequency/time synthesis, while for sustained vowels the frequency-domain method remains the strongest.
Load-bearing premise
The interpolation result is inferred from one vowel transition (/i/ to /a/ in the Portuguese word 'Tiago') produced by two speakers, plus a three-vowel 'Saiu' sequence used only in listening tests; the paper's synthesis strategy for running whispered speech assumes this pattern holds across speakers, vowel pairs, and languages.
Editorial extensions
If this is right
- An entire co-articulated vowel region can be synthesized by pitch-pulse-by-pitch-pulse interpolation between stored spectral templates of the same speaker, so no per-frame formant tracking is needed.
- For running whispered speech, the physiologically inspired GLO method is the better candidate among the tested options, since it matched the frequency-domain method and beat the combined frequency/time method on co-articulated vowels.
- For sustained vowels, the frequency-domain method remains a very efficient high-quality choice, statistically indistinguishable from the original in several listening conditions.
- Because GLO explicitly separates glottal excitation from vocal-tract filtering, voice identity and phonation type can in principle be controlled independently of articulation.
- The phase-based pitch-pulse segmentation can also serve as a monitoring tool, since its f0 contours agree with frame-based pitch estimates and it can verify correct operation of synthetic voicing algorithms.
Reading between the lines
- If the interpolation property generalizes, a whisper-restoration device could be built from a small per-speaker dictionary of sustained-vowel templates, avoiding the large paired datasets that data-driven whisper-to-speech systems require; the paper aims at this but does not test it.
- A direct extension would stress-test the model on vowel pairs with very different formant spacing (e.g., /i/–/u/), across languages and speech rates, to see whether co-articulation always tracks the sustained templates or sometimes takes a shortcut between articulatory targets.
- The observation that GLO's harmonic structure can be cleaner than the original suggests a clinical use beyond restoration: the same segmentation and synthesis machinery could quantify harmonic damage in dysphonic voices, which the paper mentions only in passing.
- Replacing the idealized Liljencrants-Fant glottal model with pulses derived from physiological recordings is named as future work in the paper; the explicit source/filter separation in GLO makes that replacement straightforward and directly testable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses two interconnected challenges in whispered-speech-to-voiced-speech restoration: (i) a new phase-based procedure for segmenting individual pitch pulses in sustained and co-articulated vowels and extracting their per-period harmonic magnitude and NRD phase structure, and (ii) three lightweight parametric synthetic-voicing alternatives (frequency-domain FRE, combined frequency/time-domain TIM, and physiologically inspired time-domain GLO), compared through spectrograms and listening tests following ITU-R BS.1116. The authors report that the harmonic magnitude/phase structure of co-articulated vowels evolves as a progressive interpolation between the structures of the same vowels produced in sustained mode, and that GLO is perceptually comparable to FRE for co-articulated vowels in word context. The paper includes public audio files of the 32 test stimuli and derives the DFT phase mathematics in detail.
Significance. If the interpolation result were established across speakers, vowel pairs, and languages, it would provide an interpretable template-based pathway for lightweight, real-time, on-the-fly whisper-to-speech restoration, which is the paper's stated goal. The work has clear strengths: the DFT phase analysis in Section 2.1 is formally derived, the listening-test protocol follows ITU-R BS.1116 with 30 listeners and randomized hidden references, the f0-contour comparison in Figure 2 provides a useful sanity check of the segmentation procedure, and the authors make the test audio files publicly available. The comparative synthesis evaluation is a useful contribution even though, as detailed below, it does not exercise the template-interpolation pathway. The significance is currently limited by the narrow evidential basis for the interpolation claim and by the mismatch between the stated synthesis strategy and the actual synthesis experiments.
major comments (4)
- [Section 2.3, Insight 3; Figs. 4 and 5] The progressive-interpolation claim is supported by only a single /i/-/a/ co-articulation in the Portuguese word 'Tiago' for each of two speakers, and the evidence is presented only as overlaid curves. No quantitative distance metric compares the observed per-period NRD/magnitude evolution against the interpolation model, and no cross-validation across other vowel pairs, speakers, or languages is provided. Because the proposed whisper-restoration strategy depends on this assumption, the claim is currently under-supported and needs either a quantitative test on the existing data or additional data.
- [Section 3.1, Eq. (10); Section 4.3] The synthesis evaluation does not implement the proposed template-interpolation pathway. In the FRE/TIM/GLO versions, each word is reconstructed using parameters extracted from the original co-articulated recording: Eq. (10) interpolates between frame-boundary parameters of the original signal, not between endpoint templates of sustained vowels. The listening tests in Tables 5-8 therefore validate reconstruction of a word from its own analysis parameters, not synthesis of co-articulation from vowel templates. Since the stated whisper-restoration strategy requires the latter, the central synthesis claim is not validated by the experiments as designed.
- [Section 2.1] The phase-based pitch-period segmentation procedure is not validated against any ground truth, such as EGG signals, manual epoch marks, or synthetic signals with known period boundaries. The consistency of overlays in Figure 3 and the agreement between 'ORI' and 'MOD' f0 contours in Figure 2 are indirect evidence; without a quantitative accuracy measure, the per-period NRD/magnitude evolution that underlies the interpolation claim is not established on firm footing.
- [Section 4.3, Tables 1-8] The listening-test analysis relies on many pairwise two-tailed t-tests without correction for multiple comparisons, and the absence of statistically significant differences is repeatedly interpreted as evidence of equivalence (e.g., FRE vs. GLO in 'Tiago' and 'Saiu'). An ANOVA or a mixed-effects model with appropriate post-hoc corrections, or explicit equivalence testing, is needed to support the comparative conclusions drawn from these data.
minor comments (6)
- [Section 4.1] The word 'representated' should be 'represented'.
- [Section 4.3, bullet list] The phrase 'the the combined frequency and time-domain voicing' contains a duplicated article.
- [Section 5] The phrase 'Is general' should be 'In general'.
- [Section 1.2] The phrase 'more that finding' should be 'more than finding'.
- [Section 3.3, Eq. (10)] The notation in Eq. (10) does not explicitly define the range of the harmonic index L or the normalization of NRD in the synthesis context; please state these definitions for reproducibility.
- [Figure 4 caption] The caption identifies the dashed magenta line as the sustained /a/ reference and the dashed blue line as the sustained /i/ reference, but the order is reversed relative to the text description in Section 2.2; please clarify the mapping.
Circularity Check
No significant circularity: the central interpolation result is an empirical observation, and the synthesis conclusions rest on external listening tests rather than on a fitted-input prediction.
full rationale
The paper's load-bearing claims are (i) that the harmonic phase/magnitude structure of co-articulated vowels can be viewed as a progressive interpolation between vowel-specific templates (Section 2.3, Insight 3), and (ii) that the three voicing-restoration methods, especially FRE and GLO, deliver perceptually acceptable quality (Section 4.3). Neither claim reduces by construction to its inputs. The interpolation insight is supported by direct pitch-period analysis of natural recordings (Figs. 4-5) and is presented as an empirical interpretation, not as a quantity forced by a fitted parameter; the later synthesis tests (Eq. 10 and Section 4.2) reconstruct each word from parameters extracted from that same word, so they do not constitute a template-interpolation prediction, but that is a scope limitation rather than circularity. The NRD/LPC/ODFT analysis tools and the FRE/TIM synthesis blocks are reused from the authors' earlier work [18,19,51], with self-citations present and normal; however, the pitch-pulse segmentation is benchmarked against a different frame-based pitch estimator (Fig. 2), and the GLO method deliberately omits NRD, giving an independent synthesis path. The comparative quality claims are anchored in ITU-R BS.1116 listening tests with 30 listeners and hidden references (Section 4.3), which are external perceptual evidence not constructed from the paper's own parameters. Accordingly, the central results have independent content, and no specific equation or fitted parameter is renamed as a prediction. Score 2 reflects only the presence of expected self-citations, none of which is load-bearing.
Assumptions & free parameters
assumptions (5)
- domain assumption The source-filter model of speech production with independent glottal source and vocal tract filter is valid.
- standard math DFT analysis of a P-point pitch period locates the f0 onset when the phase of bin k=1 is approximately -pi/2, so per-period phase segmentation is well-posed.
- domain assumption Unvoiced sounds in whispered speech do not differ appreciably from those in natural speech for the same speaker.
- domain assumption The NRD feature, combined with harmonic magnitudes, completely determines waveform shape independently of f0, time-shift, and overall magnitude.
- ad hoc to paper Co-articulation can be modeled as a progressive interpolation between sustained-vowel spectral templates.
Cite this review
Pith. "Pith review of Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods." pith.science (2026). https://pith.science/paper/KKOBXI42
@misc{pith2026250606675,
author = {Pith},
title = {Pith review of: Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/KKOBXI42}},
note = {Machine review of arXiv:2506.06675}
}
read the original abstract
Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal components that exist in certain regions of the former, called voiced regions, as a consequence of the vibration of the vocal folds, are missing in the latter. The restoration of natural speech from whispered speech requires delicate signal processing procedures that are especially useful if they can be implemented on low-resourced portable devices, in real-time, and on-the-fly, taking advantage of the established source-filter paradigm of voice production and related models. This paper addresses two challenges that are intertwined and are key in informing and making viable this envisioned technological realization. The first challenge involves characterizing and modeling the evolution of the harmonic phase/magnitude structure of a sequence of individual pitch periods in a voiced region of natural speech comprising sustained or co-articulated vowels. This paper proposes a novel algorithm segmenting individual pitch pulses, which is then used to obtain illustrative results highlighting important differences between sustained and co-articulated vowels, and suggesting practical synthetic voicing approaches. The second challenge involves model-based synthetic voicing. Three implementation alternatives are described that differ in their signal reconstruction approaches: frequency-domain, combined frequency and time-domain, and physiologically-inspired separate filtering of glottal excitation pulses individually generated. The three alternatives are compared objectively using illustrative examples, and subjectively using the results of listening tests involving synthetic voicing of sustained and co-articulated vowels in word context.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Fant, Acoustic Theory of Speech Production, The Hague , 1970
G. Fant, Acoustic Theory of Speech Production, The Hague , 1970
work page 1970
-
[2]
O’Shaughnessy, Speech Communications: Human and Mac hine, Wiley-IEEE Press, 2000
D. O’Shaughnessy, Speech Communications: Human and Mac hine, Wiley-IEEE Press, 2000
work page 2000
-
[4]
L. Rabiner, B.-H. Juang, Fundamentals of Speech Recogni tion, Prentice- Hall, Inc., 1993
work page 1993
-
[5]
A. V. Oppenheim, A. S. Willsky, S. Hamid, Signals and Syst ems, Pear- son Education Limited, 1996, 2nd Ed
work page 1996
-
[6]
A. Ferreira, On the possibility of speaker discriminati on using a glot- tal pulse phase-related feature, in: IEEE International Sy mposium on Signal Processing and Information Technology -ISSPIT, 201 4, Noida, India
- [7]
-
[8]
T. L. Eadie, P. C. Doyle, Classification of dysphonic voic e: Acoustic and auditory-perceptual measures, Journal of Voice 19 (1) ( 2005) 1–14. doi:https://doi.org/10.1016/j.jvoice.2004.02.002
-
[9]
L. M. Jesus, S. Castilho, A. Ferreira, M. Conceição Costa , Discrim- inative segmental cues to vowel height and consonantal plac e and voicing in whispered speech, Journal of Phonetics 97 (2023) 1–21. doi:https://doi.org/10.1016/j.wocn.2023.101223
arXiv 2023
Show all 64 references
-
[10]
Silva, C
J. Silva, C. Cardoso, M. Oliveira, L. Jesus, A. Ferreira , A comparative study of european portuguese stop consonants and fricative s in whis- pered speech and normal speech for real-time operation of vo ice conver- sion, in: 12th International Workshop on Models and Analysi s...
2021
-
[11]
T. Ito, K. Takeda, F. Itakura, Analysis and recognition of whispered speech, Speech Communication 45 (2) (2005) 139–1 52. doi:https://doi.org/10.1016/j.specom.2003.10.005
2005 doi
-
[12]
Zwicker, H
E. Zwicker, H. Fastl, Psychoacoustics, Facts and Model s, Springer- Verlag, 1990
1990
-
[13]
B. C. J. Moore, An Introduction to the Psychology of Hear ing, Academic Press, 1989
1989
-
[14]
S. R. Cox, C. H. Shadle, W.-R. Chen, Acoustic variabilit y in electrola- ryngeal speech, The Journal of the Acoustical Society of Ame rica 146 (4) (2019) 2921–2921. doi:10.1121/1.5137139
2019 doi
-
[15]
Michelle Cohan, This ‘artificial larynx’ prototype aim s to give cancer survivors their own voices back, online CNN document ary: <https://edition.cnn.com/2022/03/22/health/syrinx-artificial-electrolarynx-japan-spc-scn-intl/> (2022)
2022
-
[16]
Ferreira, M
A. Ferreira, M. Oliveira, V. Santos, On the mismatch bet ween the phase structure of all-pole-based synthetic vowels and natural v owels, in: IEEE Workshop on Signal Processing Systems (SiPS), 2024, pp. 1–6
2024
-
[17]
Silva, M
J. Silva, M. Oliveira, A. Ferreira, Manipulation of the fundamental fre- quency micro-variations using a fully parametric and compu tationally efficient speech model, in: 34th IEEE Workshop on Signal Proce ssing Systems (SiPS), 2020, pp. 1–6
2020
-
[18]
Ferreira, J
A. Ferreira, J. Silva, F. Brito, D. Sinha, Impact of a shi ft-invariant harmonic phase model in fully parametric harmonic voice rep resentation and time/frequency synthesis, in: IEEE International Conf erence on Acoustics, Speech and Signal Processing, 2020
2020
-
[19]
Silva, M
J. Silva, M. Oliveira, A. J. S. Ferreira, Flexible param etric implantation of voicing in whispered speech under scarce training data, i n: 28th Euro- pean Signal Processing Conference (EUSIPCO-2020), 2020, p p. 416–420
2020
-
[20]
J. ao P. Cabral, J. Kane, C. Gobl, J. Carson-Berndsen, Ev aluation of glottal epoch detection algorithms on different voice types , in: Inter- speech 2011, 2011, pp. 1989–1992. doi:10.21437/Interspee ch.2011-523. 53
2011 doi
-
[21]
Harris, D
J. Harris, D. Nelson, Glottal pulse alignment in voiced speech for pitch determination, in: IEEE International Conference on Acous- tics, Speech, and Signal Processing, Vol. 2, 1993, pp. 519–5 22. doi:10.1109/ICASSP.1993.319357
1993
-
[22]
Hagmuller, G
M. Hagmuller, G. Kubin, Poincaré pitch marks, Speech Communication 48 (12) (2006) 1650–1665. doi:https://doi.org/10.1016/j.specom.2006.07.008
2006 doi
-
[23]
Dikshit, S
P. Dikshit, S. Zahorian, S. Nagulapati, An algorithm fo r locating fun- damental frequency markers in speech signals, in: IEEE Inte rnational Conference on Acoustics, Speech, and Signal Processing, Vo l. 1, 2005, pp. I/233–I/236. doi:10.1109/ICASSP.2005.1415093
2005 arXiv
-
[24]
Cheng, D
Y. Cheng, D. O’Shaughnessy, Automatic and reliable est imation of glot- tal closure instant and period, IEEE Transactions on Acoust ics, Speech, and Signal Processing 37 (12) (1989) 1805–1815. doi:10.110 9/29.45529
1989
-
[25]
Drugman, M
T. Drugman, M. Thomas, J. Gudnason, P. Naylor, T. Dutoit , Detection of glottal closure instants from speech signals: A quantita tive review, IEEE Transactions on Audio, Speech, and Language Processin g 20 (3) (2012) 994–1006. doi:10.1109/TASL.2011.2170835
2012
-
[26]
Ananthapadmanabha, B
T. Ananthapadmanabha, B. Yegnanarayana, Epoch extrac tion from lin- ear prediction residual for identification of closed glotti s interval, IEEE Transactions on Acoustics, Speech, and Signal Processing 2 7 (4) (1979) 309–319. doi:10.1109/TASSP.1979.1163267
1979
-
[27]
Paul Boersma and David Weenink, Praat: software for spe ech analysis and synthesis, available from: <http://www.praat.org> (2 005)
-
[29]
Henrich, C
N. Henrich, C. d’Alessandro, B. Doval, M. Castellengo, On the use of the derivative of electroglottographic signals for charac terization of non- pathological phonation, The Journal of the Acoustical Soci ety of Amer- ica 115 (3) (2004) 1321–1332. doi:10.1121/1.1646401. 54
2004 doi
-
[30]
Perrotin, I
O. Perrotin, I. V. McLoughlin, Glottal flow synthesis fo r whisper-to-speech conversion, IEEE/ACM Transactions on A u- dio, Speech, and Language Processing 28 (2020) 889–900. doi:10.1109/TASLP.2020.2971417
2020
-
[31]
A. Ferreira, Implantation of voicing on whispered spee ch using frequency-domain parametric modelling of source and filter information, in: International Symposium on Signal, Image, Video and Com munica- tions (ISIVC), 2016, pp. 159–166, Tunis, Tunisia
2016
-
[32]
R. W. Morris, M. A. Clements, Reconstruction of speech f rom whispers, Medical Engineering & Physics 24 (7) (2002) 515–520
2002
-
[33]
I. V. Mcloughlin, H. R. Sharifzadeh, S. L. Tan, J. Li, Y. S ong, Re- construction of phonated speech from whispers using forman t-derived plausible pitch modulation, ACM Trans. Access. Comput. 6 (4 ) (2015) 12:1–12:21
2015
-
[34]
I. V. Mcloughlin, J. Li, Y. Song, Reconstruction of cont inuous voiced speech from whispers, in: Proceeedings of Interspeech, 201 3, pp. 1022– 1026
-
[35]
H. R. Sharifzadeh, I. V. McLoughlin, F. Ahmadi, Reconst ruction of normal sounding speech for laryngectomy patients through a modified CELP codec, IEEE Transactions on Biomedical Engineering 57 (10) (2010) 2448–2458
2010
-
[36]
Airaksinen, J
M. Airaksinen, J. L. uvela, B. B. ollepalli, J. Yamagish i, P. Alku, A comparison between straight, glottal, and sinusoidal voc oding in statistical parametric speech synthesis, IEEE/ACM Transa ctions on Audio, Speech, and Language Processing 26 (9) (2018) 1658–1 670. doi:10.1...
2018
-
[37]
T. Toda, M. Nakagiri, K. Shikano, Statistical voice con version techniques for body-conducted unvoiced speech enhancement, IEEE Tran sactions on Acoustics, Speech and Signal Processing 20 (9) (2012) 250 5–2517
2012
-
[38]
M. A. Oliveira, Machine learning approaches for whispe r to normal speech conversion: A survey, U.Porto Journal of Engineerin g 8 (2) (2022) 202–212. 55
2022
-
[39]
H. Lian, Y. Hu, J. Zhou, H. Wang, L. Tao, Whisper to normal speech based on deep neural networks with mcc and f0 features, in: 20 18 IEEE 23rd International Conference on Digital Signal Processin g (DSP), 2018, pp. 1–5. doi:10.1109/ICDSP.2018.8631888
2018
-
[40]
G. N. Meenakshi, P. K. Ghosh, Whispered speech to neutra l speech conversion using bidirectional lstms, in: Interspeech, 20 18, pp. 491–495. doi:10.21437/Interspeech.2018-1487
2018 doi
-
[41]
Konno, M
H. Konno, M. Kudo, H. Imai, M. Sugimoto, Whisper to norma l speech conversion using pitch estimated from spectrum, Speech Com munication 83 (2016) 10–20. doi:https://doi.org/10.1016/j.specom. 2016.07.001
2016 doi
-
[42]
H. Lian, Y. Hu, W. Yu, J. Zhou, W. Zheng, Whisper to nor- mal speech conversion using sequence-to-sequence mapping model with auditory attention, IEEE Access 7 (2019) 130495–13050 4. doi:10.1109/ACCESS.2019.2940700
2019
-
[43]
J. Rekimoto, Wesper: Zero-shot and realtime whisper to normal voice conversion for whisper-based speech interactions, in: Pro ceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2 023. doi:10.1145/3544548.3580706
2023
-
[44]
T. Tan, H. Ruan, X. Chen, K. Chen, Z. Lin, J. Lu, Distillw2 n: A lightweight one-shot whisper to normal voice conversion mo del using distillation of self-supervised features, in: 2025 IEEE In ternational Con- ference on Acoustics, Speech and Signal Processing (ICASSP ), 2025,...
2025
-
[45]
C. F. Yamamura, P. R. Scalassara, M. A. Oliveira, A. J. S. Fer- reira, Neural network models for whisper to normal speech co n- version, U.Porto Journal of Engineering 11 (1) (2025) 116–1 29. doi:doi.org/10.24840/2183-6493_0011-001_002739
2025 doi
-
[46]
Wagner, I
D. Wagner, I. Baumann, T. Bocklet, Generative adversar ial net- works for whispered to voiced speech conversion: a comparat ive study, International Journal of Speech Technology 27 (2024 ) 1093–1110. doi:doi.org/10.1007/s10772-024-10161-1
2024 doi
-
[47]
Pascual, A
S. Pascual, A. Bonafonte, J. Serrà, J. A. González López , Whispered-to-voiced alaryngeal speech conversion with ge nerative 56 adversarial networks, in: IberSPEECH 2018, 2018, pp. 117–1 21. doi:10.21437/IberSPEECH.2018-25
2018 doi
-
[48]
Niranjan, M
A. Niranjan, M. Sharma, S. B. C. Gutha, M. A. B. Shaik, End -to- end whisper to natural speech conversion using modified tran sformer network, preprint available from: <https://arxiv.org/ab s/2004.09347> (2021)
2021 arXiv
-
[49]
Kawahara, J
H. Kawahara, J. Estill, O. Fujimura, Aperiodicity extr action and control using mixed mode excitation and group delay manipulation fo r a high quality speech analysis, modification and synthesis system STRAIGHT, in: 2nd International Workshop on Models and Analysis of Voc al Em...
2001
-
[50]
MORISE, F
M. MORISE, F. YOKOMORI, K. OZA W A, World: A vocoder-base d high-quality speech synthesis system for real-time applic ations, IEICE Transactions on Information and Systems E99.D (7) (2016) 18 77–1884. doi:10.1587/transinf.2015EDP7457
2016 doi
-
[51]
Oliveira, V
M. Oliveira, V. Santos, A. Saraiva, A. Ferreira, Demyst ifying DFT-based harmonic phase estimation, transformation, and synthesis , Signals 5 (4) (2024) 841–868. doi:10.3390/signals5040046
2024 doi
-
[52]
A. J. Ferreira, J. M. Tribolet, A holistic glottal phase related feature, in: 21st International Conference on Digital Audio Effects ( DAFx-18), 2018, A veiro, Portugal
2018
-
[53]
A. V. Oppenheim, R. W. Schafer, Discrete-Time Signal Pr ocessing, Pearson Higher Education, Inc., 2010
2010
-
[54]
G. D. Nunes, Whispered speech segmentation based on dee p learning, Master’s thesis, Faculty of Engineering of the University o f Porto, Por- tugal, last accessed on Apr 23rd 2025. (2023). URL https://hdl.handle.net/10216/152039
2023
-
[55]
J. F. T. Costa, Adaptive phonetic segmentation in dysph onic voice, Master’s thesis, Faculty of Engineering of the University o f Porto, Portugal, last accessed on Apr 23rd 2025. (2021). URL https://repositorio-aberto.up.pt/bitstream/10216/ 133678/2/463680.pdf 57
2021
-
[56]
J. M. Silva, M. A. Oliveira, A. F. Saraiva, A. J. S. Ferrei ra, One-step discrete Fourier Transform-based sinusoid frequency esti mation under full-bandwidth quasi-harmonic interference, Acoustics 5 (3) (2023) 845–
2023
-
[57]
P. P. Vaidyanathan, Multirate Systems and Filter Banks , Prentice-Hall, 1993
1993
-
[58]
A. J. S. Ferreira, Accurate estimation in the ODFT domai n of the fre- quency, phase and magnitude of stationary sinusoids, in: 20 01 IEEE Workshop on Applications of Signal Processing to Audio and A coustics, 2001, pp. 47–50
2001
-
[59]
M. H. Hayes, Statistical Digital Signal Processing and Modeling, John Wiley & Sons Inc., 1996
1996
-
[60]
Malvar, Signal Processing with Lapped Transforms, A rtech House, Inc., 1992
H. Malvar, Signal Processing with Lapped Transforms, A rtech House, Inc., 1992
1992
-
[61]
Ferreira, D
A. Ferreira, D. Sinha, Advances to a frequency-domain p arametric coder of wideband speech, 140th Convention of the Audio Engineeri ng Soci- etyPaper 9509 (May 2016)
2016
-
[62]
Ferreira, On the physiological validity of the group delay response of all-pole vocal tract modeling, 145th Convention of the Audi o Engineer- ing SocietyPaper 10038 (October 2018)
A. Ferreira, On the physiological validity of the group delay response of all-pole vocal tract modeling, 145th Convention of the Audi o Engineer- ing SocietyPaper 10038 (October 2018)
2018
-
[63]
ITU-R Recommendation BS.1116-3, Methods for the subje ctive assess- ment of small impairments in audio systems (February 2015)
2015
-
[64]
ITU-R Recommendation BS.1284-2, General methods for t he subjective assessment of sound quality (January 2019)
2019
-
[65]
Fujisaki, K
H. Fujisaki, K. Hirose, Analysis of voice fundamental f requency contours for declarative sentences of japanese, Journal of the Acous tical Society of Japan 5 (4) (1984) 233–242. doi:10.1250/ast.5.233. 58
1984 doi
-
[869]
doi:10.3390/acoustics5030049
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.