Pith. sign in

REVIEW 4 major objections 6 minor 64 references

Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Co-articulated vowel spectra progress by interpolation between sustained-vowel templates, and a glottal-pulse synthesis method matches frequency-domain quality in perceptual tests.

desk verdict A careful, honest engineering paper with a genuinely new phase-based pitch-pulse analysis and solid listening tests, but the template-interpolation idea that motivates the whisper-restoration use case is not actually tested by the synthesis experiments. read the letter →

arxiv 2506.06675 v1 pith:KKOBXI42 submitted 2025-06-07 eess.AS cs.SD

classification eess.AScs.SD
keywords whisperedspeechpitchpulsesegmentationsyntheticvoicingharmonicphasestructureNormalizedRelativeDelayco-articulatedvowelsglottalvoicerestoration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Whispered speech lacks the periodic components of natural voice, and restoring them in real time on a portable device is the practical target. The paper tackles two connected problems: how the harmonic magnitude and phase structure of individual pitch pulses evolves in natural vowels, and which lightweight synthesis architecture best reconstructs that structure. It introduces a phase-based pitch-pulse segmentation method and reports that, in the measured transitions, co-articulated vowel spectra and phase profiles sit on a path between the same speaker's sustained-vowel templates. On that basis it compares three voicing synthesis methods and finds that for co-articulated vowels the physiologically inspired glottal-pulse method (GLO) matches the frequency-domain method perceptually and beats the combined frequency/time method. This matters because running speech consists mostly of co-articulated vowels, and the leading methods are computationally light enough for low-resource implementation.

What carries the argument

The key object is the phase-oriented pitch-pulse segmentation, which locates each pitch period by searching for the time shift at which the DFT phase of the fundamental-frequency bin equals $-\pi/2$ (the onset of $\phi_0$), then derives per-period harmonic phases $\phi_\ell=\angle X_\ell[1+\ell]+\pi/2$ and the Normalized Relative Delay vector $\mathrm{NRD}_\ell=(\phi_\ell-(\ell+1)\phi_0)/2\pi$, taken modulo 1 and unwrapped. This yields a per-pulse view of harmonic magnitude and phase that frame-based analysis cannot provide. The comparison of voicing methods rests on three synthesis architectures; the one carrying the new claim is GLO, which generates each glottal excitation pulse individually with a Liljencrants-Fant glottal pulse model (a standard parametric shape for the airflow through the glottis), filters it with a dedicated all-pole vocal-tract filter, and overlap-adds the filtered pulses, keeping the glottal source explicitly separate from the tract response.

What would settle it

Record co-articulated vowel transitions from many speakers and languages, segment them with the phase-based pitch-pulse method, and compare each intermediate pulse's NRD vector and magnitude spectrum against the interpolation of that speaker's sustained-vowel templates. If the deviation between measured and interpolated spectra exceeds the natural period-to-period scatter seen within sustained vowels, the interpolation model is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that the harmonic magnitude/phase structure across successive pitch pulses in a co-articulated vowel region can be modeled as a progressive interpolation between the natural, idiosyncratic spectral profiles of the same vowels produced in sustained mode, rather than as an arbitrary sequence of frames. The evidence comes from phase-based segmentation of the /i/–/a/ transition in the Portuguese word 'Tiago' for two speakers: the Normalized Relative Delay vectors and magnitude spectra move from the /i/ template toward the /a/ template over about 24 pitch periods. The perceptual claim is that, for co-articulated vowels in word context, GLO—which synthesizes individual glottal excitation pulses and filters each through a vocal-tract model—is rated at least as natural as frequency-domain synthesis and better than combined frequency/time synthesis, while for sustained vowels the frequency-domain method remains the strongest.

Load-bearing premise

The interpolation result is inferred from one vowel transition (/i/ to /a/ in the Portuguese word 'Tiago') produced by two speakers, plus a three-vowel 'Saiu' sequence used only in listening tests; the paper's synthesis strategy for running whispered speech assumes this pattern holds across speakers, vowel pairs, and languages.

Editorial extensions

If this is right

  • An entire co-articulated vowel region can be synthesized by pitch-pulse-by-pitch-pulse interpolation between stored spectral templates of the same speaker, so no per-frame formant tracking is needed.
  • For running whispered speech, the physiologically inspired GLO method is the better candidate among the tested options, since it matched the frequency-domain method and beat the combined frequency/time method on co-articulated vowels.
  • For sustained vowels, the frequency-domain method remains a very efficient high-quality choice, statistically indistinguishable from the original in several listening conditions.
  • Because GLO explicitly separates glottal excitation from vocal-tract filtering, voice identity and phonation type can in principle be controlled independently of articulation.
  • The phase-based pitch-pulse segmentation can also serve as a monitoring tool, since its f0 contours agree with frame-based pitch estimates and it can verify correct operation of synthetic voicing algorithms.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the interpolation property generalizes, a whisper-restoration device could be built from a small per-speaker dictionary of sustained-vowel templates, avoiding the large paired datasets that data-driven whisper-to-speech systems require; the paper aims at this but does not test it.
  • A direct extension would stress-test the model on vowel pairs with very different formant spacing (e.g., /i/–/u/), across languages and speech rates, to see whether co-articulation always tracks the sustained templates or sometimes takes a shortcut between articulatory targets.
  • The observation that GLO's harmonic structure can be cleaner than the original suggests a clinical use beyond restoration: the same segmentation and synthesis machinery could quantify harmonic damage in dysphonic voices, which the paper mentions only in passing.
  • Replacing the idealized Liljencrants-Fant glottal model with pulses derived from physiological recordings is named as future work in the paper; the explicit source/filter separation in GLO makes that replacement straightforward and directly testable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper addresses two interconnected challenges in whispered-speech-to-voiced-speech restoration: (i) a new phase-based procedure for segmenting individual pitch pulses in sustained and co-articulated vowels and extracting their per-period harmonic magnitude and NRD phase structure, and (ii) three lightweight parametric synthetic-voicing alternatives (frequency-domain FRE, combined frequency/time-domain TIM, and physiologically inspired time-domain GLO), compared through spectrograms and listening tests following ITU-R BS.1116. The authors report that the harmonic magnitude/phase structure of co-articulated vowels evolves as a progressive interpolation between the structures of the same vowels produced in sustained mode, and that GLO is perceptually comparable to FRE for co-articulated vowels in word context. The paper includes public audio files of the 32 test stimuli and derives the DFT phase mathematics in detail.

Significance. If the interpolation result were established across speakers, vowel pairs, and languages, it would provide an interpretable template-based pathway for lightweight, real-time, on-the-fly whisper-to-speech restoration, which is the paper's stated goal. The work has clear strengths: the DFT phase analysis in Section 2.1 is formally derived, the listening-test protocol follows ITU-R BS.1116 with 30 listeners and randomized hidden references, the f0-contour comparison in Figure 2 provides a useful sanity check of the segmentation procedure, and the authors make the test audio files publicly available. The comparative synthesis evaluation is a useful contribution even though, as detailed below, it does not exercise the template-interpolation pathway. The significance is currently limited by the narrow evidential basis for the interpolation claim and by the mismatch between the stated synthesis strategy and the actual synthesis experiments.

major comments (4)
  1. [Section 2.3, Insight 3; Figs. 4 and 5] The progressive-interpolation claim is supported by only a single /i/-/a/ co-articulation in the Portuguese word 'Tiago' for each of two speakers, and the evidence is presented only as overlaid curves. No quantitative distance metric compares the observed per-period NRD/magnitude evolution against the interpolation model, and no cross-validation across other vowel pairs, speakers, or languages is provided. Because the proposed whisper-restoration strategy depends on this assumption, the claim is currently under-supported and needs either a quantitative test on the existing data or additional data.
  2. [Section 3.1, Eq. (10); Section 4.3] The synthesis evaluation does not implement the proposed template-interpolation pathway. In the FRE/TIM/GLO versions, each word is reconstructed using parameters extracted from the original co-articulated recording: Eq. (10) interpolates between frame-boundary parameters of the original signal, not between endpoint templates of sustained vowels. The listening tests in Tables 5-8 therefore validate reconstruction of a word from its own analysis parameters, not synthesis of co-articulation from vowel templates. Since the stated whisper-restoration strategy requires the latter, the central synthesis claim is not validated by the experiments as designed.
  3. [Section 2.1] The phase-based pitch-period segmentation procedure is not validated against any ground truth, such as EGG signals, manual epoch marks, or synthetic signals with known period boundaries. The consistency of overlays in Figure 3 and the agreement between 'ORI' and 'MOD' f0 contours in Figure 2 are indirect evidence; without a quantitative accuracy measure, the per-period NRD/magnitude evolution that underlies the interpolation claim is not established on firm footing.
  4. [Section 4.3, Tables 1-8] The listening-test analysis relies on many pairwise two-tailed t-tests without correction for multiple comparisons, and the absence of statistically significant differences is repeatedly interpreted as evidence of equivalence (e.g., FRE vs. GLO in 'Tiago' and 'Saiu'). An ANOVA or a mixed-effects model with appropriate post-hoc corrections, or explicit equivalence testing, is needed to support the comparative conclusions drawn from these data.
minor comments (6)
  1. [Section 4.1] The word 'representated' should be 'represented'.
  2. [Section 4.3, bullet list] The phrase 'the the combined frequency and time-domain voicing' contains a duplicated article.
  3. [Section 5] The phrase 'Is general' should be 'In general'.
  4. [Section 1.2] The phrase 'more that finding' should be 'more than finding'.
  5. [Section 3.3, Eq. (10)] The notation in Eq. (10) does not explicitly define the range of the harmonic index L or the normalization of NRD in the synthesis context; please state these definitions for reproducibility.
  6. [Figure 4 caption] The caption identifies the dashed magenta line as the sustained /a/ reference and the dashed blue line as the sustained /i/ reference, but the order is reversed relative to the text description in Section 2.2; please clarify the mapping.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the central interpolation result is an empirical observation, and the synthesis conclusions rest on external listening tests rather than on a fitted-input prediction.

full rationale

The paper's load-bearing claims are (i) that the harmonic phase/magnitude structure of co-articulated vowels can be viewed as a progressive interpolation between vowel-specific templates (Section 2.3, Insight 3), and (ii) that the three voicing-restoration methods, especially FRE and GLO, deliver perceptually acceptable quality (Section 4.3). Neither claim reduces by construction to its inputs. The interpolation insight is supported by direct pitch-period analysis of natural recordings (Figs. 4-5) and is presented as an empirical interpretation, not as a quantity forced by a fitted parameter; the later synthesis tests (Eq. 10 and Section 4.2) reconstruct each word from parameters extracted from that same word, so they do not constitute a template-interpolation prediction, but that is a scope limitation rather than circularity. The NRD/LPC/ODFT analysis tools and the FRE/TIM synthesis blocks are reused from the authors' earlier work [18,19,51], with self-citations present and normal; however, the pitch-pulse segmentation is benchmarked against a different frame-based pitch estimator (Fig. 2), and the GLO method deliberately omits NRD, giving an independent synthesis path. The comparative quality claims are anchored in ITU-R BS.1116 listening tests with 30 listeners and hidden references (Section 4.3), which are external perceptual evidence not constructed from the paper's own parameters. Accordingly, the central results have independent content, and no specific equation or fitted parameter is renamed as a prediction. Score 2 reflects only the presence of expected self-citations, none of which is load-bearing.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters are fitted in this paper; the LF glottal model and LPC analysis parameters are inherited from prior work. The central claim rests on five background assumptions, the most fragile being the interpolation hypothesis, which is supported only by illustrative data. No new physical entities are introduced.

assumptions (5)
  • domain assumption The source-filter model of speech production with independent glottal source and vocal tract filter is valid.
    Invoked throughout; e.g., Sections 1.3.1 and 3.1.
  • standard math DFT analysis of a P-point pitch period locates the f0 onset when the phase of bin k=1 is approximately -pi/2, so per-period phase segmentation is well-posed.
    Eqs. (3)-(8) provide the analytic basis for this phase criterion.
  • domain assumption Unvoiced sounds in whispered speech do not differ appreciably from those in natural speech for the same speaker.
    Stated in the Introduction and supported by citations [9,10,11]; this grounds the 'implant voicing' approach.
  • domain assumption The NRD feature, combined with harmonic magnitudes, completely determines waveform shape independently of f0, time-shift, and overall magnitude.
    Adopted from [51] and used as the analysis foundation in Section 2.1.
  • ad hoc to paper Co-articulation can be modeled as a progressive interpolation between sustained-vowel spectral templates.
    This is the paper's central empirical hypothesis, stated in Section 2.3 and used to justify the interpolation-based synthesis in Section 3.3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods." pith.science (2026). https://pith.science/paper/KKOBXI42

@misc{pith2026250606675,
  author       = {Pith},
  title        = {Pith review of: Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KKOBXI42}},
  note         = {Machine review of arXiv:2506.06675}
}
read the original abstract

Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal components that exist in certain regions of the former, called voiced regions, as a consequence of the vibration of the vocal folds, are missing in the latter. The restoration of natural speech from whispered speech requires delicate signal processing procedures that are especially useful if they can be implemented on low-resourced portable devices, in real-time, and on-the-fly, taking advantage of the established source-filter paradigm of voice production and related models. This paper addresses two challenges that are intertwined and are key in informing and making viable this envisioned technological realization. The first challenge involves characterizing and modeling the evolution of the harmonic phase/magnitude structure of a sequence of individual pitch periods in a voiced region of natural speech comprising sustained or co-articulated vowels. This paper proposes a novel algorithm segmenting individual pitch pulses, which is then used to obtain illustrative results highlighting important differences between sustained and co-articulated vowels, and suggesting practical synthetic voicing approaches. The second challenge involves model-based synthetic voicing. Three implementation alternatives are described that differ in their signal reconstruction approaches: frequency-domain, combined frequency and time-domain, and physiologically-inspired separate filtering of glottal excitation pulses individually generated. The three alternatives are compared objectively using illustrative examples, and subjectively using the results of listening tests involving synthetic voicing of sustained and co-articulated vowels in word context.

Figures

Figures reproduced from arXiv: 2506.06675 by the authors.

Figure 1
Figure 1. From top to bottom: time wave and spectrogram of a vo [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Fundamental frequency contour of a sustained /a/ v [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 3
Figure 3. From left to right: overlay of pitch periods found i [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: From left to right: overlay of pitch periods found i [PITH_FULL_IMAGE:figures/full_fig_p019_4.png]
Figure 5
Figure 5. Figure 5: Magnitude spectra of 24 successive pitch periods s [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Simplified frequency-domain analysis framework t [PITH_FULL_IMAGE:figures/full_fig_p024_6.png]
Figure 7
Figure 7. Figure 7: Three parametric-oriented alternatives to synth [PITH_FULL_IMAGE:figures/full_fig_p026_7.png]
Figure 8
Figure 8. Figure 8: Reconstruction of x[n] according to the frequency-domain synthesis approach. After being multiplied by window w[n], adjacent segments including xm−1[n], xm[n], and xm+1[n] are overlapped and added yielding x[n]. described in the next two subsections. 3.3. Combined freq…
Figure 9
Figure 9. Figure 9: Reconstruction of x[n] according to the combined frequency and time-domain synthesis alternative. Pitch periods are individually shaped in frequency, synthesized and then concatenated in the time domain. x[n] = X p xp [n − np] , where np accumulates past pitch periods,…
Figure 10
Figure 10. Figure 10: Reconstruction of x[n] according to the physiologically inspired time-domain synthesis alternative. Glottal excitation periods are synthesized individually, are filtered according to a filter modeling an individual spectral envelope, and are overlapped and added to yi…
Figure 11
Figure 11. Figure 11: Spectrograms of the original (ORI) and synthetic [PITH_FULL_IMAGE:figures/full_fig_p034_11.png]
Figure 12
Figure 12. Figure 12: Spectrograms of the original (ORI) and synthetic [PITH_FULL_IMAGE:figures/full_fig_p036_12.png]
Figure 13
Figure 13. Figure 13: Spectrograms of the original (ORI) and synthetic [PITH_FULL_IMAGE:figures/full_fig_p037_13.png]
Figure 14
Figure 14. Figure 14: Spectrograms of the original (ORI) and synthetic [PITH_FULL_IMAGE:figures/full_fig_p038_14.png]
Figure 15
Figure 15. Figure 15: Spectrograms of the original (ORI) and synthetic [PITH_FULL_IMAGE:figures/full_fig_p039_15.png]
Figure 16
Figure 16. Figure 16: For each one of the 8 listening tasks, the vertical [PITH_FULL_IMAGE:figures/full_fig_p041_16.png]
Figure 17
Figure 17. Figure 17: Means and associated 95% confidence intervals of t [PITH_FULL_IMAGE:figures/full_fig_p043_17.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

64 extracted references · 57 canonical work pages

  1. [1]

    Fant, Acoustic Theory of Speech Production, The Hague , 1970

    G. Fant, Acoustic Theory of Speech Production, The Hague , 1970

  2. [2]

    O’Shaughnessy, Speech Communications: Human and Mac hine, Wiley-IEEE Press, 2000

    D. O’Shaughnessy, Speech Communications: Human and Mac hine, Wiley-IEEE Press, 2000

  3. [4]

    Rabiner, B.-H

    L. Rabiner, B.-H. Juang, Fundamentals of Speech Recogni tion, Prentice- Hall, Inc., 1993

  4. [5]

    A. V. Oppenheim, A. S. Willsky, S. Hamid, Signals and Syst ems, Pear- son Education Limited, 1996, 2nd Ed

  5. [6]

    A. Ferreira, On the possibility of speaker discriminati on using a glot- tal pulse phase-related feature, in: IEEE International Sy mposium on Signal Processing and Information Technology -ISSPIT, 201 4, Noida, India

  6. [7]

    Dellwo, M

    V. Dellwo, M. Huckvale, M. Ashby, How Is Individuality Ex pressed in Voice? An Introduction to Speech Production and Descript ion for Speaker Classification, Springer Berlin Heidelberg, Berli n, Heidelberg, 2007, pp. 1–20

  7. [8]

    T. L. Eadie, P. C. Doyle, Classification of dysphonic voic e: Acoustic and auditory-perceptual measures, Journal of Voice 19 (1) ( 2005) 1–14. doi:https://doi.org/10.1016/j.jvoice.2004.02.002

  8. [9]

    L. M. Jesus, S. Castilho, A. Ferreira, M. Conceição Costa , Discrim- inative segmental cues to vowel height and consonantal plac e and voicing in whispered speech, Journal of Phonetics 97 (2023) 1–21. doi:https://doi.org/10.1016/j.wocn.2023.101223

Show all 64 references
  1. [10]

    Silva, C

    J. Silva, C. Cardoso, M. Oliveira, L. Jesus, A. Ferreira , A comparative study of european portuguese stop consonants and fricative s in whis- pered speech and normal speech for real-time operation of vo ice conver- sion, in: 12th International Workshop on Models and Analysi s...

  2. [11]

    T. Ito, K. Takeda, F. Itakura, Analysis and recognition of whispered speech, Speech Communication 45 (2) (2005) 139–1 52. doi:https://doi.org/10.1016/j.specom.2003.10.005

  3. [12]

    Zwicker, H

    E. Zwicker, H. Fastl, Psychoacoustics, Facts and Model s, Springer- Verlag, 1990

  4. [13]

    B. C. J. Moore, An Introduction to the Psychology of Hear ing, Academic Press, 1989

  5. [14]

    S. R. Cox, C. H. Shadle, W.-R. Chen, Acoustic variabilit y in electrola- ryngeal speech, The Journal of the Acoustical Society of Ame rica 146 (4) (2019) 2921–2921. doi:10.1121/1.5137139

  6. [15]

    Michelle Cohan, This ‘artificial larynx’ prototype aim s to give cancer survivors their own voices back, online CNN document ary: <https://edition.cnn.com/2022/03/22/health/syrinx-artificial-electrolarynx-japan-spc-scn-intl/> (2022)

  7. [16]

    Ferreira, M

    A. Ferreira, M. Oliveira, V. Santos, On the mismatch bet ween the phase structure of all-pole-based synthetic vowels and natural v owels, in: IEEE Workshop on Signal Processing Systems (SiPS), 2024, pp. 1–6

  8. [17]

    Silva, M

    J. Silva, M. Oliveira, A. Ferreira, Manipulation of the fundamental fre- quency micro-variations using a fully parametric and compu tationally efficient speech model, in: 34th IEEE Workshop on Signal Proce ssing Systems (SiPS), 2020, pp. 1–6

  9. [18]

    Ferreira, J

    A. Ferreira, J. Silva, F. Brito, D. Sinha, Impact of a shi ft-invariant harmonic phase model in fully parametric harmonic voice rep resentation and time/frequency synthesis, in: IEEE International Conf erence on Acoustics, Speech and Signal Processing, 2020

  10. [19]

    Silva, M

    J. Silva, M. Oliveira, A. J. S. Ferreira, Flexible param etric implantation of voicing in whispered speech under scarce training data, i n: 28th Euro- pean Signal Processing Conference (EUSIPCO-2020), 2020, p p. 416–420

  11. [20]

    J. ao P. Cabral, J. Kane, C. Gobl, J. Carson-Berndsen, Ev aluation of glottal epoch detection algorithms on different voice types , in: Inter- speech 2011, 2011, pp. 1989–1992. doi:10.21437/Interspee ch.2011-523. 53

  12. [21]

    Harris, D

    J. Harris, D. Nelson, Glottal pulse alignment in voiced speech for pitch determination, in: IEEE International Conference on Acous- tics, Speech, and Signal Processing, Vol. 2, 1993, pp. 519–5 22. doi:10.1109/ICASSP.1993.319357

  13. [22]

    Hagmuller, G

    M. Hagmuller, G. Kubin, Poincaré pitch marks, Speech Communication 48 (12) (2006) 1650–1665. doi:https://doi.org/10.1016/j.specom.2006.07.008

  14. [23]

    Dikshit, S

    P. Dikshit, S. Zahorian, S. Nagulapati, An algorithm fo r locating fun- damental frequency markers in speech signals, in: IEEE Inte rnational Conference on Acoustics, Speech, and Signal Processing, Vo l. 1, 2005, pp. I/233–I/236. doi:10.1109/ICASSP.2005.1415093

  15. [24]

    Cheng, D

    Y. Cheng, D. O’Shaughnessy, Automatic and reliable est imation of glot- tal closure instant and period, IEEE Transactions on Acoust ics, Speech, and Signal Processing 37 (12) (1989) 1805–1815. doi:10.110 9/29.45529

  16. [25]

    Drugman, M

    T. Drugman, M. Thomas, J. Gudnason, P. Naylor, T. Dutoit , Detection of glottal closure instants from speech signals: A quantita tive review, IEEE Transactions on Audio, Speech, and Language Processin g 20 (3) (2012) 994–1006. doi:10.1109/TASL.2011.2170835

  17. [26]

    Ananthapadmanabha, B

    T. Ananthapadmanabha, B. Yegnanarayana, Epoch extrac tion from lin- ear prediction residual for identification of closed glotti s interval, IEEE Transactions on Acoustics, Speech, and Signal Processing 2 7 (4) (1979) 309–319. doi:10.1109/TASSP.1979.1163267

  18. [27]

    Paul Boersma and David Weenink, Praat: software for spe ech analysis and synthesis, available from: <http://www.praat.org> (2 005)

  19. [29]

    Henrich, C

    N. Henrich, C. d’Alessandro, B. Doval, M. Castellengo, On the use of the derivative of electroglottographic signals for charac terization of non- pathological phonation, The Journal of the Acoustical Soci ety of Amer- ica 115 (3) (2004) 1321–1332. doi:10.1121/1.1646401. 54

  20. [30]

    Perrotin, I

    O. Perrotin, I. V. McLoughlin, Glottal flow synthesis fo r whisper-to-speech conversion, IEEE/ACM Transactions on A u- dio, Speech, and Language Processing 28 (2020) 889–900. doi:10.1109/TASLP.2020.2971417

  21. [31]

    A. Ferreira, Implantation of voicing on whispered spee ch using frequency-domain parametric modelling of source and filter information, in: International Symposium on Signal, Image, Video and Com munica- tions (ISIVC), 2016, pp. 159–166, Tunis, Tunisia

  22. [32]

    R. W. Morris, M. A. Clements, Reconstruction of speech f rom whispers, Medical Engineering & Physics 24 (7) (2002) 515–520

  23. [33]

    I. V. Mcloughlin, H. R. Sharifzadeh, S. L. Tan, J. Li, Y. S ong, Re- construction of phonated speech from whispers using forman t-derived plausible pitch modulation, ACM Trans. Access. Comput. 6 (4 ) (2015) 12:1–12:21

  24. [34]

    I. V. Mcloughlin, J. Li, Y. Song, Reconstruction of cont inuous voiced speech from whispers, in: Proceeedings of Interspeech, 201 3, pp. 1022– 1026

  25. [35]

    H. R. Sharifzadeh, I. V. McLoughlin, F. Ahmadi, Reconst ruction of normal sounding speech for laryngectomy patients through a modified CELP codec, IEEE Transactions on Biomedical Engineering 57 (10) (2010) 2448–2458

  26. [36]

    Airaksinen, J

    M. Airaksinen, J. L. uvela, B. B. ollepalli, J. Yamagish i, P. Alku, A comparison between straight, glottal, and sinusoidal voc oding in statistical parametric speech synthesis, IEEE/ACM Transa ctions on Audio, Speech, and Language Processing 26 (9) (2018) 1658–1 670. doi:10.1...

  27. [37]

    T. Toda, M. Nakagiri, K. Shikano, Statistical voice con version techniques for body-conducted unvoiced speech enhancement, IEEE Tran sactions on Acoustics, Speech and Signal Processing 20 (9) (2012) 250 5–2517

  28. [38]

    M. A. Oliveira, Machine learning approaches for whispe r to normal speech conversion: A survey, U.Porto Journal of Engineerin g 8 (2) (2022) 202–212. 55

  29. [39]

    H. Lian, Y. Hu, J. Zhou, H. Wang, L. Tao, Whisper to normal speech based on deep neural networks with mcc and f0 features, in: 20 18 IEEE 23rd International Conference on Digital Signal Processin g (DSP), 2018, pp. 1–5. doi:10.1109/ICDSP.2018.8631888

  30. [40]

    G. N. Meenakshi, P. K. Ghosh, Whispered speech to neutra l speech conversion using bidirectional lstms, in: Interspeech, 20 18, pp. 491–495. doi:10.21437/Interspeech.2018-1487

  31. [41]

    Konno, M

    H. Konno, M. Kudo, H. Imai, M. Sugimoto, Whisper to norma l speech conversion using pitch estimated from spectrum, Speech Com munication 83 (2016) 10–20. doi:https://doi.org/10.1016/j.specom. 2016.07.001

  32. [42]

    H. Lian, Y. Hu, W. Yu, J. Zhou, W. Zheng, Whisper to nor- mal speech conversion using sequence-to-sequence mapping model with auditory attention, IEEE Access 7 (2019) 130495–13050 4. doi:10.1109/ACCESS.2019.2940700

  33. [43]

    J. Rekimoto, Wesper: Zero-shot and realtime whisper to normal voice conversion for whisper-based speech interactions, in: Pro ceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2 023. doi:10.1145/3544548.3580706

  34. [44]

    T. Tan, H. Ruan, X. Chen, K. Chen, Z. Lin, J. Lu, Distillw2 n: A lightweight one-shot whisper to normal voice conversion mo del using distillation of self-supervised features, in: 2025 IEEE In ternational Con- ference on Acoustics, Speech and Signal Processing (ICASSP ), 2025,...

  35. [45]

    C. F. Yamamura, P. R. Scalassara, M. A. Oliveira, A. J. S. Fer- reira, Neural network models for whisper to normal speech co n- version, U.Porto Journal of Engineering 11 (1) (2025) 116–1 29. doi:doi.org/10.24840/2183-6493_0011-001_002739

  36. [46]

    Wagner, I

    D. Wagner, I. Baumann, T. Bocklet, Generative adversar ial net- works for whispered to voiced speech conversion: a comparat ive study, International Journal of Speech Technology 27 (2024 ) 1093–1110. doi:doi.org/10.1007/s10772-024-10161-1

  37. [47]

    Pascual, A

    S. Pascual, A. Bonafonte, J. Serrà, J. A. González López , Whispered-to-voiced alaryngeal speech conversion with ge nerative 56 adversarial networks, in: IberSPEECH 2018, 2018, pp. 117–1 21. doi:10.21437/IberSPEECH.2018-25

  38. [48]

    Niranjan, M

    A. Niranjan, M. Sharma, S. B. C. Gutha, M. A. B. Shaik, End -to- end whisper to natural speech conversion using modified tran sformer network, preprint available from: <https://arxiv.org/ab s/2004.09347> (2021)

  39. [49]

    Kawahara, J

    H. Kawahara, J. Estill, O. Fujimura, Aperiodicity extr action and control using mixed mode excitation and group delay manipulation fo r a high quality speech analysis, modification and synthesis system STRAIGHT, in: 2nd International Workshop on Models and Analysis of Voc al Em...

  40. [50]

    MORISE, F

    M. MORISE, F. YOKOMORI, K. OZA W A, World: A vocoder-base d high-quality speech synthesis system for real-time applic ations, IEICE Transactions on Information and Systems E99.D (7) (2016) 18 77–1884. doi:10.1587/transinf.2015EDP7457

  41. [51]

    Oliveira, V

    M. Oliveira, V. Santos, A. Saraiva, A. Ferreira, Demyst ifying DFT-based harmonic phase estimation, transformation, and synthesis , Signals 5 (4) (2024) 841–868. doi:10.3390/signals5040046

  42. [52]

    A. J. Ferreira, J. M. Tribolet, A holistic glottal phase related feature, in: 21st International Conference on Digital Audio Effects ( DAFx-18), 2018, A veiro, Portugal

  43. [53]

    A. V. Oppenheim, R. W. Schafer, Discrete-Time Signal Pr ocessing, Pearson Higher Education, Inc., 2010

  44. [54]

    G. D. Nunes, Whispered speech segmentation based on dee p learning, Master’s thesis, Faculty of Engineering of the University o f Porto, Por- tugal, last accessed on Apr 23rd 2025. (2023). URL https://hdl.handle.net/10216/152039

  45. [55]

    J. F. T. Costa, Adaptive phonetic segmentation in dysph onic voice, Master’s thesis, Faculty of Engineering of the University o f Porto, Portugal, last accessed on Apr 23rd 2025. (2021). URL https://repositorio-aberto.up.pt/bitstream/10216/ 133678/2/463680.pdf 57

  46. [56]

    J. M. Silva, M. A. Oliveira, A. F. Saraiva, A. J. S. Ferrei ra, One-step discrete Fourier Transform-based sinusoid frequency esti mation under full-bandwidth quasi-harmonic interference, Acoustics 5 (3) (2023) 845–

  47. [57]

    P. P. Vaidyanathan, Multirate Systems and Filter Banks , Prentice-Hall, 1993

  48. [58]

    A. J. S. Ferreira, Accurate estimation in the ODFT domai n of the fre- quency, phase and magnitude of stationary sinusoids, in: 20 01 IEEE Workshop on Applications of Signal Processing to Audio and A coustics, 2001, pp. 47–50

  49. [59]

    M. H. Hayes, Statistical Digital Signal Processing and Modeling, John Wiley & Sons Inc., 1996

  50. [60]

    Malvar, Signal Processing with Lapped Transforms, A rtech House, Inc., 1992

    H. Malvar, Signal Processing with Lapped Transforms, A rtech House, Inc., 1992

  51. [61]

    Ferreira, D

    A. Ferreira, D. Sinha, Advances to a frequency-domain p arametric coder of wideband speech, 140th Convention of the Audio Engineeri ng Soci- etyPaper 9509 (May 2016)

  52. [62]

    Ferreira, On the physiological validity of the group delay response of all-pole vocal tract modeling, 145th Convention of the Audi o Engineer- ing SocietyPaper 10038 (October 2018)

    A. Ferreira, On the physiological validity of the group delay response of all-pole vocal tract modeling, 145th Convention of the Audi o Engineer- ing SocietyPaper 10038 (October 2018)

  53. [63]

    ITU-R Recommendation BS.1116-3, Methods for the subje ctive assess- ment of small impairments in audio systems (February 2015)

  54. [64]

    ITU-R Recommendation BS.1284-2, General methods for t he subjective assessment of sound quality (January 2019)

  55. [65]

    Fujisaki, K

    H. Fujisaki, K. Hirose, Analysis of voice fundamental f requency contours for declarative sentences of japanese, Journal of the Acous tical Society of Japan 5 (4) (1984) 233–242. doi:10.1250/ast.5.233. 58

  56. [869]

    doi:10.3390/acoustics5030049

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.