Pith. sign in

REVIEW 3 major objections 2 minor 54 references

Components Loss for Neural Networks in Mask-Based Speech Enhancement

T0 review · 3 major / 2 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A training loss that treats speech preservation, noise suppression, and residual-noise naturalness as separate terms gives mask-based speech-enhancement networks better perceptual quality and stronger noise attenuation than conventional…

desk verdict Clean, honest loss-function paper with a genuinely new third term, but the headline PESQ/SNR gains rest on a single small test set and need statistical backing before they carry the full weight. read the letter →

arxiv 1908.05087 v1 pith:ZWXSSPOL submitted 2019-08-14 eess.AS cs.SD

classification eess.AScs.SD
keywords speechenhancementmask-basedcomponentslossconvolutionalneuralnetworkresidualnoisequalitywhite-boxapproachPESQSNRimprovement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Mask-based speech enhancement trains a network to estimate a time–frequency mask that is applied to the noisy spectrum. This paper claims that the usual single global loss, such as mean squared error or a perceptual surrogate, is the bottleneck: it cannot separately reward keeping the speech intact, removing the noise, and leaving the remaining noise sounding natural. The proposed components loss (CL) replaces the global error with two or three weighted terms: speech-component distortion, residual-noise power, and, in the 3CL variant, similarity of the normalized residual-noise spectrum to the original noise. In experiments on one CNN architecture, networks trained with CL outperformed MSE, a perceptual-weighting-filter loss, and a differentiable PESQ loss on almost every reported metric, including at least 0.1 points higher PESQ on seen noises, about 0.2 points higher on unseen bus noise, and more than 0.5 dB higher SNR improvement. The practical point is that these gains come from changing only the loss, not the network, the data, or the training procedure.

What carries the argument

The object doing the work is the components loss (CL), a weighted sum of per-component mean-squared errors defined on the filtered speech and filtered noise spectra. For the 2CL variant, the first term $(1-\alpha)\sum_k(|\tilde{S}_\ell(k)|-|S_\ell(k)|)^2$ penalizes attenuation or distortion of the speech component, the second term $\alpha\sum_k|\tilde{D}_\ell(k)|^2$ penalizes residual noise power, and $\alpha\in[0,1]$ sets the trade-off. The 3CL variant adds a third term with weight $\beta$ that compares the normalized filtered noise spectrum $|\tilde{D}_\ell(k)|/\sqrt{\sum_\kappa|\tilde{D}_\ell(\kappa)|^2}$ with the normalized original noise spectrum, penalizing spectral reshaping of the residual noise while remaining zero for a pure fullband attenuation. The loss is naturally differentiable and is used inside the white-box training setup, where the mask is applied inside the network to the noisy magnitude spectrum and both components are available as targets.

What would settle it

Retrain the same mask-based CNN with MSE and with 3CL on a different corpus and a different network topology, then evaluate both on a held-out noise type; if 3CL's PESQ and SNR-improvement advantages over MSE fall below the reported 0.1-point and 0.5-dB thresholds, or reverse, the paper's central claim is not general.

Watch

Extended reading notes

Core claim

The central claim is that a mask-estimating CNN for single-channel speech enhancement should not be trained by comparing the enhanced spectrum to the clean spectrum alone, because that leaves the network free to mute low-SNR time–frequency bins, harming both speech detail and residual-noise naturalness. Instead, during training the known clean speech $S_\ell(k)$ and known noise $D_\ell(k)$ can each be multiplied by the estimated mask to form the filtered speech $\tilde{S}_\ell(k)$ and filtered noise $\tilde{D}_\ell(k)$, and the loss can be built from these components. The 2-component loss is $J_\ell^{\text{2CL}} = (1-\alpha)\sum_k (|\tilde{S}_\ell(k)|-|S_\ell(k)|)^2 + \alpha \sum_k |\tilde{D}_\ell(k)|^2$, with $\alpha$ trading speech preservation against noise attenuation. The 3-component loss adds a term comparing the normalized spectra of $\tilde{D}_\ell$ and $D_\ell$, so that a natural-sounding residual noise is preserved; the requirement that the weights satisfy $\alpha+\beta \le 1$ keeps the speech term from being dominated. On a fixed CNN evaluated with PESQ, POLQA, STOI, SSDR, and noise-quality measures, the paper reports that CL-trained networks give the best and most balanced performance, with speech-component quality and total enhanced-speech quality ahead of all three baseline losses.

Load-bearing premise

The load-bearing premise is that results measured on one CNN architecture and one speech/noise corpus are representative enough to support the paper's broad statement that the components loss is not restricted to any specific network topology or application.

Editorial extensions

If this is right

  • Switching from MSE to 3CL yields at least 0.1 points higher PESQ on seen noise types and about 0.2 points higher on unseen bus noise, with more than 0.5 dB higher SNR improvement in both cases.
  • Speech-component quality improves, with at least 0.5 dB higher SSDR and about 0.1 points higher PESQ on the filtered speech component for seen noises, meaning the enhanced speech retains more of the clean speech's detail.
  • Residual noise becomes more natural under 3CL than under MSE or the perceptual weighting filter loss, matching or beating the PESQ-loss baseline in WLAKR, the metric closest to musical-tone annoyance.
  • The benefits transfer without retraining the architecture or collecting new data: CL is a drop-in replacement for the loss function and is naturally differentiable.
  • Because the weights $\alpha$ and $\beta$ give explicit control over the noise-suppression versus speech-distortion trade-off, a system designer can tune the same network for more aggressive denoising or more conservative speech preservation by changing two scalars.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the paper tests only one unseen noise type and one CNN architecture, a natural extension would be to measure whether the 3CL advantage survives across several unseen noise classes and different mask-estimating architectures; the reported margins of 0.1–0.2 PESQ and 0.5 dB SNR give a concrete threshold for such a test.
  • The third 3CL term shapes the residual-noise spectrum toward the original noise; an untested corollary is that it may also act as a regularizer that reduces musical-tone artifacts beyond what WLAKR captures, so a listening study or a dedicated tonality metric would be a sharper test.
  • Since the loss needs access to clean speech and noise separately during training, it transfers most directly to fully supervised and simulation-based settings; adapting it to self-supervised or real-recording training would require an estimate of the noise component.
  • The near-balanced choices $\alpha=0.5$ and $\alpha=1-\alpha-\beta$ in the hyperparameter search suggest that equal weighting between speech preservation and noise suppression may be a robust default for other architectures, a hypothesis the paper does not test.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper proposes a components loss (CL) for training mask-based single-channel speech enhancement networks. Two variants are introduced: 2CL, which linearly combines a filtered-speech-preservation term and a residual-noise-power term, and 3CL, which adds a third term that penalizes deviation of the normalized residual-noise spectrum shape from the original noise spectrum. The loss is evaluated with one CNN architecture on Grid speech mixed with CHiME-3 noise, comparing against MSE, a perceptual weighting filter loss (PW-FILT), and a PESQ-based loss (PW-PESQ). The authors report that the CL-trained networks, especially 3CL, achieve higher PESQ, POLQA, SSDR, and ΔSNR on both seen and unseen noise types, with code provided.

Significance. If the empirical claims are reliable, the components loss is a practically useful, differentiable training objective that offers separate control over speech preservation, noise suppression, and residual noise naturalness, and it is not tied to a particular network architecture. Strengths include the coherent and differentiable loss formulation, the use of external metrics (PESQ, POLQA, STOI) for the headline claims, hyperparameter selection on validation rather than test data, and public code. The main significance risk is that the quantitative conclusions rest on a single evaluation with a small number of test speakers and no statistical uncertainty assessment, plus a confounded comparison between 2CL and 3CL.

major comments (3)
  1. [Section IV.A; Tables IV.a, IV.b, V.a, V.b] The central quantitative claims are based on a single run over a test set of only four Grid speakers (two male, two female) with no confidence intervals, bootstrap estimates, or significance tests. PESQ, POLQA, and STOI are known to be speaker- and utterance-dependent, so the reported differences (for example, roughly 0.25 PESQ for 2CL versus MSE on PED noise in Table IV.a, or about 0.2 PESQ for 3CL on unseen BUS noise in Table V.a) may not be statistically reliable. I request repeated training runs with different seeds and/or bootstrapping across speakers and utterances, together with paired significance tests for the headline metrics, and a statement of variability for every number that supports the abstract and conclusion.
  2. [Section V.B.1; Tables IV and V] The comparison between 2CL and 3CL confounds the effect of the third loss term with a change in the weighting of the first two terms: 2CL is evaluated at α=0.5 (speech weight 0.5, noise weight 0.5), while 3CL is evaluated at α=0.1, β=0.8 (speech weight 0.1, noise weight 0.1). The observed differences in ΔSNR, WLAKR, and PESQ between 2CL and 3CL therefore cannot be attributed to the third term alone. I recommend an ablation, for example comparing 2CL at α=0.1 with 3CL at α=0.1, β=0.8, or comparing 3CL with β=0 against 3CL with the same α and β>0, to isolate the contribution of the residual-noise-shape term before drawing the mechanistic conclusion that the third term improves balanced performance.
  3. [Introduction and Section VI] The paper claims that the components loss is "not restricted to any specific network topology or application," but the experiments use exactly one CNN architecture, one corpus, and one mask-estimation framework (magnitude masking with noisy phase). If the authors wish to retain the generality claim, they should either provide evidence with at least one different architecture or task, or explicitly narrow the conclusion to the tested setting; otherwise the broad phrasing in the introduction and conclusion overstates the evidentiary basis.
minor comments (2)
  1. [Figure 6 caption] The caption says the markers correspond to six SNR conditions "from 20 dB to 5 dB with a step size of −5 dB," which is inconsistent with the captions of Figures 4 and 5; it should read "from 20 dB to −5 dB."
  2. [Section V.A] The hyperparameter selection procedure states that columns with any measure at or below the baseline MSE are discarded, and then the "best performing" remaining setting is selected, but the multi-metric criterion for this final selection is not formally defined; specifying the exact ordering or scoring rule would make the selection reproducible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the components loss is a training objective validated on held-out external metrics.

full rationale

The proposed components loss is a training objective, not a fitted predictor. The paper optimizes J_2CL and J_3CL on training data (Eqs. 5 and 6) and evaluates on held-out test speakers and noise conditions using external metrics including PESQ, POLQA, STOI, SSDR, and Delta SNR. Hyperparameters alpha, beta, and lambda are selected on a 12.5% validation subset (Section V.A, Tables I-III), not on the test set, so the reported test-table gains do not reduce to fitted values. The loss terms are admittedly aligned with component metrics: Eq. (5) includes filtered-speech MSE and filtered-noise power, and Eq. (6) adds a normalized-noise-spectrum term; the paper openly attributes the Delta SNR and noise-quality improvements to these terms in Section V.B.1. This transparency is a mechanistic explanation of an empirical comparison, not circularity: the central headline claims on PESQ(hat s), POLQA(hat s), and STOI are external and not contained in the loss. Several citations are to the authors' own prior work ([27], [40], [45]), but they are used as architecture, baseline, and white-box references rather than as load-bearing evidence for the new loss. No step reduces by construction to its inputs, and no uniqueness theorem or ansatz is imported from a self-citation. The statistical robustness concern about the small test set is a correctness issue, not a circularity issue.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the standard additive noise assumption and on the availability of both components during training. No new physical or architectural entities are introduced. The only ad hoc element is the 3CL's normalized noise-shape term, whose perceptual benefit is asserted on empirical grounds.

free parameters (5)
  • alpha (2CL) = 0.5
    Trade-off between speech component quality and noise suppression in Eq. (5), selected on 12.5% of the validation set (Table I) as the most balanced setting after discarding settings with any metric below MSE.
  • alpha (3CL) = 0.1
    Noise attenuation weight in Eq. (6), selected on 12.5% of the validation set (Table II) from settings satisfying alpha = 1 - alpha - beta, choosing the best balanced PESQ, POLQA, and STOI.
  • beta (3CL) = 0.8
    Residual noise quality weight in Eq. (6), selected jointly with alpha on the validation set (Table II).
  • lambda1 (PW-PESQ baseline) = 0.2
    MSE weight for the PW-PESQ baseline loss in Eq. (12), optimized on validation for a fair comparison; not part of the proposed method.
  • lambda2 (PW-PESQ baseline) = 0.8
    PESQ-loss weight for the PW-PESQ baseline, optimized on validation; not part of the proposed method.
assumptions (4)
  • domain assumption Additive single-channel noise model y(n) = s(n) + d(n).
    Assumed throughout, e.g., Section II.A; the loss decomposition relies on the mixture being the sum of speech and noise.
  • domain assumption Clean speech and noise component spectra are available during training.
    Required to compute the filtered components in Eqs. (3)-(4) and the loss terms; stated in Section III.B.
  • domain assumption Minimizing separate errors for speech component, residual noise power, and residual noise spectral shape yields perceptually better enhancement.
    This is the design premise of the white-box approach (Section III.A); the paper acknowledges masking effects are not exploited and relies on PESQ and POLQA to validate.
  • ad hoc to paper Preserving the normalized residual noise spectrum shape improves naturalness of residual noise.
    This hypothesis motivates the third term in Eq. (6); it is not derived from psychoacoustics and is only supported by the reported WLAKR and PESQ results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Components Loss for Neural Networks in Mask-Based Speech Enhancement." pith.science (2026). https://pith.science/paper/ZWXSSPOL

@misc{pith2026190805087,
  author       = {Pith},
  title        = {Pith review of: Components Loss for Neural Networks in Mask-Based Speech Enhancement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZWXSSPOL}},
  note         = {Machine review of arXiv:1908.05087}
}
read the original abstract

Estimating time-frequency domain masks for single-channel speech enhancement using deep learning methods has recently become a popular research field with promising results. In this paper, we propose a novel components loss (CL) for the training of neural networks for mask-based speech enhancement. During the training process, the proposed CL offers separate control over preservation of the speech component quality, suppression of the residual noise component, and preservation of a naturally sounding residual noise component. We illustrate the potential of the proposed CL by evaluating a standard convolutional neural network (CNN) for mask-based speech enhancement. The new CL obtains a better and more balanced performance in almost all employed instrumental quality metrics over the baseline losses, the latter comprising the conventional mean squared error (MSE) loss and also auditory-related loss functions, such as the perceptual evaluation of speech quality (PESQ) loss and the recently proposed perceptual weighting filter loss. Particularly, applying the CL offers better speech component quality, better overall enhanced speech perceptual quality, as well as a more naturally sounding residual noise. On average, an at least 0.1 points higher PESQ score on the enhanced speech is obtained while also obtaining a higher SNR improvement by more than 0.5 dB, for seen noise types. This improvement is stronger for unseen noise types, where an about 0.2 points higher PESQ score on the enhanced speech is obtained, while also the output SNR is ahead by more than 0.5 dB. The new proposed CL is easy to implement and code is provided at https://github.com/ifnspaml/Components-Loss.

Figures

Figures reproduced from arXiv: 1908.05087 by the authors.

Figure 1
Figure 1. The NORM box in Fig. 1 represents a zero-mean and [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 1
Figure 1. Schematic of the mask-based CNN for speech spec￾trum enhancement, used for both the baseline CNN (baseline losses) and the new CNN (components loss). Details of the CNN can be seen in [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Topology details of the employed CNN in [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3: Proposed CNN training setup for speech enhancement according to the white-box approach. The hereby applied components loss (CL) is given in (5) and (6). A. White-Box Approach Since our work is inspired by the so-called white-box approach ([39], see also [40], [41]), we…
Figure 5
Figure 5. Figure 5: Noise attenuation (NAseg) vs. speech component qual￾ity (PESQ(˜s)) for different parameters α and β for the new 3CL (6) on 12.5% of the validation set. From top to bottom, the markers are corresponding to six SNR conditions from 20 dB to −5 dB with a step size of 5 dB.…
Figure 6
Figure 6. Figure 6: Noise attenuation (NAseg) vs. speech component qual￾ity (PESQ(˜s)) for different parameters λ1 and λ2 for the baseline PW-PESQ (12) on 12.5% of the validation set. From top to bottom, the markers are corresponding to six SNR conditions from 20 dB to 5 dB with a step si…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 54 canonical work pages

  1. [1]

    Speech Enhancement Using a Mini mum Mean-Square Error Short-Time Spectral Amplitude Estimato r,

    Y . Ephraim and D. Malah, “Speech Enhancement Using a Mini mum Mean-Square Error Short-Time Spectral Amplitude Estimato r,” IEEE T-ASSP, vol. 32, no. 6, pp. 1109–1121, Dec. 1984

  2. [2]

    Speech Enhancement Using a Minimum Mean-Square Err or Log- Spectral Amplitude Estimator,

    ——, “Speech Enhancement Using a Minimum Mean-Square Err or Log- Spectral Amplitude Estimator,” IEEE T-ASSP , vol. 33, no. 2, pp. 443– 445, Apr. 1985

  3. [3]

    Speech Enhancement Based on A Priori Signal to Noise Estimation,

    P . Scalart and J. V . Filho, “Speech Enhancement Based on A Priori Signal to Noise Estimation,” in Proc. of ICASSP , Atlanta, GA, USA, May 1996, pp. 629–632

  4. [4]

    Speech Enhancement by MAP Spectra l Am- plitude Estimation Using a Super-Gaussian Speech Model,

    T. Lotter and P . Vary, “Speech Enhancement by MAP Spectra l Am- plitude Estimation Using a Super-Gaussian Speech Model,” EURASIP Journal on Applied Signal Processing , vol. 2005, no. 7, pp. 1110–1126, May 2005

  5. [5]

    Speech Enhancement Using a Joint Map Estimator with Gaussian Mixture Model for (Non)-Stationar y Noise,

    B. Fodor and T. Fingscheidt, “Speech Enhancement Using a Joint Map Estimator with Gaussian Mixture Model for (Non)-Stationar y Noise,” in Proc. of ICASSP , Prague, Czech Republic, May 2011, pp. 4768–4771

  6. [6]

    Speech Enhancement Using Super-Gaussian Spe ech Models and Noncausal A Priori SNR Estimation,

    I. Cohen, “Speech Enhancement Using Super-Gaussian Spe ech Models and Noncausal A Priori SNR Estimation,” Speech Commun. , vol. 47, no. 3, pp. 336–350, Nov. 2005

  7. [7]

    Improved A Po steriori Speech Presence Probability Estimation Based on a Likeliho od Ratio with Fixed Priors,

    T. Gerkmann, C. Breithaupt, and R. Martin, “Improved A Po steriori Speech Presence Probability Estimation Based on a Likeliho od Ratio with Fixed Priors,” IEEE T-ASLP , vol. 16, no. 5, pp. 910–919, Jul. 2008

  8. [8]

    A Data-Driven Ap proach to A Priori SNR Estimation,

    S. Suhadi, C. Last, and T. Fingscheidt, “A Data-Driven Ap proach to A Priori SNR Estimation,” IEEE T-ASLP, vol. 19, no. 1, pp. 186–195, Jan. 2011

Show all 54 references
  1. [9]

    An Iterative Speech Model-Based A Priori SNR Estimator,

    S. Elshamy, N. Madhu, W. J. Tirry, and T. Fingscheidt, “An Iterative Speech Model-Based A Priori SNR Estimator,” in Proc. of Interspeech , Dresden, Germany, Sep. 2015, pp. 1740–1744

  2. [10]

    Ins tantaneous A Priori SNR Estimation by Cepstral Excitation Manipulation ,

    S. Elshamy, N. Madhu, W. Tirry, and T. Fingscheidt, “Ins tantaneous A Priori SNR Estimation by Cepstral Excitation Manipulation ,” IEEE/ACM T-ASLP, vol. 25, no. 8, pp. 1592–1605, Aug. 2017

  3. [11]

    Tracking Speech- presence Uncertainty to Improve Speech Enhancement in Non-stationa ry Noise Environments,

    D. Malah, R. V . Cox, and A. J. Accardi, “Tracking Speech- presence Uncertainty to Improve Speech Enhancement in Non-stationa ry Noise Environments,” in Proc. of ICASSP , Phoenix, AZ, USA, Mar. 1999, pp. 789–792

  4. [12]

    Data-Driven Speech Enha ncement,

    T. Fingscheidt and S. Suhadi, “Data-Driven Speech Enha ncement,” in Proc. of ITG Conf. on Speech Communication , Kiel, Germany, Apr. 2006, pp. 1–4

  5. [13]

    Environment-O ptimized Speech Enhancement,

    T. Fingscheidt, S. Suhadi, and S. Stan, “Environment-O ptimized Speech Enhancement,” IEEE T-ASLP, vol. 16, no. 4, pp. 825–834, May 2008

  6. [14]

    A general Opti mization Procedure for Spectral Speech Enhancement Methods,

    J. Erkelens, J. Jensen, and R. Heusdens, “A general Opti mization Procedure for Spectral Speech Enhancement Methods,” in Proc. of EUSIPCO, Florence, Italy, Sep. 2006, pp. 1–5

  7. [15]

    A Data-Driven Approach to Optimizing Spectral Spe ech En- hancement Methods for V arious Error Criteria,

    ——, “A Data-Driven Approach to Optimizing Spectral Spe ech En- hancement Methods for V arious Error Criteria,” Speech Communication, vol. 49, no. 7-8, pp. 530–541, Jul. 2007

  8. [16]

    On Training Targe ts for Supervised Speech Separation,

    Y . Wang, A. Narayanan, and D. L. Wang, “On Training Targe ts for Supervised Speech Separation,” IEEE/ACM T-ASLP , vol. 22, no. 12, pp. 1849–1858, Dec. 2014

  9. [17]

    Discrimina- tively Trained Recurrent Neural Networks for Single-Chann el Speech Separation,

    F. Weninger, J. R. Hershey, J. Le Roux, and B. Schuller, “ Discrimina- tively Trained Recurrent Neural Networks for Single-Chann el Speech Separation,” in Proc. of 2nd IEEE GlobalSIP , Atlanta, GA, USA, May 2014, pp. 577–581

  10. [18]

    Dee p Learning for Monaural Speech Separation,

    P . S. Huang, M. Kim, M. H. Johnson, and P . Smaragdis, “Dee p Learning for Monaural Speech Separation,” in Proc. of ICASSP , Florence, Italy, May 2014, pp. 1562–1566

  11. [19]

    P hase- Sensitive and Recognition-Boosted Speech Separation Usin g Deep Recurrent Neural Networks,

    H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “P hase- Sensitive and Recognition-Boosted Speech Separation Usin g Deep Recurrent Neural Networks,” in Proc. of ICASSP , Brisbane, QLD, Australia, Aug. 2015, pp. 708–712

  12. [20]

    A Deep Neural Network for Time-Do main Signal Reconstruction,

    Y . Wang and D. L. Wang, “A Deep Neural Network for Time-Do main Signal Reconstruction,” in Proc. of ICASSP , Brisbane, QLD, Australia, Aug. 2015, pp. 4390–4394

  13. [21]

    Complex Ratio Masking for Monaural Speech Separation,

    D. S. Williamson, Y . Wang, and D. L. Wang, “Complex Ratio Masking for Monaural Speech Separation,” IEEE/ACM T-ASLP , vol. 24, no. 3, pp. 483–492, Mar. 2016

  14. [22]

    A New Ratio Mask Representatio n for CASA-Based Speech Enhancement,

    F. Bao and W. H. Abdulla, “A New Ratio Mask Representatio n for CASA-Based Speech Enhancement,” IEEE/ACM TASLP, vol. 27, no. 1, pp. 7–19, Jan. 2019

  15. [23]

    Supervised Speech Separation Based on Deep Learning: An Overview,

    D. L. Wang and J. T. Chen, “Supervised Speech Separation Based on Deep Learning: An Overview,” IEEE/ACM T-ASLP, vol. 26, no. 10, pp. 1702–1726, Oct. 2018

  16. [24]

    A Regression Approa ch to Single-Channel Speech Separation via High-Resolution Dee p Neural Networks,

    J. Du, Y . Tu, L. R. Dai, and C. H. Lee, “A Regression Approa ch to Single-Channel Speech Separation via High-Resolution Dee p Neural Networks,” IEEE/ACM T-ASLP , vol. 24, no. 8, pp. 1424–1437, Apr. 2016

  17. [25]

    Perception Optimi zed Deep De- noising Autoencoders for Speech Enhancement

    P . G. Shivakumar and P . G. Georgiou, “Perception Optimi zed Deep De- noising Autoencoders for Speech Enhancement.” in Proc. of Interspeech, San Francisco, CA, USA, Sep. 2016, pp. 3743–3747

  18. [26]

    A Percep tually- Weighted Deep Neural Network for Monaural Speech Enhanceme nt in Various Background Noise Conditions,

    Q. J. Liu, W. Wang, P . J. B. Jackson, and Y . Tang, “A Percep tually- Weighted Deep Neural Network for Monaural Speech Enhanceme nt in Various Background Noise Conditions,” in Proc. of EUSIPCO , Kos, Greece, Aug. 2017, pp. 1270–1274

  19. [27]

    A Perceptual W eighting Filter Loss for DNN Training in Speech Enhancement,

    Z. Zhao, S. Elshamy, and T. Fingscheidt, “A Perceptual W eighting Filter Loss for DNN Training in Speech Enhancement,” arXiv preprint arXiv:1905.09754, May 2019

  20. [28]

    Lear ning to Dequantize Speech Signals by Primal-Dual Networks: An Appr oach for Acoustic Sensor Networks,

    C. Brauer, Z. Zhao, D. Lorenz, and T. Fingscheidt, “Lear ning to Dequantize Speech Signals by Primal-Dual Networks: An Appr oach for Acoustic Sensor Networks,” in Proc. of ICASSP , Brighton, UK, May 2019, pp. 7000–7004

  21. [29]

    A Deep Learning Loss Function Based on the Perceptual Evalu ation of the Speech Quality,

    J. M. Mart´ ın Do˜ nas, A. M. Gomez, J. A. Gonzalez, and A. M . Peinado, “A Deep Learning Loss Function Based on the Perceptual Evalu ation of the Speech Quality,” IEEE SPL, vol. 25, no. 11, pp. 1680–1684, Nov. 2018

  22. [30]

    DNN- Based Source Enhancement to Increase Objective Sound Quali ty As- sessment Score,

    Y . Koizumi, K. Niwa, Y . Hioka, K. Kobayashi, and Y . Haned a, “DNN- Based Source Enhancement to Increase Objective Sound Quali ty As- sessment Score,” IEEE/ACM T-ASLP , vol. 26, no. 10, pp. 1780–1792, Oct. 2018

  23. [31]

    Monaural Speech En hancement Using Deep Neural Networks by Maximizing a Short-Time Objec tive Intelligibility Measure,

    M. Kolbcek, Z. H. Tan, and J. Jensen, “Monaural Speech En hancement Using Deep Neural Networks by Maximizing a Short-Time Objec tive Intelligibility Measure,” in Proc. of ICASSP , Calgary, AB, Canada, Apr. 2018, pp. 5059–5063

  24. [32]

    Deep Neural Network Based Speech Separation Optimizing an Objective Es timator of Intelligibility for Low Latency Applications,

    G. Naithani, J. Nikunen, L. Bramslow, and T. Virtanen, “ Deep Neural Network Based Speech Separation Optimizing an Objective Es timator of Intelligibility for Low Latency Applications,” in Proc. of IWAENC , Tokyo, Japan, Sep. 2018, pp. 386–390. 12

  25. [33]

    Training Supervise d Speech Separation System to Improve STOI and PESQ Directly,

    H. Zhang, X. L. Zhang, and G. L. Gao, “Training Supervise d Speech Separation System to Improve STOI and PESQ Directly,” in Proc. of ICASSP, Calgary, AB, Canada, Apr. 2018, pp. 5374–5378

  26. [34]

    End-to- End Waveform Utterance Enhancement for Direct Evaluation M etrics Optimization by Fully Convolutional Neural Networks,

    S. W. Fu, T. W. Wang, Y . Tsao, X. Lu, and H. Kawai, “End-to- End Waveform Utterance Enhancement for Direct Evaluation M etrics Optimization by Fully Convolutional Neural Networks,” IEEE/ACM T- ASLP, vol. 26, no. 9, pp. 1570–1584, Sep. 2018

  27. [35]

    Error Concealment by Softb it Speech De- coding,

    T. Fingscheidt and P . Vary, “Error Concealment by Softb it Speech De- coding,” in Proc. of ITG-Fachtagung ”Sprachkommunikation” , Frankfurt a.M., Germany, Sep. 1996, pp. 7–10

  28. [36]

    Softbit Speech Decoding: A New Approach to Error Co nceal- ment,

    ——, “Softbit Speech Decoding: A New Approach to Error Co nceal- ment,” IEEE T-SAP, vol. 9, no. 3, pp. 240–251, Mar. 2001

  29. [37]

    A Short-Time Objective Intelligibility Measure for Time-Frequency Wei ghted Noisy Speech,

    C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A Short-Time Objective Intelligibility Measure for Time-Frequency Wei ghted Noisy Speech,” in Proc. of ICASSP , Dallas, TX, USA, Jun. 2010, pp. 4214– 4217

  30. [38]

    ITU, Rec. P.862: Perceptual Evaluation of Speech Quality (PESQ) : An Objective Method for End-To-End Speech Quality Assessme nt of Narrow-Band Telephone Networks and Speech Codecs , International Telecommunication Standardization Sector (ITU-T), Feb. 2 001

  31. [39]

    On the Optimizat ion of Speech Enhancement Systems Using Instrumental Measures,

    S. Gustafsson, R. Martin, and P . Vary, “On the Optimizat ion of Speech Enhancement Systems Using Instrumental Measures,” in Proc. of W orkshop on Qual. Assess. in Speech, Audio, and Image Comm un., Darmstadt, Germany, Mar. 1996, pp. 36–40

  32. [40]

    Quality Assessment of Sp eech Enhance- ment Systems by Separation of Enhanced Speech, Noise, and Ec ho,

    T. Fingscheidt and S. Suhadi, “Quality Assessment of Sp eech Enhance- ment Systems by Separation of Enhanced Speech, Noise, and Ec ho,” in Proc. of Interspeech , Antwerp, Belgium, Aug. 2007, pp. 818–821

  33. [41]

    A Figure of Merit for Instrume ntal Optimiza- tion of Noise Reduction Algorithms,

    H. Yu and T. Fingscheidt, “A Figure of Merit for Instrume ntal Optimiza- tion of Noise Reduction Algorithms,” in Proc. of 5th Biennial W orkshop on DSP for In-V ehicle Systems , Kiel, Germany, Sep. 2011, pp. 1–8

  34. [42]

    P.1100: Narrowband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan

    ITU, Rec. P.1100: Narrowband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan. 2019

  35. [43]

    P.1110: Wideband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan

    ——, Rec. P.1110: Wideband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan. 2015

  36. [44]

    P.1130: Subsystem Requirements for Automotive Speech Services, International Telecommunication Standardization Secto r (ITU- T), Jun

    ——, Rec. P.1130: Subsystem Requirements for Automotive Speech Services, International Telecommunication Standardization Secto r (ITU- T), Jun. 2015

  37. [45]

    Convolutional N eural Networks to Enhance Coded Speech,

    Z. Zhao, H. J. Liu, and T. Fingscheidt, “Convolutional N eural Networks to Enhance Coded Speech,” IEEE/ACM T-ASLP, vol. 27, no. 4, pp. 663– 678, Apr. 2019

  38. [46]

    MetricGAN: Ge nerative Adversarial Networks Based Black-Box Metric Scores Optimi zation for Speech Enhancement,

    S. Z. Fu, C. F. Liao, Y . Tsao, and S. D. Lin, “MetricGAN: Ge nerative Adversarial Networks Based Black-Box Metric Scores Optimi zation for Speech Enhancement,” arXiv preprint arXiv:1905.04874 , May 2019

  39. [47]

    Residual Networ ks Behave Like Ensembles of Relatively Shallow Networks,

    A. Veit, M. J. Wilber, and S. Belongie, “Residual Networ ks Behave Like Ensembles of Relatively Shallow Networks,” in Proc. of NIPS , Barcelona, Spain, Dec. 2016, pp. 550–558

  40. [48]

    An Audi o-Visual Corpus for Speech Perception and Automatic Speech Recognit ion,

    M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An Audi o-Visual Corpus for Speech Perception and Automatic Speech Recognit ion,” The Journal of the Acoustical Society of America , vol. 120, no. 5, pp. 2421– 2424, Jun. 2006

  41. [49]

    The T hird ‘CHiME’ Speech Separation and Recognition Challenge: Dataset, Tas k and Base- lines,

    J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The T hird ‘CHiME’ Speech Separation and Recognition Challenge: Dataset, Tas k and Base- lines,” in Proc. of ASRU , Scottsdale, AZ, USA, Feb. 2015, pp. 504–511

  42. [50]

    P.56: Objective Measurement of Active Speech Level , Interna- tional Telecommunication Standardization Sector (ITU-T) , Dec

    ITU, Rec. P.56: Objective Measurement of Active Speech Level , Interna- tional Telecommunication Standardization Sector (ITU-T) , Dec. 2011

  43. [51]

    ——, Rec. P .862.2: Corrigendum 1, Wideband Extension to Recomme n- dation P .862 for the Assessment of Wideband Telephone Netwo rks and Speech Codecs, International Telecommunication Standardization Secto r (ITU-T), Oct. 2017

  44. [52]

    P .863: Perceptual Objective Listening Quality Predic tion (POLQA), International Telecommunication Union, Telecommunicat ion Standardization Sector (ITU-T), Mar

    ——, Rec. P .863: Perceptual Objective Listening Quality Predic tion (POLQA), International Telecommunication Union, Telecommunicat ion Standardization Sector (ITU-T), Mar. 2018

  45. [53]

    Black Box Measurement of Musi cal Tones Produced by Noise Reduction Systems,

    H. Yu and T. Fingscheidt, “Black Box Measurement of Musi cal Tones Produced by Noise Reduction Systems,” in Proc. of ICASSP , Kyoto, Japan, Aug. 2012, pp. 4573–4576

  46. [54]

    14) , 3GPP; TSG SA, Mar

    3GPP, Mandatory Speech Codec Speech Processing Functions; Adapt ive Multi-Rate (AMR) Speech Codec; Transcoding Functions (3GP P TS 26.090, Rel. 14) , 3GPP; TSG SA, Mar. 2017

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.