REVIEW 2 major objections 5 minor 34 references
Soft-Constrained Spatially Selective Active Noise Control for Open-fitting Hearables
T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A single trade-off parameter connects full noise cancellation to distortion-free speech preservation in hearable ANC.
desk verdict A clean, correctly derived extension of SSANC with a tunable trade-off; the math is solid, but the practical claims outrun the single idealized simulation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the soft-constrained cost $\min_w E\{e^2(n)\}+\beta\|w\|^2+\mu\|H(q+Gw)-\delta_\Delta\|^2$, whose normal equations produce the time-domain filter (21). In the frequency domain the identical relaxation produces a regularized projection: the scalar $1/\mu$ sits inside the denominator of Eq. (27), so $\mu$ controls how much of the distortionless speech-preservation term is admixed into the pure noise-cancellation solution. The matrix $H$ is the convolution matrix of relative impulse responses (ReIRs), i.e., the transfer paths from each outer microphone to the reference microphone; $q+Gw$ represents the total path from input to the inner error microphone, so $H(q+Gw)-\delta_\Delta$ measures the speech deviation from the desired delayed reference, and $\mu$ converts that deviation from a hard equality constraint into a quadratic penalty.
What would settle it
Run the same open-fitting hearable simulation while deliberately corrupting the secondary-path estimate, for example by scaling $\hat{g}$ by 1.1 or delaying it by one sample, and check whether the intermediate operating point $\log_{10}\mu=-2$ still beats the hard-constrained baseline on SNR improvement. If it does not, the claimed practical advantage depends on exact secondary-path knowledge. Alternatively, replace the anechoic impulse responses with reverberant ones and test whether the reported PESQ and ESTOI gains at that $\mu$ persist.
Extended reading notes
Core claim
The central claim is that the optimal soft-constrained SSANC filter is an interpolation between two known filters. In the frequency domain the solution is $w_{\mathrm{soft}}(\omega)=\frac{1}{G^{*}(\omega)}\left[-q_{\omega}+\frac{\Phi_x^{-1}(\omega)h(\omega)e^{i\omega\Delta}}{1/\mu+h^{H}(\omega)\Phi_x^{-1}(\omega)h(\omega)}\right]$, so that $\lim_{\mu\to 0}w_{\mathrm{soft}}(\omega)=w_{\mathrm{ANC}}(\omega)$ and $\lim_{\mu\to\infty}w_{\mathrm{soft}}(\omega)=w_{\mathrm{hard}}(\omega)$. The same interpolation is expressed in the time domain by Eq. (21), where the regularized covariance $\Phi_{rr}+\mu G^{T}H^{T}HG$ appears. Because the distortionless constraint is relaxed rather than enforced, the optimizer is free to remove more of the leakage noise; simulations on measured anechoic impulse responses from an open-fitting earpiece confirm that for intermediate $\mu$ (e.g., $\log_{10}\mu=-2$) the system reaches 20.7 dB noise reduction at the cost of $-14.7$ dB intelligibility-weighted spectral distortion, yielding an SNR improvement of 17.2 dB, a PESQ gain of 0.54, and an ESTOI gain of 0.39, each larger than the hard-constrained baseline.
Load-bearing premise
The load-bearing premise is that the acoustic path from the loudspeaker to the inner error microphone (the secondary path) is known exactly, $\hat{g}=g$; if that estimate is wrong, the extracted leakage $\hat{b}_p(n)$ is biased, the equality $p(n)=q^T x(n)$ breaks down, and the optimized soft-constrained filter, like the hard-constrained one, no longer implements the intended trade-off between speech distortion and noise reduction.
Editorial extensions
If this is right
- Conventional ANC and hard-constrained SSANC are not rival designs but limiting cases of a single filter parameterized by $\mu$, so one implementation covers both operating modes.
- There is a broad intermediate range of $\mu$ in which SNR improvement, PESQ, and ESTOI all exceed the hard-constrained baseline, meaning zero speech distortion is not the best operating point for perceived quality in this setup.
- The soft constraint requires no extra microphones or hardware; it changes only the optimization criterion used to compute the control filter.
- The frequency-domain identity holds for any positive-definite input covariance and known secondary path, so the interpolation between conventional ANC and hard-constrained SSANC is not tied to the particular simulation setup.
Reading between the lines
- The paper fixes $\mu$ to a single frequency-independent value; a natural extension is to let $\mu$ vary across frequency bands, tuning the trade-off where speech intelligibility matters most and full cancellation elsewhere.
- The derivation assumes a perfectly known secondary path; because the soft penalty is not an equality constraint, the design may tolerate secondary-path mismatch better than the hard-constrained one, a robustness that could be tested by perturbing the path estimate.
- The same soft-constraint cost transfers directly to closed-fitting and open-ear ANC devices and to feedback controllers, since it needs only an inner error signal, a leakage estimate, and relative impulse responses.
- The reported gains come from an anechoic, stationary scenario; a reverberant or moving-talker test would show whether a single fixed $\mu$ generalizes across scenes or whether an adaptive $\mu$ scheduler is needed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a soft-constrained spatially selective active noise control (SSANC) formulation for open-fitting hearables. A single scalar parameter mu is added as a quadratic penalty on speech distortion, yielding a closed-form time-domain filter (Eq. (21)) and a frequency-domain filter (Eq. (27)). The authors show analytically that conventional ANC and hard-constrained SSANC are recovered as mu approaches 0 and infinity, respectively (Eq. (28)), and they report a simulation study with one anechoic scenario in which an intermediate mu yields improved SNR, PESQ, and ESTOI relative to the hard-constrained design.
Significance. The derivation is self-contained and the algebra checks out: the proposed filters are valid least-squares solutions of a well-defined objective, and the limiting cases follow directly from the equations rather than from an assumed outcome. The simulation uses realistic measured impulse responses for a KEMAR-based open-fitting hearable, and the data provenance (impulse-response database, VCTK, NOISEX-92) is described well enough to support reproducibility. The main value is a simple, interpretable trade-off parameter that unifies two existing designs. The practical claims, however, currently rest on an ideal secondary-path estimate and on a single simulation scenario, which limits the strength of the conclusions until those points are addressed.
major comments (2)
- [Section 2, Eq. (8)] The identity p(n)=q^T x(n), on which the entire derivation rests, assumes a perfect secondary-path estimate, i.e., \hat g = g, as stated after Eq. (6). When \hat g ≠ g, the estimated leakage becomes \hat p(n)=p(n)+(g-\hat g)^T y(n), so the stacked input vector x(n) is control-dependent, Eq. (8) no longer holds, and the soft-constraint penalty in Eq. (20) penalizes a biased response. In particular, the mu→∞ limit in Eq. (28) need not preserve the desired speech. Section 6.1 uses the measured impulse response as the secondary-path estimate, so the simulation does not probe this assumption, which is critical for open-fitting hearables whose secondary paths vary with insertion and anatomy. Please add a robustness study, e.g., re-running Figure 3 with perturbed \hat g (gain error, delay error, or re-insertion impulse responses) and reporting the resulting NR, SD, SNR, PESQ, and ESTOI curves, or explicitly restrict the practical claims to perfectly known secondary paths.
- [Section 6.3, Fig. 3] The conclusion that a 'broad range' of mu provides substantial improvements over the hard-constrained design is supported by only a single anechoic scenario: one speech source, two noise sources, one input SNR (−5 dB), one set of relative impulse responses, and no error bars or multiple realizations. This evidence is insufficient to establish the generalizability of the practical claim. Please add variations in source positions, input SNR, reverberation, and repeated hearable insertions, or temper the conclusions to the specific configuration tested.
minor comments (5)
- [Section 6.1] The sentence 'As the secondary path estimate we used the measured impulse response between the outer receiver and the inner error microphone' needs a comma after 'estimate'; also, 'the outer receiver at the right ear as the secondary source' could be phrased more precisely as 'the right-ear outer receiver served as the secondary source.'
- [Section 5, Eqs. (27)-(28)] The notation is inconsistent: h appears without its frequency argument in the denominator terms h^H Φ_x^{-1}(ω) h, while h(ω) is used elsewhere. Please harmonize the notation throughout Section 5.
- [Eq. (16c)] The definition of δ_Δ uses underbraced length expressions that are difficult to parse; rewriting with explicit index positions (e.g., 1 at index L_a+Δ) would improve clarity.
- [Section 6.2 and Fig. 3] The intelligibility-weighted spectral distortion SD_intellig is reported in negative dB, with more negative values meaning less distortion; this sign convention should be stated explicitly in Section 6.2 or in the Figure 3 caption to avoid ambiguity.
- [Section 6.3] No confidence intervals or sensitivity analyses are provided for the reported metrics; even in a single-scenario study, rerunning with a different speech utterance or noise realization would strengthen the claim that the observed improvements are not sample-specific.
Circularity Check
No circularity: the soft-constrained filter follows directly from the stated objective, and the claimed limiting cases are algebraic limits of the derived expression, not fitted predictions or self-citational premises.
full rationale
The central derivation is self-contained. The proposed filter wsoft in Eq. (21) is the minimizer of the explicitly stated optimization problem in Eq. (20), which augments the conventional ANC objective with the quadratic penalty μ||H(q+Gw)−δΔ||^2. The frequency-domain solution wsoft(ω) in Eq. (27) is obtained from the corresponding quadratic objective in Eq. (26). Equation (28) then evaluates the limits of this single expression: as μ→0 the penalty term vanishes and the solution reduces to the conventional ANC solution wANC(ω)=−qω/G*(ω), while as μ→∞ the term 1/μ→0 and the solution reduces to the hard-constrained form in Eq. (25). These are direct algebraic consequences of Eq. (27), not empirical predictions or renamed inputs. The paper does rely on the authors' earlier hard-constrained SSANC work ([8], [20], [21]) for the relative-impulse-response constraint H(q+Gw)=δΔ, but those equations are restated in the paper, and the soft-constrained extension is a genuine generalization whose limiting behavior is derived here. The explicit assumption of a perfect secondary-path estimate (bg=g after Eq. (6)) is a real robustness limitation, but it is a modeling assumption, not a circular step: the derivations do not assume the conclusions they claim to establish. No load-bearing argument reduces to a self-citation, and no fitted parameter is renamed as a prediction, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (4)
- mu (trade-off parameter) =
swept, log10 mu from -5 to 0.5; highlighted mu = 0.01
- beta (control filter regularization) =
lambda_max(G^T Phi_xx G) / (4 x 10^5)
- rho (matrix inversion regularization) =
lambda_max(H G Phi_rr^{-1} G^T H^T) / (4 x 10^5)
- Delay and filter lengths =
Delta = 32 samples, Lw = 600, Lg = 280, La = 22, Lh = 262
assumptions (6)
- domain assumption Perfect secondary path estimate bg = g
- domain assumption Acoustic feedback paths from loudspeaker to outer microphones are known and perfectly canceled
- domain assumption Relative impulse responses (ReIRs) for the desired speech source are known or accurately estimated
- standard math Standard convex quadratic optimization and matrix inversion lemma
- standard math Input covariance matrix Phi_x(omega) is positive definite
- domain assumption Anechoic environment with measured impulse responses
Cite this review
Pith. "Pith review of Soft-Constrained Spatially Selective Active Noise Control for Open-fitting Hearables." pith.science (2026). https://pith.science/paper/EEJH3YD4
@misc{pith2026250712122,
author = {Pith},
title = {Pith review of: Soft-Constrained Spatially Selective Active Noise Control for Open-fitting Hearables},
year = {2026},
howpublished = {\url{https://pith.science/paper/EEJH3YD4}},
note = {Machine review of arXiv:2507.12122}
}
read the original abstract
Recent advances in spatially selective active noise control (SSANC) using multiple microphones have enabled hearables to suppress undesired noise while preserving desired speech from a specific direction. Aiming to achieve minimal speech distortion, a hard constraint has been used in previous work in the optimization problem to compute the control filter. In this work, we propose a soft-constrained SSANC system that uses a frequency-independent parameter to trade off between speech distortion and noise reduction. We derive both time- and frequency-domain formulations, and show that conventional active noise control and hard-constrained SSANC represent two limiting cases of the proposed design. We evaluate the system through simulations using a pair of open-fitting hearables in an anechoic environment with one speech source and two noise sources. The simulation results validate the theoretical derivations and demonstrate that for a broad range of the trade-off parameter, the signal-to-noise ratio and the speech quality and intelligibility in terms of PESQ and ESTOI can be substantially improved compared to the hard-constrained design.
Reference graph
Works this paper leans on
-
[14]
——, “A speech distortion weighting based approach to integrated active noise control and noise reduction in hearing aids,” Signal Processing , vol. 93, no. 9, pp. 2440–2452, 2013
work page 2013
-
[1]
S. J. Elliott, Signal processing for active control . Academic Press, 2000
work page 2000
- [2]
-
[3]
Recent advances on active noise control: open issues and innovative applications,
Y . Kajikawa, W.-S. Gan, and S. M. Kuo, “Recent advances on active noise control: open issues and innovative applications,” APSIPA Transactions on Signal and Information Processing , vol. 1, e3, pp. 1–21, 2012
work page 2012
-
[4]
Listening in a noisy environment: Integration of active noise control in audio products,
C.-Y . Chang, A. Siswanto, C.-Y . Ho, T.-K. Yeh, Y .-R. Chen, and S. M. Kuo, “Listening in a noisy environment: Integration of active noise control in audio products,” IEEE Consumer Electronics Magazine , vol. 5, no. 4, pp. 34–43, 2016
work page 2016
-
[5]
Augmented/mixed reality audio for hearables: sensing, control, and rendering,
R. Gupta, J. He, R. Ranjan, W.-S. Gan, F. Klein, C. Schneiderwind, A. Neidhardt, K. Brandenburg, and V . V ¨alim¨aki, “Augmented/mixed reality audio for hearables: sensing, control, and rendering,” IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 63–89, 2022
work page 2022
-
[6]
Integrated active noise control and noise reduction in hearing aids,
R. Serizel, M. Moonen, J. Wouters, and S. H. Jensen, “Integrated active noise control and noise reduction in hearing aids,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 6, pp. 1137–1146, 2010
work page 2010
-
[7]
V . Patel, J. Cheer, and S. Fontana, “Design and implementation of an active noise control headphone with directional hear-through capability,” IEEE Transactions on Consumer Electronics , vol. 66, no. 1, pp. 32–40, Feb. 2020
work page 2020
Show all 34 references
-
[8]
Spatially selective active noise control systems,
T. Xiao, B. Xu, and C. Zhao, “Spatially selective active noise control systems,” The Journal of the Acoustical Society of America , vol. 153, no. 5, pp. 2733–2744, May 2023
2023
-
[9]
Optimization of a fixed virtual sensing feedback ANC controller for in-ear headphones with multiple loudspeakers,
P. R. Benois, R. Roden, M. Blau, and S. Doclo, “Optimization of a fixed virtual sensing feedback ANC controller for in-ear headphones with multiple loudspeakers,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 8717–8721
2022
-
[10]
Beamforming: a versatile approach to spatial filtering,
B. Van Veen and K. Buckley, “Beamforming: a versatile approach to spatial filtering,” IEEE ASSP Magazine , vol. 5, no. 2, pp. 4–24, 1988
1988
-
[11]
Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diversity using multiple microphones,
S. Doclo, W. Kellermann, S. Makino, and S. E. Nordholm, “Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diversity using multiple microphones,” IEEE Signal Processing Magazine, vol. 32, no. 2, pp. 18–30, Mar. 2015
2015
-
[12]
A consoli- dated perspective on multimicrophone speech enhancement and source separation,
S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A consoli- dated perspective on multimicrophone speech enhancement and source separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, no. 4, pp. 692–730, 2017
2017
-
[13]
Output SNR analysis of integrated active noise control and noise reduction in hearing aids under a single speech source scenario,
R. Serizel, M. Moonen, J. Wouters, and S. H. Jensen, “Output SNR analysis of integrated active noise control and noise reduction in hearing aids under a single speech source scenario,” Signal Processing, vol. 91, no. 8, pp. 1719–1729, 2011
2011
-
[15]
Combined feedforward-feedback noise reduction schemes for open-fitting hearing aids,
D. Dalga and S. Doclo, “Combined feedforward-feedback noise reduction schemes for open-fitting hearing aids,” in Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , New Paltz, NY , USA, 2011, pp. 185–188
2011
-
[16]
Theoretical performance analysis of ANC-motivated noise re- duction algorithms for open-fitting hearing aids,
——, “Theoretical performance analysis of ANC-motivated noise re- duction algorithms for open-fitting hearing aids,” in Proc. International Workshop on Acoustic Signal Enhancement (IWAENC) , Aachen, Germany, 2012, pp. 1–4
2012
-
[17]
Integrated active noise control for open-fit hearing aids with customized filter,
C.-Y . Ho, K.-K. Shyu, C.-Y . Chang, and S. M. Kuo, “Integrated active noise control for open-fit hearing aids with customized filter,” Applied Acoustics, vol. 137, pp. 1–8, 2018
2018
-
[18]
A time- domain multi-channel directional active noise control system,
H. Zhang, J. Zhang, F. Ma, P. N. Samarasinghe, and H. Sun, “A time- domain multi-channel directional active noise control system,” in Proc. European Signal Processing Conference (EUSIPCO) , Helsinki, Finland, 2023, pp. 376–380
2023
-
[19]
A spherical-harmonic domain selective spatial active noise control system based on sound field reproduction,
H. Zhang, H. J. Sun, J. A. Zhang, P. Samarasinghe, and Y . A. Zhang, “A spherical-harmonic domain selective spatial active noise control system based on sound field reproduction,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Hyde...
2025
-
[20]
Effect of target signals and delays on spatially selective active noise control for open-fitting hearables,
T. Xiao and S. Doclo, “Effect of target signals and delays on spatially selective active noise control for open-fitting hearables,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea, 2024, pp. 1056–1060
2024
-
[21]
Spatially selective active noise control for open-fitting hearables with acausal optimization,
——, “Spatially selective active noise control for open-fitting hearables with acausal optimization,” in Proc. F orum Acusticum Euronoise 2025 , M´alaga, Spain, Jun. 2025
2025
-
[22]
GSVD-based optimal filtering for single and multimicrophone speech enhancement,
S. Doclo and M. Moonen, “GSVD-based optimal filtering for single and multimicrophone speech enhancement,” IEEE Transactions on Signal Processing, vol. 50, no. 9, pp. 2230–2244, 2002
2002
-
[23]
Spatially pre-processed speech distortion weighted multi-channel Wiener filtering for noise reduction,
A. Spriet, M. Moonen, and J. Wouters, “Spatially pre-processed speech distortion weighted multi-channel Wiener filtering for noise reduction,” Signal Processing, vol. 84, no. 12, pp. 2367–2387, 2004
2004
-
[24]
Adaptive microphone array employing spatial quadratic soft constraints and spectral shaping,
S. Nordholm, H. Q. Dam, N. Grbic, and S. Y . Low, “Adaptive microphone array employing spatial quadratic soft constraints and spectral shaping,” Signals and Communication Technology , pp. 229–246, 2005
2005
-
[25]
Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction,
S. Doclo, A. Spriet, J. Wouters, and M. Moonen, “Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction,” Speech Communication , vol. 49, no. 7, pp. 636–656, 2007
2007
-
[26]
Residual noise control using a parametric multichannel Wiener filter,
S. Braun, K. Kowalczyk, and E. A. P. Habets, “Residual noise control using a parametric multichannel Wiener filter,” in Proc. IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP) , South Brisbane, QLD, Australia, 2015, pp. 360–364
2015
-
[27]
Acoustic feedback suppression for multi-microphone hearing devices using a soft-constrained null- steering beamformer,
H. Schepker, S. Nordholm, and S. Doclo, “Acoustic feedback suppression for multi-microphone hearing devices using a soft-constrained null- steering beamformer,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 929–940, 2020
2020
-
[28]
A one-size-fits-all earpiece with multiple microphones and drivers for hearing device research,
F. Denk, M. Lettau, H. Schepker, S. Doclo, R. Roden, M. Blau, J.-H. Bach, J. Wellmann, and B. Kollmeier, “A one-size-fits-all earpiece with multiple microphones and drivers for hearing device research,” in Proc. AES International Conference on Headphone Technology , San Franci...
2019
-
[29]
The hearpiece database of individual transfer functions of an in-the-ear earpiece for hearing device research,
F. Denk and B. Kollmeier, “The hearpiece database of individual transfer functions of an in-the-ear earpiece for hearing device research,” Acta Acustica, vol. 5, no. 2, pp. 1–16, 2021
2021
-
[30]
CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2017. [Online]. Available: https://doi.org/10.7488/ds/2645
2017 doi
-
[31]
Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,
A. Varga and H. J. Steeneken, “Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,” Speech Communication, vol. 12, no. 3, pp. 247–251, 1993
1993
-
[32]
Methods for Calculation of the Speech Intelligibility Index,
Acoustical Society of America (ASA), “Methods for Calculation of the Speech Intelligibility Index,” American National Standards Institute (ANSI), ANSI/ASA S3.5-1997 Standard, 1997
1997
-
[33]
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (ICASS...
2001
-
[34]
An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,
J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 11, pp. 2009–2022, 2016
2009
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.