REVIEW 2 major objections 4 minor 30 references
Spatially Selective Active Noise Control for Open-fitting Hearables with Acausal Optimization
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that letting the desired-source response be modeled by acausal relative impulse responses turns spatially selective active noise control for open-fitting hearables from a speech-distorting noise canceller into a true…
desk verdict A clean, honest extension of the authors' own causal SSANC work to acausal ReIRs, with large simulated gains that are real but narrower than the 'consistently outperforms' claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the acausal relative impulse response (ReIR) matrix $H$, built from per-channel filters $h_k(l)$ that carry an $L_a$-tap anti-causal part ($l = -L_a, \ldots, -1$) in addition to their causal taps. Each outer-microphone speech component is expressed as a convolution of the reference speech signal with its ReIR, and stacking these relations yields $x_s(n) = H^T x_{ref,s}(n)$. The speech-preservation constraint $e_s(n) = x_{ref,s}(n - \Delta)$ then becomes the linear equation $H(q + Gw) = \delta_\Delta$. The control filter $w$ is obtained by minimizing the inner error power plus a regularization term $\beta w^T w$ subject to that constraint, with the same closed-form structure as the earlier causal SSANC; setting $L_a = 0$ recovers the previous method. The acausal taps provide extra degrees of freedom to match the desired-source response, which is the mechanism that allows low speech distortion together with high noise reduction.
What would settle it
Run the same two-source simulation with a perturbed secondary-path estimate (for example, a 10% gain error) or with the speech source moved to 10 degrees off axis while the ReIRs are still the 0-degree white-noise estimates; the central claim fails if the acausal design no longer consistently beats the causal design on speech distortion and SNR improvement across the tested delays.
Extended reading notes
Core claim
The central discovery is that the bottleneck of spatially selective ANC is not the control effort but the causal modeling of the desired source. In the causal formulation, preserving the delayed speech component of a reference microphone at the inner error microphone forces a delay at least as large as the acoustic propagation delay, and large delays degrade noise reduction; with tight regularization the causal system gives up on speech preservation entirely and behaves like conventional ANC. The paper shows that replacing the causal ReIR matrix $H$ with an acausal one — having an $L_a$-tap anti-causal part — lets the constrained optimizer satisfy the speech-preservation constraint $H(q + Gw) = \delta_\Delta$ much more accurately while the control filter $w$ itself remains causal. As a result, for a two-source scenario and small $\beta$, the acausal design ($L_a = 22$) achieves speech distortion of about $-28$ dB, noise reduction around 21 dB, and SNR improvement above 17 dB across most delays, whereas the causal design at the same $\beta$ achieves only about $-2$ dB speech distortion. In a harder five-babble-source scenario, speech distortion improves from about $-15$ dB to $-26$ dB and SNR improvement from about 3 dB to 6 dB. The paper attributes the gain to acausal ReIRs characterizing the desired-source response more accurately, and shows that the benefit saturates once $L_a$ reaches roughly 12.
Load-bearing premise
The results rest on perfect knowledge of the secondary path and acoustic feedback paths, and on ReIRs estimated from white noise at the same 0-degree direction as the evaluated speech; if the real conditions deviate, the reported 26 dB gain may shrink or disappear.
Editorial extensions
If this is right
- Speech distortion can be reduced by roughly 26 dB (from about $-2$ dB to $-28$ dB) without giving up noise reduction, which stays near 21 dB.
- Performance becomes robust to the choice of delay: with $L_a = 22$ and low $\beta$, all three metrics stay essentially flat for delays from 4 to 80 samples, removing the need to tune $\Delta$ to the acoustic delay.
- A modest degree of acausality suffices: speech distortion and SNR improvement stabilize once $L_a$ reaches about 12 taps in the two-source scenario.
- In the more demanding five-babble-source scenario the same qualitative gains hold: speech distortion improves by about 11 dB and SNR improvement roughly doubles from 3 dB to 6 dB.
- Because the control filter itself remains causal, the acausal ReIR formulation can be implemented in real time with a fixed reference delay, making the approach practical for hearables.
Reading between the lines
- If the ReIRs are estimated online or re-estimated when the talker moves, the acausal approach could track a moving desired source; the paper only tests the static case with matched estimation and evaluation directions.
- The large gain at small $\beta$ suggests that in causal SSANC the regularization is fighting the wrong enemy: the constraint is infeasible because causal ReIRs cannot represent the desired-source response, not because the secondary source is too weak.
- A natural testable extension is to replace the anechoic, matched-condition ReIRs with estimates from reverberant or slightly off-axis conditions to quantify how much of the 26 dB improvement survives model mismatch.
- The same acausal-ReIR trick could be applied to the noise-reduction beamformers used in hearables, separating the spatial filter design from the causality constraint.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an extension of spatially selective active noise control (SSANC) for open-fitting hearables by allowing the relative impulse responses (ReIRs) used in the desired-speech constraint to have an acausal part. Section 3 formulates the design as a regularized quadratic program with the equality constraint H(q+Gw)=δ_Δ (Eq. (17)), gives a closed-form solution (Eq. (18)), and notes that La=0 recovers the causal method. Section 4 evaluates the method in two anechoic KEMAR scenarios (one speech interferer at 45°; five babble sources) using intelligibility-weighted spectral distortion, noise reduction, and SNR improvement. The results show large gains for the acausal design in the first scenario with β=λ1/(2×10^6), and similar but smaller gains in the second scenario.
Significance. If the reported gains persist under realistic conditions, acausal ReIRs are a useful extension of SSANC: they reduce the minimum delay needed for causality, lower speech distortion at a given noise-reduction level, and make performance less sensitive to delay selection. The algebraic derivation is standard and appears correct, and the simulations use a measured KEMAR database with clearly specified filter lengths and source geometry. The main uncertainty is that the evidence is obtained under oracle conditions and with a small set of scenarios; the paper does not provide code or machine-checked proofs, but the optimization itself is standard and reproducible from the description.
major comments (2)
- [Section 4.3, Fig. 4] The comparison in the second scenario is confounded: the causal design (La=0) uses β=λ1/(4×10^3) whereas the acausal design (La=22) uses β=λ1/(4×10^7). Since β is the regularization weight in the objective (Eq. (17a)) and directly controls the trade-off between error minimization and control effort, a change of four orders of magnitude moves the two systems to very different operating points. The reported gains in speech distortion, noise reduction, and SNR improvement are therefore not attributable to acausality alone. Please either use the same β for both designs (as in the first scenario) or sweep β over a common range and show that the acausal design dominates at every β.
- [Sections 2 and 4.1] The simulations are run under oracle conditions: Section 2 assumes a perfect secondary-path estimate (bg=g) and known acoustic feedback paths, and Section 4.1 estimates the acausal ReIRs from white noise at exactly the same 0° source position and the same anechoic impulse responses used in the evaluation. Under these conditions the constraint in Eq. (16) can be satisfied almost exactly, which explains the very low speech distortion (about -28 dB) in Fig. 3(b). The paper does not test robustness to source-direction mismatch, secondary-path estimation error, or reverberation. I do not regard the ReIR estimation as circular, because it does not use the evaluation metrics, but the matched-condition advantage is significant. The abstract's claim that the method 'consistently outperforms the causal approach across all metrics and scenarios' is broader than the evidence. Please add a robustness test (e.g., a ±10° source shift or a few percent secondary-path error) or qualify the claim.
minor comments (4)
- [Eq. (22)] The symbol '·∆SNR' contains a stray centered dot; it should read 'ΔSNR'.
- [Fig. 6 caption] Please state the La value used for the plotted ReIRs and clarify whether these are the estimated ReIRs employed in the experiments.
- [Section 4.3] All conclusions are drawn from a single 5-second speech utterance and one noise realization per scenario; at least one additional realization or a noise bootstrap would support the use of the word 'consistently'.
- [Section 1 and 4.3] The claimed 'optimal range of delays' is only implicit in the figures; a concise statement of the range (e.g., Δ≈4–30 samples for scenario 1) would be helpful.
Circularity Check
No significant circularity: the acausal SSANC derivation is a standard constrained optimization, and the reported metrics are computed on separate signals rather than being forced by fitted parameters.
full rationale
The paper's derivation is self-contained. The acausal ReIR matrix H in Eq. (13) is obtained by adaptive system identification from white noise at the 0° source position (Section 4.1), and the equality constraint H(q+Gw)=δΔ in Eq. (16) is the mathematical reformulation of the desired signal-preservation condition e_s(n)=x_ref,s(n−Δ), not an expression of the evaluation metrics. The speech-distortion, noise-reduction, and SNR-improvement metrics (Eqs. 20-22) are evaluated on separate speech and noise signals, not on the identification signal used to estimate H. The lower speech distortion of the acausal design follows from the increased ability of an acausal H to represent the measured relative impulse responses, which is the intended design objective rather than a circular prediction. The perfect-secondary-path assumption and the matched source position are idealizations that limit generality, but they do not make the derivation circular. Self-citations [15,17] are used only to define the causal baseline and to note similarity of the solution form; they are not load-bearing for the acausal claim. No step in the paper reduces by construction to its own input, so the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Regularization factor beta =
lambda1/(5e3), lambda1/(2e6), lambda1/(4e3), lambda1/(4e7) in different runs
- Regularization factor rho =
lambda_max(HG Phi_rr^{-1} G^T H^T)/(1e5) or /(2e5)
- Acausal tap count La =
22 in main comparison; swept 0..40 in Fig. 5; performance stabilizes around 12
- Delay Delta =
swept 0..80 samples; optimal around 4..30 for causal, wider for acausal
assumptions (4)
- domain assumption Perfect secondary path estimate (bg = g) and known acoustic feedback paths so that feedback is canceled
- domain assumption Speech and noise components are uncorrelated
- domain assumption Desired speech in every microphone can be modeled as the convolution of the reference speech with a time-invariant relative impulse response of length La+Lh, possibly acausal
- domain assumption The acoustic environment is anechoic and the impulse responses are exactly measured
Cite this review
Pith. "Pith review of Spatially Selective Active Noise Control for Open-fitting Hearables with Acausal Optimization." pith.science (2026). https://pith.science/paper/ZRWG3BEO
@misc{pith2026250510372,
author = {Pith},
title = {Pith review of: Spatially Selective Active Noise Control for Open-fitting Hearables with Acausal Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZRWG3BEO}},
note = {Machine review of arXiv:2505.10372}
}
read the original abstract
Recent advances in active noise control have enabled the development of hearables with spatial selectivity, which actively suppress undesired noise while preserving desired sound from specific directions. In this work, we propose an improved approach to spatially selective active noise control that incorporates acausal relative impulse responses into the optimization process, resulting in significantly improved performance over the causal design. We evaluate the system through simulations using a pair of open-fitting hearables with spatially localized speech and noise sources in an anechoic environment. Performance is evaluated in terms of speech distortion, noise reduction, and signal-to-noise ratio improvement across different delays and degrees of acausality. Results show that the proposed acausal optimization consistently outperforms the causal approach across all metrics and scenarios, as acausal filters more effectively characterize the response of the desired source.
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION Active noise control (ANC) hearables create a quiet envi- ronment by using secondary sources to generate anti-noise, which minimizes sound at specific positions when superim- posed on the primary noise (leakage) [1, 2]. Based on their fit, hearables can be categorized as closed-fitting (com- pletely occluding the ear), open-fitting (partially...
work page Pith review arXiv 2025
-
[2]
1, we consider a hearable withK outer microphones
SIGNAL MODEL As shown in Fig. 1, we consider a hearable withK outer microphones. Without loss of generality, we consider one loudspeaker as the secondary source and one inner error microphone, resulting in a total of K + 1 microphones. We assume that the acoustic feedback paths between the loudspeaker and the outer microphones are known, such that acousti...
work page 2025
-
[3]
This can be achieved by imposing the constraint es(n) =xref,s(n− ∆)
ACAUSAL OPTIMIZA TION In the SSANC system, the objective is to minimize the power of the inner error microphone signal while preserv- ing the delayed desired speech component of an outer refer- ence microphone signal. This can be achieved by imposing the constraint es(n) =xref,s(n− ∆). (10) For each channel, the speech component can be rep- resented as th...
-
[4]
SIMULA TIONS 4.1 Setup For the simulation, we considered a pair of open-fitting hearables [18,19] inserted into both ears of a GRAS 45BB- 12 KEMAR Head & Torso simulator, as shown in Fig. 2(a). We used four outer microphones (entrance microphones and concha microphones at the left and right ears, labeled as #1– #4), one inner error microphone (located at ...
work page 2025
-
[5]
CONCLUSION This paper has presented an improved approach to spatially selective active noise control that incorporates acausal rela- tive impulse responses into the optimization process, lead- ing to a substantial performance improvement over the causal optimization approach. Using a pair of open-fitting hearables in two acoustic scenarios, we demonstrate...
work page 2025
-
[6]
ACKNOWLEDGMENTS This research was funded by the Deutsche Forschungsge- meinschaft (DFG, German Research Foundation) – Project- ID 352015383 – SFB 1330 C1
-
[7]
S. J. Elliott, Signal processing for active control. Aca- demic Press, 2000
work page 2000
- [8]
Show all 30 references
-
[9]
Recent ad- vances on active noise control: open issues and innova- tive applications,
Y . Kajikawa, W.-S. Gan, and S. M. Kuo, “Recent ad- vances on active noise control: open issues and innova- tive applications,” APSIPA Transactions on Signal and Information Processing, vol. 1, e3, pp. 1–21, 2012
2012
-
[10]
Listening in a noisy environ- ment: Integration of active noise control in audio prod- ucts,
C.-Y . Chang, A. Siswanto, C.-Y . Ho, T.-K. Yeh, Y .-R. Chen, and S. M. Kuo, “Listening in a noisy environ- ment: Integration of active noise control in audio prod- ucts,” IEEE Consumer Electronics Magazine , vol. 5, no. 4, pp. 34–43, 2016
2016
-
[11]
Augmented/mixed reality audio for hear- ables: sensing, control, and rendering,
R. Gupta, J. He, R. Ranjan, W.-S. Gan, F. Klein, C. Schneiderwind, A. Neidhardt, K. Brandenburg, and V . V¨alim¨aki, “Augmented/mixed reality audio for hear- ables: sensing, control, and rendering,” IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 63–89, 2022
2022
-
[12]
Beamforming: a versa- tile approach to spatial filtering,
B. Van Veen and K. Buckley, “Beamforming: a versa- tile approach to spatial filtering,”IEEE ASSP Magazine, vol. 5, no. 2, pp. 4–24, 1988
1988
-
[13]
Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diver- sity using multiple microphones,
S. Doclo, W. Kellermann, S. Makino, and S. E. Nord- holm, “Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diver- sity using multiple microphones,” IEEE Signal Pro- cessing Magazine, vol. 32, pp. 18–30, Mar. 2015
2015
-
[14]
A consolidated perspective on multimi- crophone speech enhancement and source separation,
S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A consolidated perspective on multimi- crophone speech enhancement and source separation,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 25, no. 4, pp. 692–730, 2017
2017
-
[15]
Integrated active noise control and noise reduction in hearing aids,
R. Serizel, M. Moonen, J. Wouters, and S. H. Jensen, “Integrated active noise control and noise reduction in hearing aids,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 6, pp. 1137– 1146, 2010
2010
-
[16]
Output SNR analysis of integrated active noise con- trol and noise reduction in hearing aids under a single speech source scenario,
R. Serizel, M. Moonen, J. Wouters, and S. H. Jensen, “Output SNR analysis of integrated active noise con- trol and noise reduction in hearing aids under a single speech source scenario,” Signal Processing, vol. 91, no. 8, pp. 1719–1729, 2011
2011
-
[17]
Combined feedforward- feedback noise reduction schemes for open-fitting hear- ing aids,
D. Dalga and S. Doclo, “Combined feedforward- feedback noise reduction schemes for open-fitting hear- ing aids,” in Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), (New Paltz, NY , USA), pp. 185–188, 2011
2011
-
[18]
Theoretical performance anal- ysis of ANC-motivated noise reduction algorithms for open-fitting hearing aids,
D. Dalga and S. Doclo, “Theoretical performance anal- ysis of ANC-motivated noise reduction algorithms for open-fitting hearing aids,” in Proc. International Workshop on Acoustic Signal Enhancement (IWAENC), (Aachen, Germany), pp. 1–4, 2012
2012
-
[19]
Integrated active noise control for open-fit hearing aids with customized filter,
C.-Y . Ho, K.-K. Shyu, C.-Y . Chang, and S. M. Kuo, “Integrated active noise control for open-fit hearing aids with customized filter,” Applied Acoustics, vol. 137, pp. 1–8, 2018
2018
-
[20]
Design and imple- mentation of an active noise control headphone with directional hear-through capability,
V . Patel, J. Cheer, and S. Fontana, “Design and imple- mentation of an active noise control headphone with directional hear-through capability,” IEEE Transac- tions on Consumer Electronics , vol. 66, pp. 32–40, Feb. 2020
2020
-
[21]
Spatially selective active noise control systems,
T. Xiao, B. Xu, and C. Zhao, “Spatially selective active noise control systems,” The Journal of the Acoustical Society of America , vol. 153, pp. 2733–2744, May 2023
2023
-
[22]
A time-domain multi-channel directional ac- tive noise control system,
H. Zhang, J. Zhang, F. Ma, P. N. Samarasinghe, and H. Sun, “A time-domain multi-channel directional ac- tive noise control system,” in Proc. European Signal Processing Conference (EUSIPCO) , (Helsinki, Fin- land), pp. 376–380, 2023
2023
-
[23]
Effect of target signals and delays on spatially selective active noise control for open-fitting hearables,
T. Xiao and S. Doclo, “Effect of target signals and delays on spatially selective active noise control for open-fitting hearables,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Process- ing (ICASSP), (Seoul, Republic of Korea), pp. 1056– 1060, 2024
2024
-
[24]
A one-size-fits-all earpiece with multiple microphones and drivers for hearing device research,
F. Denk, M. Lettau, H. Schepker, S. Doclo, R. Roden, M. Blau, J.-H. Bach, J. Wellmann, and B. Kollmeier, “A one-size-fits-all earpiece with multiple microphones and drivers for hearing device research,” in Proc. AES International Conference on Headphone Technology, (San Franci...
2019
-
[25]
The hearpiece database of individual transfer functions of an in-the-ear earpiece for hearing device research,
F. Denk and B. Kollmeier, “The hearpiece database of individual transfer functions of an in-the-ear earpiece for hearing device research,” Acta Acustica , vol. 5, no. 2, pp. 1–16, 2021
2021
-
[26]
CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2017
2017
-
[27]
Assessment for auto- matic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,
A. Varga and H. J. Steeneken, “Assessment for auto- matic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,” Speech Communica- tion, vol. 12, no. 3, pp. 247–251, 1993
1993
-
[28]
Spatially pre- processed speech distortion weighted multi-channel Wiener filtering for noise reduction,
A. Spriet, M. Moonen, and J. Wouters, “Spatially pre- processed speech distortion weighted multi-channel Wiener filtering for noise reduction,” Signal Process- ing, vol. 84, no. 12, pp. 2367–2387, 2004
2004
-
[29]
Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction,
S. Doclo, A. Spriet, J. Wouters, and M. Moonen, “Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction,” Speech Communication , vol. 49, no. 7, pp. 636–656, 2007
2007
-
[30]
Methods for Calculation of the Speech Intelligibility Index,
Acoustical Society of America (ASA), “Methods for Calculation of the Speech Intelligibility Index,” ANSI/ASA S3.5-1997 Standard, American National Standards Institute (ANSI), 1997
1997
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.