Pith. sign in

REVIEW 2 major objections 4 minor 30 references

Spatially Selective Active Noise Control for Open-fitting Hearables with Acausal Optimization

T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that letting the desired-source response be modeled by acausal relative impulse responses turns spatially selective active noise control for open-fitting hearables from a speech-distorting noise canceller into a true…

desk verdict A clean, honest extension of the authors' own causal SSANC work to acausal ReIRs, with large simulated gains that are real but narrower than the 'consistently outperforms' claim. read the letter →

arxiv 2505.10372 v1 pith:ZRWG3BEO submitted 2025-05-15 eess.AS cs.SDcs.SYeess.SPeess.SY

classification eess.AScs.SDcs.SYeess.SPeess.SY PACS 43.50.Ki
keywords activenoisecontrolopen-fittinghearablesspatialselectivityrelativeimpulseresponseacausalfilteringspeechdistortionreductionSNRimprovement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that spatially selective active noise control (SSANC) for open-fitting hearables is substantially improved when the relative impulse responses (ReIRs) of the desired speech source are allowed to be acausal instead of strictly causal. With a small amount of acausality ($L_a = 22$ taps) and a lightly regularized control filter, speech distortion at the inner error microphone drops from about $-2$ dB to about $-28$ dB while noise reduction stays near 21 dB and SNR improvement rises above 17 dB. The acausal design also makes performance nearly flat across a wide range of delays, removing the need to tune the delay to the acoustic propagation delay of the desired source. This matters because the prior causal approach, when tuned for strong noise reduction, tended to suppress the desired speech along with the noise.

What carries the argument

The central object is the acausal relative impulse response (ReIR) matrix $H$, built from per-channel filters $h_k(l)$ that carry an $L_a$-tap anti-causal part ($l = -L_a, \ldots, -1$) in addition to their causal taps. Each outer-microphone speech component is expressed as a convolution of the reference speech signal with its ReIR, and stacking these relations yields $x_s(n) = H^T x_{ref,s}(n)$. The speech-preservation constraint $e_s(n) = x_{ref,s}(n - \Delta)$ then becomes the linear equation $H(q + Gw) = \delta_\Delta$. The control filter $w$ is obtained by minimizing the inner error power plus a regularization term $\beta w^T w$ subject to that constraint, with the same closed-form structure as the earlier causal SSANC; setting $L_a = 0$ recovers the previous method. The acausal taps provide extra degrees of freedom to match the desired-source response, which is the mechanism that allows low speech distortion together with high noise reduction.

What would settle it

Run the same two-source simulation with a perturbed secondary-path estimate (for example, a 10% gain error) or with the speech source moved to 10 degrees off axis while the ReIRs are still the 0-degree white-noise estimates; the central claim fails if the acausal design no longer consistently beats the causal design on speech distortion and SNR improvement across the tested delays.

Watch

Extended reading notes

Core claim

The central discovery is that the bottleneck of spatially selective ANC is not the control effort but the causal modeling of the desired source. In the causal formulation, preserving the delayed speech component of a reference microphone at the inner error microphone forces a delay at least as large as the acoustic propagation delay, and large delays degrade noise reduction; with tight regularization the causal system gives up on speech preservation entirely and behaves like conventional ANC. The paper shows that replacing the causal ReIR matrix $H$ with an acausal one — having an $L_a$-tap anti-causal part — lets the constrained optimizer satisfy the speech-preservation constraint $H(q + Gw) = \delta_\Delta$ much more accurately while the control filter $w$ itself remains causal. As a result, for a two-source scenario and small $\beta$, the acausal design ($L_a = 22$) achieves speech distortion of about $-28$ dB, noise reduction around 21 dB, and SNR improvement above 17 dB across most delays, whereas the causal design at the same $\beta$ achieves only about $-2$ dB speech distortion. In a harder five-babble-source scenario, speech distortion improves from about $-15$ dB to $-26$ dB and SNR improvement from about 3 dB to 6 dB. The paper attributes the gain to acausal ReIRs characterizing the desired-source response more accurately, and shows that the benefit saturates once $L_a$ reaches roughly 12.

Load-bearing premise

The results rest on perfect knowledge of the secondary path and acoustic feedback paths, and on ReIRs estimated from white noise at the same 0-degree direction as the evaluated speech; if the real conditions deviate, the reported 26 dB gain may shrink or disappear.

Editorial extensions

If this is right

  • Speech distortion can be reduced by roughly 26 dB (from about $-2$ dB to $-28$ dB) without giving up noise reduction, which stays near 21 dB.
  • Performance becomes robust to the choice of delay: with $L_a = 22$ and low $\beta$, all three metrics stay essentially flat for delays from 4 to 80 samples, removing the need to tune $\Delta$ to the acoustic delay.
  • A modest degree of acausality suffices: speech distortion and SNR improvement stabilize once $L_a$ reaches about 12 taps in the two-source scenario.
  • In the more demanding five-babble-source scenario the same qualitative gains hold: speech distortion improves by about 11 dB and SNR improvement roughly doubles from 3 dB to 6 dB.
  • Because the control filter itself remains causal, the acausal ReIR formulation can be implemented in real time with a fixed reference delay, making the approach practical for hearables.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the ReIRs are estimated online or re-estimated when the talker moves, the acausal approach could track a moving desired source; the paper only tests the static case with matched estimation and evaluation directions.
  • The large gain at small $\beta$ suggests that in causal SSANC the regularization is fighting the wrong enemy: the constraint is infeasible because causal ReIRs cannot represent the desired-source response, not because the secondary source is too weak.
  • A natural testable extension is to replace the anechoic, matched-condition ReIRs with estimates from reverberant or slightly off-axis conditions to quantify how much of the 26 dB improvement survives model mismatch.
  • The same acausal-ReIR trick could be applied to the noise-reduction beamformers used in hearables, separating the spatial filter design from the causality constraint.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes an extension of spatially selective active noise control (SSANC) for open-fitting hearables by allowing the relative impulse responses (ReIRs) used in the desired-speech constraint to have an acausal part. Section 3 formulates the design as a regularized quadratic program with the equality constraint H(q+Gw)=δ_Δ (Eq. (17)), gives a closed-form solution (Eq. (18)), and notes that La=0 recovers the causal method. Section 4 evaluates the method in two anechoic KEMAR scenarios (one speech interferer at 45°; five babble sources) using intelligibility-weighted spectral distortion, noise reduction, and SNR improvement. The results show large gains for the acausal design in the first scenario with β=λ1/(2×10^6), and similar but smaller gains in the second scenario.

Significance. If the reported gains persist under realistic conditions, acausal ReIRs are a useful extension of SSANC: they reduce the minimum delay needed for causality, lower speech distortion at a given noise-reduction level, and make performance less sensitive to delay selection. The algebraic derivation is standard and appears correct, and the simulations use a measured KEMAR database with clearly specified filter lengths and source geometry. The main uncertainty is that the evidence is obtained under oracle conditions and with a small set of scenarios; the paper does not provide code or machine-checked proofs, but the optimization itself is standard and reproducible from the description.

major comments (2)
  1. [Section 4.3, Fig. 4] The comparison in the second scenario is confounded: the causal design (La=0) uses β=λ1/(4×10^3) whereas the acausal design (La=22) uses β=λ1/(4×10^7). Since β is the regularization weight in the objective (Eq. (17a)) and directly controls the trade-off between error minimization and control effort, a change of four orders of magnitude moves the two systems to very different operating points. The reported gains in speech distortion, noise reduction, and SNR improvement are therefore not attributable to acausality alone. Please either use the same β for both designs (as in the first scenario) or sweep β over a common range and show that the acausal design dominates at every β.
  2. [Sections 2 and 4.1] The simulations are run under oracle conditions: Section 2 assumes a perfect secondary-path estimate (bg=g) and known acoustic feedback paths, and Section 4.1 estimates the acausal ReIRs from white noise at exactly the same 0° source position and the same anechoic impulse responses used in the evaluation. Under these conditions the constraint in Eq. (16) can be satisfied almost exactly, which explains the very low speech distortion (about -28 dB) in Fig. 3(b). The paper does not test robustness to source-direction mismatch, secondary-path estimation error, or reverberation. I do not regard the ReIR estimation as circular, because it does not use the evaluation metrics, but the matched-condition advantage is significant. The abstract's claim that the method 'consistently outperforms the causal approach across all metrics and scenarios' is broader than the evidence. Please add a robustness test (e.g., a ±10° source shift or a few percent secondary-path error) or qualify the claim.
minor comments (4)
  1. [Eq. (22)] The symbol '·∆SNR' contains a stray centered dot; it should read 'ΔSNR'.
  2. [Fig. 6 caption] Please state the La value used for the plotted ReIRs and clarify whether these are the estimated ReIRs employed in the experiments.
  3. [Section 4.3] All conclusions are drawn from a single 5-second speech utterance and one noise realization per scenario; at least one additional realization or a noise bootstrap would support the use of the word 'consistently'.
  4. [Section 1 and 4.3] The claimed 'optimal range of delays' is only implicit in the figures; a concise statement of the range (e.g., Δ≈4–30 samples for scenario 1) would be helpful.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the acausal SSANC derivation is a standard constrained optimization, and the reported metrics are computed on separate signals rather than being forced by fitted parameters.

full rationale

The paper's derivation is self-contained. The acausal ReIR matrix H in Eq. (13) is obtained by adaptive system identification from white noise at the 0° source position (Section 4.1), and the equality constraint H(q+Gw)=δΔ in Eq. (16) is the mathematical reformulation of the desired signal-preservation condition e_s(n)=x_ref,s(n−Δ), not an expression of the evaluation metrics. The speech-distortion, noise-reduction, and SNR-improvement metrics (Eqs. 20-22) are evaluated on separate speech and noise signals, not on the identification signal used to estimate H. The lower speech distortion of the acausal design follows from the increased ability of an acausal H to represent the measured relative impulse responses, which is the intended design objective rather than a circular prediction. The perfect-secondary-path assumption and the matched source position are idealizations that limit generality, but they do not make the derivation circular. Self-citations [15,17] are used only to define the causal baseline and to note similarity of the solution form; they are not load-bearing for the acausal claim. No step in the paper reduces by construction to its own input, so the circularity score is 0.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central comparison rests on ideal-path assumptions such as a perfect secondary path estimate and known acoustic feedback cancellation, a single-source time-invariant ReIR model fitted to the exact desired-source position, and hand-selected regularization parameters. No new physical entities are introduced; acausal ReIRs are a modeling construct. The regularization parameters beta and rho are the main hand-chosen numbers, and the anti-causal tap count La is selected from the data.

free parameters (4)
  • Regularization factor beta = lambda1/(5e3), lambda1/(2e6), lambda1/(4e3), lambda1/(4e7) in different runs
    Controls the trade-off between noise reduction and speech distortion; selected by hand per scenario. The paper emphasizes the acausal design allows a smaller beta without distortion.
  • Regularization factor rho = lambda_max(HG Phi_rr^{-1} G^T H^T)/(1e5) or /(2e5)
    Small diagonal loading for the constraint matrix inverse; chosen by hand.
  • Acausal tap count La = 22 in main comparison; swept 0..40 in Fig. 5; performance stabilizes around 12
    Number of anti-causal taps in the ReIR model; the paper identifies about 12 as sufficient, making it a design parameter selected from the data.
  • Delay Delta = swept 0..80 samples; optimal around 4..30 for causal, wider for acausal
    Processing delay in samples; a design variable evaluated in the study.
assumptions (4)
  • domain assumption Perfect secondary path estimate (bg = g) and known acoustic feedback paths so that feedback is canceled
    Stated in Section 2 before Eq. (1) and after Eq. (6); central to the derivation of e(n)=q^T x(n).
  • domain assumption Speech and noise components are uncorrelated
    Used in Eq. (9) to separate es and ev.
  • domain assumption Desired speech in every microphone can be modeled as the convolution of the reference speech with a time-invariant relative impulse response of length La+Lh, possibly acausal
    Eq. (11); assumes a single desired source and known or stationary acoustic paths.
  • domain assumption The acoustic environment is anechoic and the impulse responses are exactly measured
    Simulation setup uses measured anechoic impulse responses from [19]; the central comparison is conducted only in these conditions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spatially Selective Active Noise Control for Open-fitting Hearables with Acausal Optimization." pith.science (2026). https://pith.science/paper/ZRWG3BEO

@misc{pith2026250510372,
  author       = {Pith},
  title        = {Pith review of: Spatially Selective Active Noise Control for Open-fitting Hearables with Acausal Optimization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZRWG3BEO}},
  note         = {Machine review of arXiv:2505.10372}
}
read the original abstract

Recent advances in active noise control have enabled the development of hearables with spatial selectivity, which actively suppress undesired noise while preserving desired sound from specific directions. In this work, we propose an improved approach to spatially selective active noise control that incorporates acausal relative impulse responses into the optimization process, resulting in significantly improved performance over the causal design. We evaluate the system through simulations using a pair of open-fitting hearables with spatially localized speech and noise sources in an anechoic environment. Performance is evaluated in terms of speech distortion, noise reduction, and signal-to-noise ratio improvement across different delays and degrees of acausality. Results show that the proposed acausal optimization consistently outperforms the causal approach across all metrics and scenarios, as acausal filters more effectively characterize the response of the desired source.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 28 canonical work pages

  1. [1]

    INTRODUCTION Active noise control (ANC) hearables create a quiet envi- ronment by using secondary sources to generate anti-noise, which minimizes sound at specific positions when superim- posed on the primary noise (leakage) [1, 2]. Based on their fit, hearables can be categorized as closed-fitting (com- pletely occluding the ear), open-fitting (partially...

  2. [2]

    1, we consider a hearable withK outer microphones

    SIGNAL MODEL As shown in Fig. 1, we consider a hearable withK outer microphones. Without loss of generality, we consider one loudspeaker as the secondary source and one inner error microphone, resulting in a total of K + 1 microphones. We assume that the acoustic feedback paths between the loudspeaker and the outer microphones are known, such that acousti...

  3. [3]

    This can be achieved by imposing the constraint es(n) =xref,s(n− ∆)

    ACAUSAL OPTIMIZA TION In the SSANC system, the objective is to minimize the power of the inner error microphone signal while preserv- ing the delayed desired speech component of an outer refer- ence microphone signal. This can be achieved by imposing the constraint es(n) =xref,s(n− ∆). (10) For each channel, the speech component can be rep- resented as th...

  4. [4]

    p361 005

    SIMULA TIONS 4.1 Setup For the simulation, we considered a pair of open-fitting hearables [18,19] inserted into both ears of a GRAS 45BB- 12 KEMAR Head & Torso simulator, as shown in Fig. 2(a). We used four outer microphones (entrance microphones and concha microphones at the left and right ears, labeled as #1– #4), one inner error microphone (located at ...

  5. [5]

    CONCLUSION This paper has presented an improved approach to spatially selective active noise control that incorporates acausal rela- tive impulse responses into the optimization process, lead- ing to a substantial performance improvement over the causal optimization approach. Using a pair of open-fitting hearables in two acoustic scenarios, we demonstrate...

  6. [6]

    ACKNOWLEDGMENTS This research was funded by the Deutsche Forschungsge- meinschaft (DFG, German Research Foundation) – Project- ID 352015383 – SFB 1330 C1

  7. [7]

    S. J. Elliott, Signal processing for active control. Aca- demic Press, 2000

  8. [8]

    Hansen, S

    C. Hansen, S. Snyder, X. Qiu, L. Brooks, and D. Moreau, Active Control of Noise and Vibration . CRC Press, 2 ed., Nov. 2012

Show all 30 references
  1. [9]

    Recent ad- vances on active noise control: open issues and innova- tive applications,

    Y . Kajikawa, W.-S. Gan, and S. M. Kuo, “Recent ad- vances on active noise control: open issues and innova- tive applications,” APSIPA Transactions on Signal and Information Processing, vol. 1, e3, pp. 1–21, 2012

  2. [10]

    Listening in a noisy environ- ment: Integration of active noise control in audio prod- ucts,

    C.-Y . Chang, A. Siswanto, C.-Y . Ho, T.-K. Yeh, Y .-R. Chen, and S. M. Kuo, “Listening in a noisy environ- ment: Integration of active noise control in audio prod- ucts,” IEEE Consumer Electronics Magazine , vol. 5, no. 4, pp. 34–43, 2016

  3. [11]

    Augmented/mixed reality audio for hear- ables: sensing, control, and rendering,

    R. Gupta, J. He, R. Ranjan, W.-S. Gan, F. Klein, C. Schneiderwind, A. Neidhardt, K. Brandenburg, and V . V¨alim¨aki, “Augmented/mixed reality audio for hear- ables: sensing, control, and rendering,” IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 63–89, 2022

  4. [12]

    Beamforming: a versa- tile approach to spatial filtering,

    B. Van Veen and K. Buckley, “Beamforming: a versa- tile approach to spatial filtering,”IEEE ASSP Magazine, vol. 5, no. 2, pp. 4–24, 1988

  5. [13]

    Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diver- sity using multiple microphones,

    S. Doclo, W. Kellermann, S. Makino, and S. E. Nord- holm, “Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diver- sity using multiple microphones,” IEEE Signal Pro- cessing Magazine, vol. 32, pp. 18–30, Mar. 2015

  6. [14]

    A consolidated perspective on multimi- crophone speech enhancement and source separation,

    S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A consolidated perspective on multimi- crophone speech enhancement and source separation,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 25, no. 4, pp. 692–730, 2017

  7. [15]

    Integrated active noise control and noise reduction in hearing aids,

    R. Serizel, M. Moonen, J. Wouters, and S. H. Jensen, “Integrated active noise control and noise reduction in hearing aids,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 6, pp. 1137– 1146, 2010

  8. [16]

    Output SNR analysis of integrated active noise con- trol and noise reduction in hearing aids under a single speech source scenario,

    R. Serizel, M. Moonen, J. Wouters, and S. H. Jensen, “Output SNR analysis of integrated active noise con- trol and noise reduction in hearing aids under a single speech source scenario,” Signal Processing, vol. 91, no. 8, pp. 1719–1729, 2011

  9. [17]

    Combined feedforward- feedback noise reduction schemes for open-fitting hear- ing aids,

    D. Dalga and S. Doclo, “Combined feedforward- feedback noise reduction schemes for open-fitting hear- ing aids,” in Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), (New Paltz, NY , USA), pp. 185–188, 2011

  10. [18]

    Theoretical performance anal- ysis of ANC-motivated noise reduction algorithms for open-fitting hearing aids,

    D. Dalga and S. Doclo, “Theoretical performance anal- ysis of ANC-motivated noise reduction algorithms for open-fitting hearing aids,” in Proc. International Workshop on Acoustic Signal Enhancement (IWAENC), (Aachen, Germany), pp. 1–4, 2012

  11. [19]

    Integrated active noise control for open-fit hearing aids with customized filter,

    C.-Y . Ho, K.-K. Shyu, C.-Y . Chang, and S. M. Kuo, “Integrated active noise control for open-fit hearing aids with customized filter,” Applied Acoustics, vol. 137, pp. 1–8, 2018

  12. [20]

    Design and imple- mentation of an active noise control headphone with directional hear-through capability,

    V . Patel, J. Cheer, and S. Fontana, “Design and imple- mentation of an active noise control headphone with directional hear-through capability,” IEEE Transac- tions on Consumer Electronics , vol. 66, pp. 32–40, Feb. 2020

  13. [21]

    Spatially selective active noise control systems,

    T. Xiao, B. Xu, and C. Zhao, “Spatially selective active noise control systems,” The Journal of the Acoustical Society of America , vol. 153, pp. 2733–2744, May 2023

  14. [22]

    A time-domain multi-channel directional ac- tive noise control system,

    H. Zhang, J. Zhang, F. Ma, P. N. Samarasinghe, and H. Sun, “A time-domain multi-channel directional ac- tive noise control system,” in Proc. European Signal Processing Conference (EUSIPCO) , (Helsinki, Fin- land), pp. 376–380, 2023

  15. [23]

    Effect of target signals and delays on spatially selective active noise control for open-fitting hearables,

    T. Xiao and S. Doclo, “Effect of target signals and delays on spatially selective active noise control for open-fitting hearables,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Process- ing (ICASSP), (Seoul, Republic of Korea), pp. 1056– 1060, 2024

  16. [24]

    A one-size-fits-all earpiece with multiple microphones and drivers for hearing device research,

    F. Denk, M. Lettau, H. Schepker, S. Doclo, R. Roden, M. Blau, J.-H. Bach, J. Wellmann, and B. Kollmeier, “A one-size-fits-all earpiece with multiple microphones and drivers for hearing device research,” in Proc. AES International Conference on Headphone Technology, (San Franci...

  17. [25]

    The hearpiece database of individual transfer functions of an in-the-ear earpiece for hearing device research,

    F. Denk and B. Kollmeier, “The hearpiece database of individual transfer functions of an in-the-ear earpiece for hearing device research,” Acta Acustica , vol. 5, no. 2, pp. 1–16, 2021

  18. [26]

    CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,

    C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2017

  19. [27]

    Assessment for auto- matic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,

    A. Varga and H. J. Steeneken, “Assessment for auto- matic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,” Speech Communica- tion, vol. 12, no. 3, pp. 247–251, 1993

  20. [28]

    Spatially pre- processed speech distortion weighted multi-channel Wiener filtering for noise reduction,

    A. Spriet, M. Moonen, and J. Wouters, “Spatially pre- processed speech distortion weighted multi-channel Wiener filtering for noise reduction,” Signal Process- ing, vol. 84, no. 12, pp. 2367–2387, 2004

  21. [29]

    Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction,

    S. Doclo, A. Spriet, J. Wouters, and M. Moonen, “Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction,” Speech Communication , vol. 49, no. 7, pp. 636–656, 2007

  22. [30]

    Methods for Calculation of the Speech Intelligibility Index,

    Acoustical Society of America (ASA), “Methods for Calculation of the Speech Intelligibility Index,” ANSI/ASA S3.5-1997 Standard, American National Standards Institute (ANSI), 1997

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.