REVIEW 3 major objections 5 minor 25 references
RTF-steered binaural MVDR beamforming incorporating multiple external microphones
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Combining per-microphone RTF estimates via the output-SNR-maximizing eigenvector gives the best binaural MVDR beamformer and beats covariance whitening in a moving-speaker experiment.
desk verdict A clean, small extension of SC-based RTF estimation to multiple external microphones; the mSNR derivation is correct, but the experimental support is a single recording with no held-out evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the normalized linear combination $\mathbf{a}_L^{\mathrm{SC-C}} = \mathbf{A}_L^{\mathrm{SC}}\mathbf{c} / (\mathbf{e}_L^T \mathbf{A}_L^{\mathrm{SC}}\mathbf{c})$, where $\mathbf{A}_L^{\mathrm{SC}}$ stacks the per-external-microphone spatial-coherence RTF estimates and $\mathbf{c}$ is a per-frequency combination vector. The paper expresses the BMVDR output SNR as the generalized Rayleigh quotient $\mathrm{SNR}_{\mathrm{out}} = \mathbf{c}^H\Lambda_1\mathbf{c}/(\mathbf{c}^H\Lambda_2\mathbf{c}-1)$ and shows that maximizing it is achieved by the principal eigenvector of $\Lambda_2^{-1}\Lambda_1$. This identity turns the practical question of how to combine RTF estimates into a small eigenvalue problem, with dimension equal to the number of external microphones rather than the full microphone count.
What would settle it
Place the external microphones close together or use directional noise so their noise components become correlated, then rerun the moving-speaker experiment; if the mSNR combination no longer outperforms covariance whitening, the uncorrelated-noise assumption is the load-bearing premise.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that linearly combining several spatial-coherence RTF estimates so as to maximize the narrowband output SNR of the BMVDR yields a better beamformer than either selecting a single external-microphone estimate, averaging all estimates, or applying the covariance whitening method. Per frequency, the maximizing combination vector is the principal eigenvector of $\Lambda_2^{-1}\Lambda_1$, where $\Lambda_1$ and $\Lambda_2$ are built from the stacked RTF estimates, the noise covariance matrix, and the noisy input covariance matrix. In the reported dynamic scenario—moving speaker, roughly 400 ms reverberation, pseudo-diffuse noise, and three external microphones—this mSNR combination reaches an average binaural SNR improvement of 10.7 dB, against 10.4 dB for covariance whitening, 10.3 dB for input-SNR selection, and 8.9 dB for averaging. The claim includes a complexity benefit: the needed eigenvalue decomposition has dimension $M_E$ (here 3) rather than the full $M=7$ dimension of the covariance whitening method.
Load-bearing premise
The load-bearing premise is that noise in each external microphone is uncorrelated with noise in every other microphone, which holds only approximately for diffuse noise with well-separated microphones.
Editorial extensions
If this is right
- mSNR gives the largest binaural SNR improvement (10.7 dB) among all compared methods, beating covariance whitening by 0.3 dB.
- The iSNR combination nearly ties covariance whitening (10.3 dB) while requiring no eigenvalue decomposition at all.
- Averaging the estimates is the worst procedure (8.9 dB), consistent with non-uniform estimation errors across external microphones.
- The mSNR combination is computationally lighter than covariance whitening: a 3×3 eigenvalue decomposition instead of a 7×7 one.
- The mSNR advantage holds across most time instances of the 30-second moving-speaker recording, not only on average.
Reading between the lines
- The mSNR construction does not depend on how the RTF estimates were obtained; the same generalized-Rayleigh-quotient combination should maximize output SNR for any collection of RTF estimates, so it could be reused with other estimators.
- Since each frequency is combined independently, a smoothness constraint across frequency might prevent discontinuities and could improve broadband speech quality beyond the reported SNR metric.
- The comparison rests on a single acoustic scenario; repeating it with different talkers, reverberation times, and external-microphone placements would determine how general the 0.3 dB advantage is.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper considers a binaural MVDR beamformer for a hearing-device configuration with multiple external microphones, where the desired-speech RTF vector is estimated per external microphone using the spatial-coherence (SC) method. Three combination procedures are proposed: selection of the estimate with the highest input SNR (iSNR), simple averaging (AV), and a per-frequency linear combination that maximizes the narrowband output SNR of the BMVDR (mSNR). The mSNR combination is derived in closed form as the principal eigenvector of the matrix Λ2^{-1}Λ1 of the generalized Rayleigh quotient for the BMVDR output SNR. An experimental evaluation in a single dynamic scenario with a moving speaker and pseudo-diffuse noise shows that mSNR yields the largest average binaural SNR improvement (10.7 dB), slightly outperforming the covariance-whitening method (10.4 dB).
Significance. If the reported comparative advantage were validated, the mSNR combination would be a useful, low-complexity alternative to covariance whitening for binaural beamforming with multiple external microphones. The derivation in Eqs. (20)–(23) is clean and correct, and it introduces no new free parameters: the combination vector is obtained directly from the estimated covariance matrices. The paper also demonstrates a computational advantage (a 3×3 EVD instead of a 7×7 EVD) and provides online audio demonstrations. The main weakness is experimental: the evaluation rests on a single recording, the mSNR combination is optimized on the same data used to measure its performance, and no uncertainty quantification is provided. These limitations leave the central claim of superiority over the state of the art insufficiently supported.
major comments (3)
- [Section 6, Eq. (23), Fig. 3] The mSNR combination vector is defined per frequency as the principal eigenvector of Λ2^{-1}Λ1, which maximizes the plug-in output SNR in Eq. (20) computed from covariance estimates on the test recording. The reported ΔBSNR in Fig. 3 is therefore an in-sample quantity for mSNR, and the 0.3 dB margin over the CW method is not evidence of generalization to new recordings. Please add a held-out evaluation, for example by cross-validating the combination vector across time segments, using multiple speaker trajectories or noise realizations, or reporting performance on a separate test recording, along with error bars or confidence intervals.
- [Section 6.1 and 6.2] Only one acoustic scenario is evaluated: one moving speaker, one pseudo-diffuse noise field, and one reverberation time. Figure 4 shows time traces for a single run, and no error bars or statistical tests are provided to support the claim that mSNR outperforms iSNR and AV 'for almost all time instances' or that the average differences are significant. Additional independent trials or a statistical analysis are needed to support the comparative conclusions.
- [Section 5, Eq. (15)] The SC estimator in Eq. (15) assumes that the noise component in each external microphone is uncorrelated with the noise components in all other microphones. In the pseudo-diffuse field generated by four loudspeakers, this assumption holds only approximately, and the bias is acknowledged in [12] to be real-valued and input-SNR-dependent. The magnitude of this bias in the present setup is not quantified. A sensitivity analysis or an evaluation in a more truly diffuse field would help establish the robustness of the proposed combination procedures.
minor comments (5)
- [Figure 3] The y-axis spans only 6–11 dB, which visually amplifies differences of about 0.3 dB; please consider starting the axis at 0 dB or adding error bars to provide an honest visual scale.
- [Notation] The abbreviation 'A V' appears with a space in several places (e.g., Eqs. (19), Fig. 4); please use the consistent notation 'AV'.
- [Eq. (18)] The iSNR selection rule in Eq. (18) requires both R_y and R_n; the sentence 'this only requires an estimate of R_y (and not R_x)' is misleading because R_n is also needed, although it is already available for the BMVDR.
- [Section 5.1] The statement about the bias in the SC estimator is brief; a short sentence giving the bias expression or its dependence on input SNR would help the reader assess the approximation.
- [Reproducibility] The audio demo link is a good resource; making the processing scripts available would further improve reproducibility.
Circularity Check
mSNR combination is fit to maximize output SNR and then evaluated by the same output-SNR metric on the same recording, making the headline ranking partly in-sample.
-
fitted input called prediction
[Section 5.2, Eqs. (20) and (23); Section 6, Eq. (24) and Fig. 3.]
"As a more sophisticated procedure, denoted as mSNR, we propose to combine the SC-based RTF vector estimates (per frequency) such that the narrowband output SNR of the BMVDR is maximized. ... cmSNR = arg maxc SNRout BMVDR,L =P{Λ−1 2 Λ1} ... The SNR-maximizing combination (mSNR) yields an average ∆BSNR of 10.7 dB, hence outperforming all other combination procedures and RTF vector estimation methods."
The mSNR combination vector is estimated per frequency from the same recording's estimated covariance matrices Ry and Rn to maximize the narrowband output SNR of the BMVDR (Eq. 20/23). The reported performance metric, ∆BSNR in Eq. (24), is the same output-SNR quantity (averaged binaurally over time and frequency) computed on that same recording. Thus the ranking mSNR > iSNR/AV is not an independent experimental finding: c_mSNR is by construction the maximizer of the very objective being reported, so the in-sample comparison against the other proposed combinations is statistically forced.
full rationale
The paper's genuine contribution is the combination framework for multiple SC-based RTF estimates (iSNR, AV, mSNR), and the SC estimator itself is taken from the authors' prior work [11,12] with an explicitly stated noise-uncorrelatedness assumption. That self-citation is a normal building block rather than a circular load-bearing argument, because the current paper does not derive its result from the truth of the SC estimator alone; the combination strategies are the novel part. However, the experimental validation of mSNR is partially circular: mSNR is defined as the combination maximizing narrowband output SNR (Eq. 23), and the evaluation metric ∆BSNR (Eq. 24) is an output-SNR measure computed on the same single dynamic recording used to fit c_mSNR. Therefore the claimed superiority of mSNR over iSNR and AV is largely a restatement of the optimization objective, not an independent empirical confirmation. The comparison with the covariance-whitening method provides some external grounding, but it is still in-sample and lacks held-out evaluation or error bars. This warrants a moderate circularity score rather than a higher one: the paper is not a purely self-referential derivation, but a key reported 'prediction' reduces to the fitting criterion on the test recording.
Assumptions & free parameters
free parameters (2)
- Recursive estimation time constants =
250 ms (Ry), 1.5 s (Rn)
- Speech presence probability threshold =
Not specified in the paper
assumptions (4)
- domain assumption Single desired speech source
- domain assumption Statistical independence between speech and noise
- domain assumption Diffuse noise field with uncorrelated noise between external and other microphones
- domain assumption Separately recorded speech and noise enable oracle covariance computation for evaluation
Cite this review
Pith. "Pith review of RTF-steered binaural MVDR beamforming incorporating multiple external microphones." pith.science (2026). https://pith.science/paper/UZ3GNK3B
@misc{pith2026190804848,
author = {Pith},
title = {Pith review of: RTF-steered binaural MVDR beamforming incorporating multiple external microphones},
year = {2026},
howpublished = {\url{https://pith.science/paper/UZ3GNK3B}},
note = {Machine review of arXiv:1908.04848}
}
read the original abstract
The binaural minimum-variance distortionless-response (BMVDR) beamformer is a well-known noise reduction algorithm that can be steered using the relative transfer function (RTF) vector of the desired speech source. Exploiting the availability of an external microphone that is spatially separated from the head-mounted microphones, an efficient method has been recently proposed to estimate the RTF vector in a diffuse noise field. When multiple external microphones are available, different RTF vector estimates can be obtained by using this method for each external microphone. In this paper, we propose several procedures to combine these RTF vector estimates, either by selecting the estimate corresponding to the highest input SNR, by averaging the estimates or by combining the estimates in order to maximize the output SNR of the BMVDR beamformer. Experimental results for a moving speaker and diffuse noise in a reverberant environment show that the output SNR-maximizing combination yields the largest binaural SNR improvement and also outperforms the state-of-the art covariance whitening method.
Figures
Reference graph
Works this paper leans on
-
[12]
Robust distributed noise reduction in hearing aids with external acoustic sensor nodes,
A. Bertrand and M. Moonen, “Robust distributed noise reduction in hearing aids with external acoustic sensor nodes,” EURASIP Journal on Advances in Signal Processing , vol. 2009, p. 14 pages, Jan. 2009
work page 2009
-
[1]
INTRODUCTION Noise reduction algorithms for head-mounted assistive listening devices (e.g., hearing aids, earbuds, headsets) are crucial to improve speech intelligibility and speech quality in noisy environments. Binaural noise reduction algorithms, which exploit the information captured by all microphones on both sides of the head [1, 2], do not only all...
-
[2]
RTF-steered binaural MVDR beamforming incorporating multiple external microphones
CONFIGURA TION AND NOTA TION Consider the binaural hearing device configuration depicted in Figure 1, consisting of a left and a right hearing device (each equipped with MD microphones), and ME external microphones that are spatially separated from the head-mounted microphones, i.e. M = 2MD + ME microphones in total. In the frequency-domain, the m-th micro...
work page Pith review arXiv 1908
-
[3]
BINAURAL MVDR BEAMFORMER The BMVDR [2, 15] aims at minimizing the output noise PSD while preserving the desired speech component in the reference microphone signals (xL and xR), hence preserving the binaural cues of the desired speech source. The optimization problem for the left filter vector wL is given by minwL wH L RnwL subject to wH L aL = 1 . (10) Th...
-
[4]
Using the Cholesky decomposition of the noise covariance matrix, i.e
COV ARIANCE WHITENING METHOD The covariance whitening (CW) method [13, 14] is based on the generalized eigenvalue decomposition of the noisy input covariance matrix Ry and the noise covariance matrixRn. Using the Cholesky decomposition of the noise covariance matrix, i.e. Rn = RH/2 n R1/2 n , (12) the pre-whitened noisy input covariance matrix is defined a...
-
[5]
SPA TIAL COHERENCE METHOD In this section, we propose RTF vector estimation methods that assume that the noise component in each external microphone signal 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics October 20-23, 2019, New Paltz, NY is uncorrelated with the noise components in all other microphone signals. This can, e....
work page 2019
-
[6]
EXPERIMENTAL RESULTS For a dynamic acoustic scenario with a moving speaker in a reverberant room, in this section we compare the performance of the BMVDR using the different RTF vector estimation meth- ods described in Sections 4 and 5 for a binaural hearing device incorporating three external microphones. 6.1. Recording setup and implementation All signa...
work page 2019
-
[7]
Each external microphone was used to obtain an SC-based RTF vector estimate
CONCLUSIONS In this paper, we proposed to use the SC-based RTF vector estimation method for a scenario where multiple external microphones are incorporated into the BMVDR processing of a binaural hearing device. Each external microphone was used to obtain an SC-based RTF vector estimate. We proposed to linearly combine the different RTF vector estimates u...
work page 2019
Show all 25 references
-
[8]
Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diversity using multiple microphones,
S. Doclo, W. Kellermann, S. Makino, and S. E. Nordholm, “Multichannel signal enhancement algorithms for assisted listening devices: Exploiting spatial diversity using multiple microphones,” IEEE Signal Processing Magazine , vol. 32, no. 2, pp. 18–30, Mar. 2015
2015
-
[9]
Binaural speech processing with application to hearing devices,
S. Doclo, S. Gannot, D. Marquardt, and E. Hadad, “Binaural speech processing with application to hearing devices,” in Audio Source Separation and Speech Enhancement . Wiley, 2018, ch. 18, pp. 413–442
2018
-
[10]
Theoretical analysis of binaural multi- microphone noise reduction techniques,
B. Cornelis, S. Doclo, T. Van den Bogaert, J. Wouters, and M. Moonen, “Theoretical analysis of binaural multi- microphone noise reduction techniques,” IEEE Transactions on Audio, Speech and Language Processing, vol. 18, no. 2, pp. 342–355, Feb. 2010
2010
-
[11]
Signal enhance- ment using beamforming and nonstationarity with applications to speech,
S. Gannot, D. Burshtein, and E. Weinstein, “Signal enhance- ment using beamforming and nonstationarity with applications to speech,” IEEE Transactions on Signal Processing , vol. 49, no. 8, pp. 1614–1626, 2001
2001
-
[13]
Binaural noise cue preservation in a binaural noise reduction system with a remote microphone signal,
J. Szurley, A. Bertrand, B. van Dijk, and M. Moonen, “Binaural noise cue preservation in a binaural noise reduction system with a remote microphone signal,” IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 24, no. 5, pp. 952–966, May 2016
2016
-
[14]
Informed sound source localization using relative transfer functions for hearing aid applications,
M. Farmani, M. S. Pedersen, Z.-H. Tan, and J. Jensen, “Informed sound source localization using relative transfer functions for hearing aid applications,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing , vol. 25, no. 3, pp. 611–623, Mar. 2017
2017
-
[15]
A noise reduction post-filter for binaurally-linked single-microphone hearing aids utilizing a nearby external microphone,
D. Yee, H. Kamkar-Parsi, R. Martin, and H. Puder, “A noise reduction post-filter for binaurally-linked single-microphone hearing aids utilizing a nearby external microphone,” IEEE/ACM Transactions on Audio Speech and Language Processing, vol. 26, no. 1, pp. 5–18, Jan. 2018
2018
-
[16]
Generalised sidelobe canceller for noise reduction in hearing devices using an external microphone,
R. Ali, T. van Waterschoot, and M. Moonen, “Generalised sidelobe canceller for noise reduction in hearing devices using an external microphone,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, Canada, Apr. 2018, pp. 521–525
2018
-
[17]
Completing the RTF vector for an MVDR beam- former as applied to a local microphone array and an external microphone,
——, “Completing the RTF vector for an MVDR beam- former as applied to a local microphone array and an external microphone,” in Proc. International Workshop on Acoustic Signal Enhancement (IWAENC), Tokyo, Japan, Sep. 2018, pp. 211–215
2018
-
[18]
Relative transfer function estimation exploiting spatially separated microphones in a diffuse noise field,
N. G ¨oßling and S. Doclo, “Relative transfer function estimation exploiting spatially separated microphones in a diffuse noise field,” inProc. International Workshop on Acoustic Signal En- hancement (IWAENC), Tokyo, Japan, Sep. 2018, pp. 146–150
2018
-
[19]
RTF-steered binaural MVDR beamforming incorporat- ing an external microphone for dynamic acoustic scenarios,
——, “RTF-steered binaural MVDR beamforming incorporat- ing an external microphone for dynamic acoustic scenarios,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Brighton, UK, May 2019, pp. 416–420
2019
-
[20]
Multichannel eigenspace beamforming in a reverberant noisy environment with multiple interfering speech signals,
S. Markovich, S. Gannot, and I. Cohen, “Multichannel eigenspace beamforming in a reverberant noisy environment with multiple interfering speech signals,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 17, no. 6, pp. 1071–1086, Aug. 2009
2009
-
[21]
Performance analysis of the covariance subtraction method for relative transfer function estimation and comparison to the covariance whiten- ing method,
S. Markovich-Golan and S. Gannot, “Performance analysis of the covariance subtraction method for relative transfer function estimation and comparison to the covariance whiten- ing method,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASS...
2015
-
[22]
Acoustic beamforming for hearing aid applications,
S. Doclo, S. Gannot, M. Moonen, and A. Spriet, “Acoustic beamforming for hearing aid applications,” in Handbook on Array Processing and Sensor Networks . Wiley, 2010, pp. 269–302
2010
-
[23]
Unbiased MMSE-based noise power estimation with low complexity and low tracking delay,
T. Gerkmann and R. C. Hendriks, “Unbiased MMSE-based noise power estimation with low complexity and low tracking delay,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, no. 4, pp. 1383–1393, May 2012
2012
-
[24]
Reference microphone selection for MWF-based noise reduction using distributed microphone arrays,
T. C. Lawin-Ore and S. Doclo, “Reference microphone selection for MWF-based noise reduction using distributed microphone arrays,” in Proc. ITG Conference on Speech Communication, Braunschweig, Germany, Sep. 2012, pp. 1–4
2012
-
[25]
G ¨oßling, W
N. G ¨oßling, W. Middelberg, and S. Doclo. (2019) RTF-steered binaural MVDR beamforming incorporating multiple external microphones. [Online]. Available: https://uol.de/en/sigproc/ research/audio-demos/binaural-noise-reduction/
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.