REVIEW 3 major objections 6 minor 11 references
Improved in-car sound pick-up using multichannel Wiener filter
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Extracting each talker with a multichannel Wiener filter before summing removes echo notches and cuts background noise.
desk verdict A standard MWF application with a genuinely new use case, but the SIR metric that anchors its central claim is ill-defined under the paper's own non-overlapping speech protocol. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the frequency-domain multichannel Wiener filter: for each target source (driver at microphone 1, passenger at microphone 2), a filter $\mathbf{w}(\omega)=\mathbf{R}^{-1}(\omega)\mathbf{p}(\omega)$ is computed from the multichannel correlation matrix $\mathbf{R}$ of the microphone signals and the cross-correlation vector $\mathbf{p}$ with the target source, with Eqs. (11) and (13) extracting the speech-only correlations from periods when only one source is active plus noise-only statistics. Regularization $\delta\,\mathrm{tr}(\mathbf{R})/\mathrm{size}(\mathbf{R})\,\mathbf{I}$ keeps the matrix inverse stable, and a moving-average forgetting factor $\lambda$ lets the filter track changing source positions. The two extracted sources are then simply added, which is what removes the notch filtering and reduces background noise.
What would settle it
Run the identical two-microphone setup with measured car impulse responses, replace the perfect classifier with a standard voice-activity detector, and include overlapping driver-and-passenger speech segments; the central claim is settled by checking whether SNR, SIR, and DNSMOS gains still exceed the simple microphone sum, which loses about 1 dB in SNR.
Extended reading notes
Core claim
The central assertion is that an adaptive multichannel Wiener filter, applied before any mixing, can extract each talker's contribution at its own dedicated microphone and suppress the cross-talk path, so the two extracted signals can be added without the comb-filter notches that appear when raw microphone signals are summed. In the tested simulated car cabin, MWF gives SNR gains of roughly 10 dB for white noise and 6-9 dB for red, pink, green, and hoth noises, while raw microphone summing loses about 1 dB; DNSMOS speech and noise scores are consistently higher with MWF. The paper further states that the filter works with one or two speakers active and that speaker head movements of 0.1-0.15 m cause almost no performance drop as long as adaptation continues every 8 ms.
Load-bearing premise
The load-bearing premise, stated in Section III, is that a perfect classifier always knows which speaker is active, and the test signals never overlap; if voice-activity detection is imperfect or driver and passenger speak simultaneously, the correlation matrices feeding the MWF become biased and the reported gains may not hold.
Editorial extensions
If this is right
- Hands-free telephony and voice-command systems can place MWF directly after microphone pickup, before acoustic echo cancellation, with no direction-of-arrival or propagation model.
- Because the extracted sources are combined before transmission, the downstream echo path stays stable, avoiding the discontinuities caused by microphone switching.
- The same filtering handles single-talker and simultaneous-talker scenarios, and in the simultaneous case it can act as interference cancellation between driver and passenger.
- Larger frame sizes (up to 100 ms) remove nearly all spectral notches for stationary sources, while speech needs smaller frames (8 ms) to track changing statistics; both regimes are demonstrated.
- When loudspeaker or media signals are present, MWF must be disabled or paired with per-microphone acoustic echo cancellation, since the number of extractable sources is capped by the number of microphones.
Reading between the lines
- If real voice-activity detection replaces the paper's perfect-classification assumption, correlation estimates will be biased and the reported SNR, SIR, and DNSMOS gains are likely to shrink; the perfect-detection setup thus marks an upper bound for practical systems.
- The simulated cabin uses symmetric impulse responses with 2 dB cross-side attenuation; real cabins with stronger or asymmetric cross-talk could show either larger benefits from MWF or require different frame sizes and regularization.
- The paper claims MWF handles simultaneous talkers but only tests non-overlapping speech segments; a direct overlapping-speech experiment would settle whether the two filters truly separate talkers or merely noise-reduce the mixture.
- With more than two microphones, the same correlation-estimation framework could extract additional sources or improve robustness, an extension the paper mentions but does not test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript applies an adaptive multichannel Wiener filter (MWF) to two in-car microphones, with the goal of extracting the driver and passenger speech sources separately and then summing the two extracted signals. The authors argue that this avoids the spectral notch filtering and noise-power doubling that occur when the two raw microphone signals are simply added. After presenting the standard frequency-domain MWF equations with recursive correlation estimation and diagonal loading, the paper describes a simulated rectangular-car experiment with perfect active-speaker detection, non-overlapping driver and passenger speech, and five noise types. Results in Tables II and III report SNR gain, SIR gain, and DNSMOS for the sum of MWF outputs versus the raw microphone sum, and Figs. 4-7 address notch removal and head-movement adaptation. The paper concludes that MWF gives significant improvements over simple mixing.
Significance. If the reported results withstand scrutiny, the paper provides a useful engineering demonstration that a standard MWF can be inserted before microphone-signal combining in an in-car two-talker setup, reducing notch distortion and background noise while allowing the two extracted signals to be added without echo-path switching. The comparison against a simple-sum baseline is appropriate, the MWF formulation is standard and self-consistent, and the authors are transparent about the perfect active-speaker-detection assumption. The DNSMOS evaluation and the head-movement adaptation study are useful additions. However, the SIR metric as defined is not computable under the non-overlapping-speech protocol, and the hyperparameters are selected on the same data used for evaluation; both issues currently prevent the quantitative claims from being taken at face value. The single simulated room and the absence of overlapping speech also limit generalization.
major comments (3)
- [III, 'Evaluation metrics'; Tables II and III] The reported finite SIR gains, e.g., 8.28 dB for the driver with white noise in Table II, are not consistent with the stated definition and protocol. Because the driver and passenger speech segments do not overlap, in any driver-only segment the passenger source is silent and the input SIR is infinite; the same holds for passenger-only segments. With the manuscript's statement that 'SIR and SNR gain for driver and passenger speech are calculated from their respective time segments,' a finite output SIR cannot produce the finite SIR gains shown in Tables II and III. If the authors instead computed SIR over the full recording, the input SIR would be close to 0 dB because the two non-overlapping equal-power speech segments each act as interference for the other, which contradicts the 'respective time segments' wording. Since SIR is the only metric that directly quantifies source decoupling, this ambiguity undermines the central quantitative evidence for the claimed interference-mitigation benefit. Please redefine the SIR gain with a well-specified reference (e.g., the full recording, or a defined floor for absent competing sources) and recompute Tables II and III, or add an overlapping-speech experiment in which the input SIR is finite by construction.
- [IV.B] The frame size (8 ms), the regularization factor delta, and the forgetting factor lambda were selected as the values that 'provided the best noise reduction' on the same recordings used for Tables II and III, and a hoth-noise-specific low-frequency regularization was also chosen. This is test-set hyperparameter fitting, so the reported SNR, SIR, and DNSMOS gains over simple mixing are optimistic and the sensitivity of the conclusions to these settings is unknown. Please add a development/test split, or alternatively report the metric values over the parameter grid and show that the qualitative improvement over simple microphone summing holds for a range of reasonable settings.
- [II, III, and VI] The paper claims in Sections II and VI that the MWF can extract driver and passenger speech whether only one talker is active or both are talking simultaneously, and that it can be used for interference cancellation in the simultaneous case. However, Section III explicitly states that 'there is no overlap between speech segments in the experiments,' and no experiment in Section IV exercises two active speakers. The non-overlapping protocol therefore provides no evidence for the simultaneous-speaker or interference-cancellation claims. Please either temper these claims to the single-talker-plus-noise conditions actually tested, or add an overlapping-speech condition with different utterances for the two sources and report SNR gain, SIR gain, and DNSMOS for that condition.
minor comments (6)
- [II, Eq. (9)] The second element of the cross-correlation vector p(omega) appears to be printed as Phi_{x1}(omega), but from the right-hand side and from Eq. (12) it should be Phi_{x2 d}(omega); please correct this typo.
- [II, Eqs. (5)-(7)] The notation relating w, w^H, and the conjugated filter coefficients in Eqs. (5)-(7) is inconsistent and should be normalized so that the filter vector used in the inner product is unambiguously defined.
- [IV.C] The text states that with white noise and 5 dB input SNR the driver's SNR gain decreases from 7.68 to 7.47 dB, but Table II reports a driver SNR gain of 10.14 dB for the same nominal condition; please clarify which parameter settings or averaging windows produce each number.
- [Fig. 2] The time-axis labels in Fig. 2 appear garbled (e.g., '0 1 02 03 04 05 06 07 0') and are inconsistent with the 68-second intervals described in Section IV.C; please regenerate the figure with correct axis formatting.
- [IV.C] There is a grammatical typo in 'Test were also performed with adaptation paused after 36 seconds'; it should be 'Tests were also performed.'
- [III, 'Evaluation metrics'] The sentence defining SIR says it 'assesses the decoupling of the speech and driver sources,' but it should refer to the driver and passenger sources; please fix the wording.
Circularity Check
No significant circularity; the MWF derivation and evaluation are self-contained, though the SIR metric definition is internally inconsistent (a correctness concern, not a circular one).
full rationale
The derivation chain is self-contained: Eq. (6), w = R^-1 p, is the standard MWF solution, with R and p estimated from microphone statistics in Eqs. (10)-(13), and the evaluated output is the sum of the two MWF source estimates compared against an independent 'simple mixing' baseline. Nothing in the equations forces the claimed SNR, SIR, or DNSMOS improvements; the comparison could in principle have favored the baseline, so the central comparison is not circular. The only self-referential aspect is that MWF is designed to estimate the source at a single reference microphone, so removing the second propagation path (and hence the notch) is the mechanism of the method, not a hidden equivalence between input and output. Section III explicitly states the perfect active-speaker-detection assumption ('a perfect classification assumption is made'), and Section IV.B selects the frame size, regularization factor, and forgetting factor from the test data; these are validity limitations, not circular steps. A separate, non-circular correctness concern: Section III defines SIR gain as the ratio of output SIR to input SIR and states that 'SIR and SNR gain for driver and passenger speech are calculated from their respective time segments,' while also stating that 'there is no overlap between speech segments in the experiments.' For a driver-only segment the passenger source is silent, so the input SIR is infinite; finite SIR gains in Tables II and III (e.g., 8.28 dB) cannot follow from the stated definition unless an unstated full-recording computation is used. This makes the SIR-based evidence internally ambiguous, but it is a metric-definition issue rather than a circular derivation.
Assumptions & free parameters
free parameters (3)
- MWF regularization factor delta =
1.0 (100 below 312.5 Hz for hoth noise)
- MWF forgetting factor lambda =
0.96
- Frame size =
8 ms
assumptions (5)
- domain assumption Driver and passenger speech sources are mutually uncorrelated and uncorrelated with background noise (Eqs. 3 and 4).
- ad hoc to paper Perfect active speaker detection is assumed in the experiments.
- ad hoc to paper Driver and passenger speech segments never overlap in the test data.
- domain assumption A rectangular image-method simulation accurately approximates the car cabin, with specified dimensions, reverberation time, and cardioid microphones.
- domain assumption The number of acoustic sources must not exceed the number of microphones for MWF to work effectively, so loudspeaker signals are absent or require AEC.
Cite this review
Pith. "Pith review of Improved in-car sound pick-up using multichannel Wiener filter." pith.science (2026). https://pith.science/paper/NBD7LJMR
@misc{pith2026250611157,
author = {Pith},
title = {Pith review of: Improved in-car sound pick-up using multichannel Wiener filter},
year = {2026},
howpublished = {\url{https://pith.science/paper/NBD7LJMR}},
note = {Machine review of arXiv:2506.11157}
}
read the original abstract
With advancements in automotive electronics and sensors, the sound pick-up using multiple microphones has become feasible for hands-free telephony and voice command in-car applications. However, challenges remain in effectively processing multiple microphone signals due to bandwidth or processing limitations. This work explores the use of the Multichannel Wiener Filter algorithm with a two-microphone in-car system, to enhance speech quality for driver and passenger voice, i.e., to mitigate notch-filtering effects caused by echoes and improve background noise reduction. We evaluate its performance under various noise conditions using modern objective metrics like Deep Noise Suppression Mean Opinion Score. The effect of head movements of driver/passenger is also investigated. The proposed method is shown to provide significant improvements over a simple mixing of microphone signals.
Figures
Reference graph
Works this paper leans on
-
[1]
J. Meyer and K. U. Simmer, "Multi-channel speech enhancement in a car environment using Wiener filtering and spectral subtraction," in Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Munich, Germany, 1997, pp. 1167-1170
work page 1997
-
[2]
A subband hybrid beamforming for in-car speech e nhancement,
C. Fox, G. Vitte, M. Charbit, J. Pr ado, R. Badeau and B. David, "A subband hybrid beamforming for in-car speech e nhancement," in Proc. 20th European Signal Processing Conference (EUSIPCO 2012) , Bucharest, Romania, 2012, pp. 11-15
work page 2012
-
[3]
On the Speech Distortion Weighted Multichannel Wiener Filter for diffuse Noise,
S. Stenzel and J. Freudenberger, "On the Speech Distortion Weighted Multichannel Wiener Filter for diffuse Noise," in Proc. 10th ITG Symposium on Speech Communication, Braunschweig, Germany, 2012, pp. 1-4
work page 2012
-
[4]
A Multichannel Spatial Hands- Free Application for In-Car Communication Systems,
M. Gimm, F. Kühne and G. Schmidt, "A Multichannel Spatial Hands- Free Application for In-Car Communication Systems," in Towards Human-Vehicle Harmonization , Germany: De Gruyter, 2023, ch. 10, pp.129-140
work page 2023
-
[5]
B. Kaulen, J. Abshagen, and G. Schm idt, "Multichannel Wiener filter in active sound-navigation-and-ranging systems - A joint beamformer and matched filter approach, " IET Radar, Sonar & Navigation, vol. 18, no. 9, pp.1554-1569, Sept. 2024
work page 2024
-
[6]
Adaptive Beamforming and Postfiltering,
S. Gannot and I. Cohen, "Adaptive Beamforming and Postfiltering," in Springer Handbook of Speech Processing, Germany: Springer, 2008, ch. 47, pp. 945-978
work page 2008
-
[7]
T. Van den Bogaert, S. Doclo, J. Wouters and M. Moonen, "Speech enhancement with multichannel Wiener filter techniques in multimicrophone binaural hearing aids," The Journal of the Acoustical Society of America, vol. 125, no.1, pp. 360-371, Jan. 2009
work page 2009
-
[8]
Available: https://github.com/ehabets/RIR- Generator
RIR-Generator [Online]. Available: https://github.com/ehabets/RIR- Generator
Show all 11 references
-
[9]
Image method for efficiently simulating small-room acoustics,
J.B. Allen and D.A. Berkley, "Image method for efficiently simulating small-room acoustics," The Journal of the Acoustical Society of America, vol. 65, no. 4, pp. 943-950, April 1979
1979
-
[10]
Room noise spectra at subscribers' telephone locations,
D.F. Hoth, "Room noise spectra at subscribers' telephone locations," The Journal of the Acoustical Society of America, vol. 12, no. 3, pp. 499-504, Apr. 1941
1941
-
[11]
DNSMOS P.835: A Non- Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,
C. K. A. Reddy, V. Gopal, and R. Cutler, "DNSMOS P.835: A Non- Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors," arXiv preprint arXiv:2110.01763 , pp.1-5, Feb. 2022. [Online]. Available: https://arxiv.org/abs/2110.01763 DNSMOS SIG DNSMOS BAK ...
2022 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.