Pith. sign in

REVIEW 3 major objections 6 minor 11 references

Improved in-car sound pick-up using multichannel Wiener filter

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Extracting each talker with a multichannel Wiener filter before summing removes echo notches and cuts background noise.

desk verdict A standard MWF application with a genuinely new use case, but the SIR metric that anchors its central claim is ill-defined under the paper's own non-overlapping speech protocol. read the letter →

arxiv 2506.11157 v1 pith:NBD7LJMR submitted 2025-06-11 eess.AS eess.SP

classification eess.ASeess.SP
keywords multichannelWienerfilterin-carspeechenhancementnotchfilteringeffectnoisereductionDNSMOShands-freetelephonysourceextractionmicrophonearray
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes using a multichannel Wiener filter (MWF) on a two-microphone in-car setup to extract driver and passenger speech separately, then sum the extracted sources. It aims to show that this beats the common practice of simply adding the two microphone signals, which creates a comb-like notch filtering distortion from the two delayed acoustic paths and doubles background noise. The paper claims the MWF approach removes most of these spectral notches, especially with larger processing frame sizes, and improves signal-to-noise ratio, signal-to-interference ratio, and perceptual DNSMOS scores across white, red, pink, green, and hoth noise. It also claims the filter adapts quickly enough to tolerate head movements, provided adaptation is never paused.

What carries the argument

The machinery is the frequency-domain multichannel Wiener filter: for each target source (driver at microphone 1, passenger at microphone 2), a filter $\mathbf{w}(\omega)=\mathbf{R}^{-1}(\omega)\mathbf{p}(\omega)$ is computed from the multichannel correlation matrix $\mathbf{R}$ of the microphone signals and the cross-correlation vector $\mathbf{p}$ with the target source, with Eqs. (11) and (13) extracting the speech-only correlations from periods when only one source is active plus noise-only statistics. Regularization $\delta\,\mathrm{tr}(\mathbf{R})/\mathrm{size}(\mathbf{R})\,\mathbf{I}$ keeps the matrix inverse stable, and a moving-average forgetting factor $\lambda$ lets the filter track changing source positions. The two extracted sources are then simply added, which is what removes the notch filtering and reduces background noise.

What would settle it

Run the identical two-microphone setup with measured car impulse responses, replace the perfect classifier with a standard voice-activity detector, and include overlapping driver-and-passenger speech segments; the central claim is settled by checking whether SNR, SIR, and DNSMOS gains still exceed the simple microphone sum, which loses about 1 dB in SNR.

Watch

Extended reading notes

Core claim

The central assertion is that an adaptive multichannel Wiener filter, applied before any mixing, can extract each talker's contribution at its own dedicated microphone and suppress the cross-talk path, so the two extracted signals can be added without the comb-filter notches that appear when raw microphone signals are summed. In the tested simulated car cabin, MWF gives SNR gains of roughly 10 dB for white noise and 6-9 dB for red, pink, green, and hoth noises, while raw microphone summing loses about 1 dB; DNSMOS speech and noise scores are consistently higher with MWF. The paper further states that the filter works with one or two speakers active and that speaker head movements of 0.1-0.15 m cause almost no performance drop as long as adaptation continues every 8 ms.

Load-bearing premise

The load-bearing premise, stated in Section III, is that a perfect classifier always knows which speaker is active, and the test signals never overlap; if voice-activity detection is imperfect or driver and passenger speak simultaneously, the correlation matrices feeding the MWF become biased and the reported gains may not hold.

Editorial extensions

If this is right

  • Hands-free telephony and voice-command systems can place MWF directly after microphone pickup, before acoustic echo cancellation, with no direction-of-arrival or propagation model.
  • Because the extracted sources are combined before transmission, the downstream echo path stays stable, avoiding the discontinuities caused by microphone switching.
  • The same filtering handles single-talker and simultaneous-talker scenarios, and in the simultaneous case it can act as interference cancellation between driver and passenger.
  • Larger frame sizes (up to 100 ms) remove nearly all spectral notches for stationary sources, while speech needs smaller frames (8 ms) to track changing statistics; both regimes are demonstrated.
  • When loudspeaker or media signals are present, MWF must be disabled or paired with per-microphone acoustic echo cancellation, since the number of extractable sources is capped by the number of microphones.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If real voice-activity detection replaces the paper's perfect-classification assumption, correlation estimates will be biased and the reported SNR, SIR, and DNSMOS gains are likely to shrink; the perfect-detection setup thus marks an upper bound for practical systems.
  • The simulated cabin uses symmetric impulse responses with 2 dB cross-side attenuation; real cabins with stronger or asymmetric cross-talk could show either larger benefits from MWF or require different frame sizes and regularization.
  • The paper claims MWF handles simultaneous talkers but only tests non-overlapping speech segments; a direct overlapping-speech experiment would settle whether the two filters truly separate talkers or merely noise-reduce the mixture.
  • With more than two microphones, the same correlation-estimation framework could extract additional sources or improve robustness, an extension the paper mentions but does not test.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript applies an adaptive multichannel Wiener filter (MWF) to two in-car microphones, with the goal of extracting the driver and passenger speech sources separately and then summing the two extracted signals. The authors argue that this avoids the spectral notch filtering and noise-power doubling that occur when the two raw microphone signals are simply added. After presenting the standard frequency-domain MWF equations with recursive correlation estimation and diagonal loading, the paper describes a simulated rectangular-car experiment with perfect active-speaker detection, non-overlapping driver and passenger speech, and five noise types. Results in Tables II and III report SNR gain, SIR gain, and DNSMOS for the sum of MWF outputs versus the raw microphone sum, and Figs. 4-7 address notch removal and head-movement adaptation. The paper concludes that MWF gives significant improvements over simple mixing.

Significance. If the reported results withstand scrutiny, the paper provides a useful engineering demonstration that a standard MWF can be inserted before microphone-signal combining in an in-car two-talker setup, reducing notch distortion and background noise while allowing the two extracted signals to be added without echo-path switching. The comparison against a simple-sum baseline is appropriate, the MWF formulation is standard and self-consistent, and the authors are transparent about the perfect active-speaker-detection assumption. The DNSMOS evaluation and the head-movement adaptation study are useful additions. However, the SIR metric as defined is not computable under the non-overlapping-speech protocol, and the hyperparameters are selected on the same data used for evaluation; both issues currently prevent the quantitative claims from being taken at face value. The single simulated room and the absence of overlapping speech also limit generalization.

major comments (3)
  1. [III, 'Evaluation metrics'; Tables II and III] The reported finite SIR gains, e.g., 8.28 dB for the driver with white noise in Table II, are not consistent with the stated definition and protocol. Because the driver and passenger speech segments do not overlap, in any driver-only segment the passenger source is silent and the input SIR is infinite; the same holds for passenger-only segments. With the manuscript's statement that 'SIR and SNR gain for driver and passenger speech are calculated from their respective time segments,' a finite output SIR cannot produce the finite SIR gains shown in Tables II and III. If the authors instead computed SIR over the full recording, the input SIR would be close to 0 dB because the two non-overlapping equal-power speech segments each act as interference for the other, which contradicts the 'respective time segments' wording. Since SIR is the only metric that directly quantifies source decoupling, this ambiguity undermines the central quantitative evidence for the claimed interference-mitigation benefit. Please redefine the SIR gain with a well-specified reference (e.g., the full recording, or a defined floor for absent competing sources) and recompute Tables II and III, or add an overlapping-speech experiment in which the input SIR is finite by construction.
  2. [IV.B] The frame size (8 ms), the regularization factor delta, and the forgetting factor lambda were selected as the values that 'provided the best noise reduction' on the same recordings used for Tables II and III, and a hoth-noise-specific low-frequency regularization was also chosen. This is test-set hyperparameter fitting, so the reported SNR, SIR, and DNSMOS gains over simple mixing are optimistic and the sensitivity of the conclusions to these settings is unknown. Please add a development/test split, or alternatively report the metric values over the parameter grid and show that the qualitative improvement over simple microphone summing holds for a range of reasonable settings.
  3. [II, III, and VI] The paper claims in Sections II and VI that the MWF can extract driver and passenger speech whether only one talker is active or both are talking simultaneously, and that it can be used for interference cancellation in the simultaneous case. However, Section III explicitly states that 'there is no overlap between speech segments in the experiments,' and no experiment in Section IV exercises two active speakers. The non-overlapping protocol therefore provides no evidence for the simultaneous-speaker or interference-cancellation claims. Please either temper these claims to the single-talker-plus-noise conditions actually tested, or add an overlapping-speech condition with different utterances for the two sources and report SNR gain, SIR gain, and DNSMOS for that condition.
minor comments (6)
  1. [II, Eq. (9)] The second element of the cross-correlation vector p(omega) appears to be printed as Phi_{x1}(omega), but from the right-hand side and from Eq. (12) it should be Phi_{x2 d}(omega); please correct this typo.
  2. [II, Eqs. (5)-(7)] The notation relating w, w^H, and the conjugated filter coefficients in Eqs. (5)-(7) is inconsistent and should be normalized so that the filter vector used in the inner product is unambiguously defined.
  3. [IV.C] The text states that with white noise and 5 dB input SNR the driver's SNR gain decreases from 7.68 to 7.47 dB, but Table II reports a driver SNR gain of 10.14 dB for the same nominal condition; please clarify which parameter settings or averaging windows produce each number.
  4. [Fig. 2] The time-axis labels in Fig. 2 appear garbled (e.g., '0 1 02 03 04 05 06 07 0') and are inconsistent with the 68-second intervals described in Section IV.C; please regenerate the figure with correct axis formatting.
  5. [IV.C] There is a grammatical typo in 'Test were also performed with adaptation paused after 36 seconds'; it should be 'Tests were also performed.'
  6. [III, 'Evaluation metrics'] The sentence defining SIR says it 'assesses the decoupling of the speech and driver sources,' but it should refer to the driver and passenger sources; please fix the wording.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the MWF derivation and evaluation are self-contained, though the SIR metric definition is internally inconsistent (a correctness concern, not a circular one).

full rationale

The derivation chain is self-contained: Eq. (6), w = R^-1 p, is the standard MWF solution, with R and p estimated from microphone statistics in Eqs. (10)-(13), and the evaluated output is the sum of the two MWF source estimates compared against an independent 'simple mixing' baseline. Nothing in the equations forces the claimed SNR, SIR, or DNSMOS improvements; the comparison could in principle have favored the baseline, so the central comparison is not circular. The only self-referential aspect is that MWF is designed to estimate the source at a single reference microphone, so removing the second propagation path (and hence the notch) is the mechanism of the method, not a hidden equivalence between input and output. Section III explicitly states the perfect active-speaker-detection assumption ('a perfect classification assumption is made'), and Section IV.B selects the frame size, regularization factor, and forgetting factor from the test data; these are validity limitations, not circular steps. A separate, non-circular correctness concern: Section III defines SIR gain as the ratio of output SIR to input SIR and states that 'SIR and SNR gain for driver and passenger speech are calculated from their respective time segments,' while also stating that 'there is no overlap between speech segments in the experiments.' For a driver-only segment the passenger source is silent, so the input SIR is infinite; finite SIR gains in Tables II and III (e.g., 8.28 dB) cannot follow from the stated definition unless an unstated full-recording computation is used. This makes the SIR-based evidence internally ambiguous, but it is a metric-definition issue rather than a circular derivation.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim depends on a standard algorithm, but the reported best-case performance is obtained after tuning frame size, forgetting factor, and regularization on the evaluation data. The simulation also assumes perfect voice-activity detection, non-overlapping speech, and an idealized rectangular cabin impulse response. No new entities are introduced.

free parameters (3)
  • MWF regularization factor delta = 1.0 (100 below 312.5 Hz for hoth noise)
    Chosen to give the best noise reduction across noise types; tuned on the evaluation data (Section IV-B).
  • MWF forgetting factor lambda = 0.96
    Selected as providing the best noise reduction in the reported simulations (Section IV-B).
  • Frame size = 8 ms
    Selected as best for noise reduction; larger frames remove notches better but adapt more slowly (Section IV-B).
assumptions (5)
  • domain assumption Driver and passenger speech sources are mutually uncorrelated and uncorrelated with background noise (Eqs. 3 and 4).
    Required for the MWF correlation matrix estimates to separate sources; reasonable for the simulation but not validated in real car conditions.
  • ad hoc to paper Perfect active speaker detection is assumed in the experiments.
    Stated in Section III to decouple detection errors from MWF performance; real cars will have imperfect detection, especially with overlapping speech.
  • ad hoc to paper Driver and passenger speech segments never overlap in the test data.
    Stated in Section III as a design choice; the conclusion nevertheless claims simultaneous-talker capability that is never tested.
  • domain assumption A rectangular image-method simulation accurately approximates the car cabin, with specified dimensions, reverberation time, and cardioid microphones.
    Standard room acoustics simulation, but it may not capture real cabin geometry, trim materials, or microphone variability.
  • domain assumption The number of acoustic sources must not exceed the number of microphones for MWF to work effectively, so loudspeaker signals are absent or require AEC.
    Stated in the introduction and conclusion; it limits applicability when media is playing unless additional processing is used.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improved in-car sound pick-up using multichannel Wiener filter." pith.science (2026). https://pith.science/paper/NBD7LJMR

@misc{pith2026250611157,
  author       = {Pith},
  title        = {Pith review of: Improved in-car sound pick-up using multichannel Wiener filter},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NBD7LJMR}},
  note         = {Machine review of arXiv:2506.11157}
}
read the original abstract

With advancements in automotive electronics and sensors, the sound pick-up using multiple microphones has become feasible for hands-free telephony and voice command in-car applications. However, challenges remain in effectively processing multiple microphone signals due to bandwidth or processing limitations. This work explores the use of the Multichannel Wiener Filter algorithm with a two-microphone in-car system, to enhance speech quality for driver and passenger voice, i.e., to mitigate notch-filtering effects caused by echoes and improve background noise reduction. We evaluate its performance under various noise conditions using modern objective metrics like Deep Noise Suppression Mean Opinion Score. The effect of head movements of driver/passenger is also investigated. The proposed method is shown to provide significant improvements over a simple mixing of microphone signals.

Figures

Figures reproduced from arXiv: 2506.11157 by the authors.

Figure 7
Figure 7. DNSMOS performance scores for MWF outputs sum and simple [PITH_FULL_IMAGE:figures/full_fig_p006_7.png] view at source ↗
Figure 6
Figure 6. Performance scores for MWF outputs sum with head movement for different input SNRs and driver movements, MWF adaptation paused after 36 seconds. single AEC. To resolve these issues, we investigated the use of an adaptive MWF filter to separate driver and passenger signals before combining them, thereby reducing notch filtering effects and minimizing background noise. The MWF proved effective in a car environment, wh… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references · 10 canonical work pages

  1. [1]

    Multi-channel speech enhancement in a car environment using Wiener filtering and spectral subtraction,

    J. Meyer and K. U. Simmer, "Multi-channel speech enhancement in a car environment using Wiener filtering and spectral subtraction," in Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Munich, Germany, 1997, pp. 1167-1170

  2. [2]

    A subband hybrid beamforming for in-car speech e nhancement,

    C. Fox, G. Vitte, M. Charbit, J. Pr ado, R. Badeau and B. David, "A subband hybrid beamforming for in-car speech e nhancement," in Proc. 20th European Signal Processing Conference (EUSIPCO 2012) , Bucharest, Romania, 2012, pp. 11-15

  3. [3]

    On the Speech Distortion Weighted Multichannel Wiener Filter for diffuse Noise,

    S. Stenzel and J. Freudenberger, "On the Speech Distortion Weighted Multichannel Wiener Filter for diffuse Noise," in Proc. 10th ITG Symposium on Speech Communication, Braunschweig, Germany, 2012, pp. 1-4

  4. [4]

    A Multichannel Spatial Hands- Free Application for In-Car Communication Systems,

    M. Gimm, F. Kühne and G. Schmidt, "A Multichannel Spatial Hands- Free Application for In-Car Communication Systems," in Towards Human-Vehicle Harmonization , Germany: De Gruyter, 2023, ch. 10, pp.129-140

  5. [5]

    Multichannel Wiener filter in active sound-navigation-and-ranging systems - A joint beamformer and matched filter approach,

    B. Kaulen, J. Abshagen, and G. Schm idt, "Multichannel Wiener filter in active sound-navigation-and-ranging systems - A joint beamformer and matched filter approach, " IET Radar, Sonar & Navigation, vol. 18, no. 9, pp.1554-1569, Sept. 2024

  6. [6]

    Adaptive Beamforming and Postfiltering,

    S. Gannot and I. Cohen, "Adaptive Beamforming and Postfiltering," in Springer Handbook of Speech Processing, Germany: Springer, 2008, ch. 47, pp. 945-978

  7. [7]

    Speech enhancement with multichannel Wiener filter techniques in multimicrophone binaural hearing aids,

    T. Van den Bogaert, S. Doclo, J. Wouters and M. Moonen, "Speech enhancement with multichannel Wiener filter techniques in multimicrophone binaural hearing aids," The Journal of the Acoustical Society of America, vol. 125, no.1, pp. 360-371, Jan. 2009

  8. [8]

    Available: https://github.com/ehabets/RIR- Generator

    RIR-Generator [Online]. Available: https://github.com/ehabets/RIR- Generator

Show all 11 references
  1. [9]

    Image method for efficiently simulating small-room acoustics,

    J.B. Allen and D.A. Berkley, "Image method for efficiently simulating small-room acoustics," The Journal of the Acoustical Society of America, vol. 65, no. 4, pp. 943-950, April 1979

  2. [10]

    Room noise spectra at subscribers' telephone locations,

    D.F. Hoth, "Room noise spectra at subscribers' telephone locations," The Journal of the Acoustical Society of America, vol. 12, no. 3, pp. 499-504, Apr. 1941

  3. [11]

    DNSMOS P.835: A Non- Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,

    C. K. A. Reddy, V. Gopal, and R. Cutler, "DNSMOS P.835: A Non- Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors," arXiv preprint arXiv:2110.01763 , pp.1-5, Feb. 2022. [Online]. Available: https://arxiv.org/abs/2110.01763 DNSMOS SIG DNSMOS BAK ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.