{"id":"0fdd9452-d59d-4a45-b4fa-4e6683d851a1","arxiv_id":"2506.11157","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In a simulated car, a multichannel Wiener filter separates two speakers and suppresses background noise, outperforming the simple sum of two microphone signals.","lead":"This paper tests a multichannel Wiener filter in a simulated two-microphone car cabin, showing it can separate driver and passenger voices while reducing echo-induced distortion and background noise. The method could make hands-free calls and voice commands in vehicles sound clearer without requiring expensive hardware.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SIR gain metric is ill-defined under the non-overlapping-speech protocol, leaving the paper's quantitative evidence for its novel source-decoupling claim unverifiable.","rationale":"The reader's verdict is CONDITIONAL, and my concern reinforces that conditionality rather than overturning it. The central claim is that the sum of MWF outputs significantly improves over simple microphone mixing. SNR and DNSMOS results do support some improvement, so the paper should not be rejected outright. However, the SIR gain, which is the metric most specific to the paper's novel claim of decoupling and notch mitigation, is not well-defined under the stated non-overlapping-speech protocol. This is an internal inconsistency in the evaluation, not merely a question of real-world generalization. The reader's weakest assumption focused on perfect voice-activity detection and untested simultaneous speech; my concern is related but distinct, as it affects even the non-overlapping results that were actually run. A concrete recomputation with a well-defined input SIR would settle whether the reported SIR gains are meaningful. Pending that check, the appropriate verdict remains CONDITIONAL, so no change from the reader's verdict is needed.","tokens_in":8891,"tokens_out":9464,"duration_ms":123631,"concrete_test":"Recompute the SIR gains from the same simulated microphone signals with a well-defined input SIR. Recommended: (1) For each target segment, add the competing speaker's signal at equal power (or at the same input SIR used for SNR) so the denominator is nonzero, then compute output SIR for the MWF-output sum and for the microphone sum; (2) Alternatively, compute SIR over the full recording using total source energies. If the MWF-versus-mic SIR gain is not reproduced under either well-defined definition, the reported SIR gains are an artifact of the metric rather than evidence for the method.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III defines SIR as the ratio between output SIR and input SIR, where each SIR is the ratio of target speech power over competing speech power. Section III/IV then states that driver and passenger speech are intermittent and that there is no overlap between speech segments. For a driver-only time segment, the competing passenger source is silent, so the input SIR is infinite; the same holds for the passenger-only segment. The finite SIR gains reported in Tables II and III (e.g., 8.28 dB for the driver with white noise) therefore cannot follow from the stated definition. If the authors instead computed SIR over the full recording, where the two equal-length non-overlapping source segments make the input SIR roughly 0 dB, that contradicts the statement that gains are 'calculated from their respective time segments.' The SIR metric is the only quantitative measure that directly targets the paper's claimed contribution: cross-path/notch mitigation and speaker decoupling. Its current definition leaves the central evidence ambiguous. The perfect-VAD assumption and test-set hyperparameter tuning, noted by the reader, are additional limitations, but this SIR issue is internal to the reported results and should be resolved before the improvement claim is accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript applies an adaptive multichannel Wiener filter (MWF) to two in-car microphones, with the goal of extracting the driver and passenger speech sources separately and then summing the two extracted signals. The authors argue that this avoids the spectral notch filtering and noise-power doubling that occur when the two raw microphone signals are simply added. After presenting the standard frequency-domain MWF equations with recursive correlation estimation and diagonal loading, the paper describes a simulated rectangular-car experiment with perfect active-speaker detection, non-overlapping driver and passenger speech, and five noise types. Results in Tables II and III report SNR gain, SIR gain, and DNSMOS for the sum of MWF outputs versus the raw microphone sum, and Figs. 4-7 address notch removal and head-movement adaptation. The paper concludes that MWF gives significant improvements over simple mixing.","tokens_in":9113,"tokens_out":8075,"duration_ms":99450,"significance":"If the reported results withstand scrutiny, the paper provides a useful engineering demonstration that a standard MWF can be inserted before microphone-signal combining in an in-car two-talker setup, reducing notch distortion and background noise while allowing the two extracted signals to be added without echo-path switching. The comparison against a simple-sum baseline is appropriate, the MWF formulation is standard and self-consistent, and the authors are transparent about the perfect active-speaker-detection assumption. The DNSMOS evaluation and the head-movement adaptation study are useful additions. However, the SIR metric as defined is not computable under the non-overlapping-speech protocol, and the hyperparameters are selected on the same data used for evaluation; both issues currently prevent the quantitative claims from being taken at face value. The single simulated room and the absence of overlapping speech also limit generalization.","major_comments":[{"comment":"The reported finite SIR gains, e.g., 8.28 dB for the driver with white noise in Table II, are not consistent with the stated definition and protocol. Because the driver and passenger speech segments do not overlap, in any driver-only segment the passenger source is silent and the input SIR is infinite; the same holds for passenger-only segments. With the manuscript's statement that 'SIR and SNR gain for driver and passenger speech are calculated from their respective time segments,' a finite output SIR cannot produce the finite SIR gains shown in Tables II and III. If the authors instead computed SIR over the full recording, the input SIR would be close to 0 dB because the two non-overlapping equal-power speech segments each act as interference for the other, which contradicts the 'respective time segments' wording. Since SIR is the only metric that directly quantifies source decoupling, this ambiguity undermines the central quantitative evidence for the claimed interference-mitigation benefit. Please redefine the SIR gain with a well-specified reference (e.g., the full recording, or a defined floor for absent competing sources) and recompute Tables II and III, or add an overlapping-speech experiment in which the input SIR is finite by construction.","section":"III, 'Evaluation metrics'; Tables II and III"},{"comment":"The frame size (8 ms), the regularization factor delta, and the forgetting factor lambda were selected as the values that 'provided the best noise reduction' on the same recordings used for Tables II and III, and a hoth-noise-specific low-frequency regularization was also chosen. This is test-set hyperparameter fitting, so the reported SNR, SIR, and DNSMOS gains over simple mixing are optimistic and the sensitivity of the conclusions to these settings is unknown. Please add a development/test split, or alternatively report the metric values over the parameter grid and show that the qualitative improvement over simple microphone summing holds for a range of reasonable settings.","section":"IV.B"},{"comment":"The paper claims in Sections II and VI that the MWF can extract driver and passenger speech whether only one talker is active or both are talking simultaneously, and that it can be used for interference cancellation in the simultaneous case. However, Section III explicitly states that 'there is no overlap between speech segments in the experiments,' and no experiment in Section IV exercises two active speakers. The non-overlapping protocol therefore provides no evidence for the simultaneous-speaker or interference-cancellation claims. Please either temper these claims to the single-talker-plus-noise conditions actually tested, or add an overlapping-speech condition with different utterances for the two sources and report SNR gain, SIR gain, and DNSMOS for that condition.","section":"II, III, and VI"}],"minor_comments":[{"comment":"The second element of the cross-correlation vector p(omega) appears to be printed as Phi_{x1}(omega), but from the right-hand side and from Eq. (12) it should be Phi_{x2 d}(omega); please correct this typo.","section":"II, Eq. (9)"},{"comment":"The notation relating w, w^H, and the conjugated filter coefficients in Eqs. (5)-(7) is inconsistent and should be normalized so that the filter vector used in the inner product is unambiguously defined.","section":"II, Eqs. (5)-(7)"},{"comment":"The text states that with white noise and 5 dB input SNR the driver's SNR gain decreases from 7.68 to 7.47 dB, but Table II reports a driver SNR gain of 10.14 dB for the same nominal condition; please clarify which parameter settings or averaging windows produce each number.","section":"IV.C"},{"comment":"The time-axis labels in Fig. 2 appear garbled (e.g., '0 1 02 03 04 05 06 07 0') and are inconsistent with the 68-second intervals described in Section IV.C; please regenerate the figure with correct axis formatting.","section":"Fig. 2"},{"comment":"There is a grammatical typo in 'Test were also performed with adaptation paused after 36 seconds'; it should be 'Tests were also performed.'","section":"IV.C"},{"comment":"The sentence defining SIR says it 'assesses the decoupling of the speech and driver sources,' but it should refer to the driver and passenger sources; please fix the wording.","section":"III, 'Evaluation metrics'"}],"recommendation":"major_revision","confidential_remarks":"Editor: The paper is a straightforward engineering evaluation of a standard algorithm and is potentially suitable for an applied signal-processing venue. The main gatekeeping issues are the internally inconsistent SIR definition under the non-overlapping-speech protocol and the test-set hyperparameter selection, both of which affect the headline quantitative claims. I would ask the authors to recompute or redefine SIR, add a development/test separation or parameter sensitivity analysis, and preferably include at least one overlapping-speech experiment before reconsidering the paper. I do not see a novelty or citation-practice concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate, modest engineering contribution. The idea is to use two MWFs to extract driver and passenger speech separately before summing, avoiding the notch filtering that occurs when raw microphone signals are mixed and reducing background noise. The math is standard, and the notch-mitigation plots in Fig. 4 do show the effect. The DNSMOS gains over simple mixing are plausible. But there is an internal metric problem that needs fixing before the numbers can be trusted, and the experimental setup has idealizations that deserve fuller acknowledgment.\n\nThe big issue is the SIR gain definition. The paper defines SIR as ratio of target speech power to competing speech power, and says gains are calculated from the respective time segments. Since driver and passenger never speak at the same time in the experiments, a driver-only segment has infinite input SIR (passenger silent), and the same holds for passenger-only segments. Finite gains like 8.28 dB cannot come out of that formula. If instead the SIR were computed over the full recording, the input SIR would be roughly 0 dB because the two non-overlapping segments are equal length, contradicting the stated \"respective time segments.\" Either way, the reported SIR numbers don't follow from the text, and the stress-test note is right. This matters because SIR is the only metric that directly targets the paper's claimed contribution: separating the two sources and removing cross-path distortion.\n\nOther weaknesses are the ones the reader flagged. Hyperparameters (frame size, delta, lambda) were selected on the same data used for evaluation. The perfect-VAD assumption is explicitly stated, which is honest, but it means real-device behavior will be worse. Overlapping speech is never tested even though the conclusion claims simultaneous-talker capability. There is no code or data, and only one simulated room.\n\nCredit where it is due: the paper is transparent about the VAD assumption, the head-movement experiment with adaptation paused versus continuous is a useful stress test, and the application itself is new relative to the cited literature.\n\nNet: this is not a theoretical advance and the evidence is not conclusive, but the idea is sound enough to warrant proper review. A serious referee could get the SIR metric clarified and push for a validation split, overlapping-speech tests, and ideally real recordings. I would accept it for peer review with a strong request for revision. I would not cite it in my own work until the SIR issue is resolved, and I would not bring it to a general reading group.","headline":"A standard MWF application with a genuinely new use case, but the SIR metric that anchors its central claim is ill-defined under the paper's own non-overlapping speech protocol.","tokens_in":9613,"tokens_out":2322,"would_cite":false,"duration_ms":27432,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Extracting each talker with a multichannel Wiener filter before summing removes echo notches and cuts background noise.","keywords":["multichannel Wiener filter","in-car speech enhancement","notch filtering effect","noise reduction","DNSMOS","hands-free telephony","source extraction","microphone array"],"falsifier":"Run the identical two-microphone setup with measured car impulse responses, replace the perfect classifier with a standard voice-activity detector, and include overlapping driver-and-passenger speech segments; the central claim is settled by checking whether SNR, SIR, and DNSMOS gains still exceed the simple microphone sum, which loses about 1 dB in SNR.","tokens_in":8674,"feed_emoji":"🎙️","tokens_out":6497,"duration_ms":58796,"temperature":0.7,"pith_summary":"This paper proposes using a multichannel Wiener filter (MWF) on a two-microphone in-car setup to extract driver and passenger speech separately, then sum the extracted sources. It aims to show that this beats the common practice of simply adding the two microphone signals, which creates a comb-like notch filtering distortion from the two delayed acoustic paths and doubles background noise. The paper claims the MWF approach removes most of these spectral notches, especially with larger processing frame sizes, and improves signal-to-noise ratio, signal-to-interference ratio, and perceptual DNSMOS scores across white, red, pink, green, and hoth noise. It also claims the filter adapts quickly enough to tolerate head movements, provided adaptation is never paused.","feed_headline":"Wiener filter beats raw mic mixing for in-car voice","feed_subtitle":"Extract driver and passenger speech separately, then add: simulations show notch distortion drops and noise scores rise.","key_machinery":"The machinery is the frequency-domain multichannel Wiener filter: for each target source (driver at microphone 1, passenger at microphone 2), a filter $\\mathbf{w}(\\omega)=\\mathbf{R}^{-1}(\\omega)\\mathbf{p}(\\omega)$ is computed from the multichannel correlation matrix $\\mathbf{R}$ of the microphone signals and the cross-correlation vector $\\mathbf{p}$ with the target source, with Eqs. (11) and (13) extracting the speech-only correlations from periods when only one source is active plus noise-only statistics. Regularization $\\delta\\,\\mathrm{tr}(\\mathbf{R})/\\mathrm{size}(\\mathbf{R})\\,\\mathbf{I}$ keeps the matrix inverse stable, and a moving-average forgetting factor $\\lambda$ lets the filter track changing source positions. The two extracted sources are then simply added, which is what removes the notch filtering and reduces background noise.","core_discovery":"The central assertion is that an adaptive multichannel Wiener filter, applied before any mixing, can extract each talker's contribution at its own dedicated microphone and suppress the cross-talk path, so the two extracted signals can be added without the comb-filter notches that appear when raw microphone signals are summed. In the tested simulated car cabin, MWF gives SNR gains of roughly 10 dB for white noise and 6-9 dB for red, pink, green, and hoth noises, while raw microphone summing loses about 1 dB; DNSMOS speech and noise scores are consistently higher with MWF. The paper further states that the filter works with one or two speakers active and that speaker head movements of 0.1-0.15 m cause almost no performance drop as long as adaptation continues every 8 ms.","pith_inferences":["If real voice-activity detection replaces the paper's perfect-classification assumption, correlation estimates will be biased and the reported SNR, SIR, and DNSMOS gains are likely to shrink; the perfect-detection setup thus marks an upper bound for practical systems.","The simulated cabin uses symmetric impulse responses with 2 dB cross-side attenuation; real cabins with stronger or asymmetric cross-talk could show either larger benefits from MWF or require different frame sizes and regularization.","The paper claims MWF handles simultaneous talkers but only tests non-overlapping speech segments; a direct overlapping-speech experiment would settle whether the two filters truly separate talkers or merely noise-reduce the mixture.","With more than two microphones, the same correlation-estimation framework could extract additional sources or improve robustness, an extension the paper mentions but does not test."],"forward_implications":["Hands-free telephony and voice-command systems can place MWF directly after microphone pickup, before acoustic echo cancellation, with no direction-of-arrival or propagation model.","Because the extracted sources are combined before transmission, the downstream echo path stays stable, avoiding the discontinuities caused by microphone switching.","The same filtering handles single-talker and simultaneous-talker scenarios, and in the simultaneous case it can act as interference cancellation between driver and passenger.","Larger frame sizes (up to 100 ms) remove nearly all spectral notches for stationary sources, while speech needs smaller frames (8 ms) to track changing statistics; both regimes are demonstrated.","When loudspeaker or media signals are present, MWF must be disabled or paired with per-microphone acoustic echo cancellation, since the number of extractable sources is capped by the number of microphones."],"supporting_citations":[{"why":"Prior MWF in-car speech enhancement for a single talker, which this work extends to two talkers.","marker":"[1]"},{"why":"An in-car beamforming approach whose assumptions MWF avoids, since no direction-of-arrival is needed.","marker":"[2]"},{"why":"Speech-distortion-weighted MWF for diffuse noise, providing a baseline for noise-reduction behavior.","marker":"[3]"},{"why":"Establishes the equivalence between MWF and beamforming, justifying the no-propagation-model formulation.","marker":"[6]"},{"why":"Generates the simulated car room impulse responses used in all experiments.","marker":"[8]"},{"why":"Image method that the RIR generator is based on, supplying the simulated acoustic paths.","marker":"[9]"},{"why":"Standard hoth noise spectrum used as one of the tested background noise conditions.","marker":"[10]"},{"why":"Provides the DNSMOS non-intrusive perceptual scores used to evaluate output quality.","marker":"[11]"}],"fun_headline_variants":["MWF reduces in-car echo notches, improves clarity","Multichannel Wiener filter lifts in-car speech in noise","Adaptive MWF beats mic summing for driver and passenger","In-car voice: MWF improves noise scores and cuts notches","Wiener filter wins over raw mic mix for in-car speech"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise, stated in Section III, is that a perfect classifier always knows which speaker is active, and the test signals never overlap; if voice-activity detection is imperfect or driver and passenger speak simultaneously, the correlation matrices feeding the MWF become biased and the reported gains may not hold.","fun_headline_variants_meta":{"raw":{"variants":["MWF reduces in-car echo notches, improves clarity","Multichannel Wiener filter lifts in-car speech in noise","Adaptive MWF beats mic summing for driver and passenger","In-car voice: MWF improves noise scores and cuts notches","Wiener filter wins over raw mic mix for in-car speech"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000806,"raw_usage":{"total_tokens":3478,"prompt_tokens":823,"completion_tokens":2655,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":2571}},"tokens_in":439,"tokens_out":2655,"duration_ms":19856,"temperature":1.0,"reasoning_tokens":2571,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:37:48.633008+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical two-microphone setup with measured car impulse responses, replace the perfect classifier with a standard voice-activity detector, and include overlapping driver-and-passenger speech segments; the central claim is settled by checking whether SNR, SIR, and DNSMOS gains still exceed the simple microphone sum, which loses about 1 dB in SNR.","supporting_citations":[{"cited_title":"Multi-channel speech enhancement in a car environment using Wiener filtering and spectral subtraction,","cited_arxiv_id":null,"evidence_quote":"Prior MWF in-car speech enhancement for a single talker, which this work extends to two talkers."},{"cited_title":"A subband hybrid beamforming for in-car speech e nhancement,","cited_arxiv_id":null,"evidence_quote":"An in-car beamforming approach whose assumptions MWF avoids, since no direction-of-arrival is needed."},{"cited_title":"On the Speech Distortion Weighted Multichannel Wiener Filter for diffuse Noise,","cited_arxiv_id":null,"evidence_quote":"Speech-distortion-weighted MWF for diffuse noise, providing a baseline for noise-reduction behavior."},{"cited_title":"Adaptive Beamforming and Postfiltering,","cited_arxiv_id":null,"evidence_quote":"Establishes the equivalence between MWF and beamforming, justifying the no-propagation-model formulation."},{"cited_title":"Available: https://github.com/ehabets/RIR- Generator","cited_arxiv_id":null,"evidence_quote":"Generates the simulated car room impulse responses used in all experiments."},{"cited_title":"Image method for efficiently simulating small-room acoustics,","cited_arxiv_id":null,"evidence_quote":"Image method that the RIR generator is based on, supplying the simulated acoustic paths."},{"cited_title":"Room noise spectra at subscribers' telephone locations,","cited_arxiv_id":null,"evidence_quote":"Standard hoth noise spectrum used as one of the tested background noise conditions."}],"review_version":1}