{"id":"bc052658-fee4-413a-ab94-4129013cd673","arxiv_id":"2507.18350","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A dual-path time-frequency MCLP filter cascade plus a multi-norm (l2 and l1) beamformer improves speech enhancement metrics over classical baselines in simulated reverberant rooms, especially at high T60.","lead":"This paper proposes a speech enhancement method that combines dual-path multi-channel linear prediction filters with a beamformer that suppresses both the power and the l1 norm of the output. It reports improved PESQ and SI-SNR over classical baselines in simulated reverberant environments, with gains largest at high reverberation times.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed gains may reflect test-set-tuned prediction orders: Sec. 5 fixes K_t and K_f per T60 without showing they come from the Sec. 4 Pearson rule, so the high-reverberation advantage is not yet established.","rationale":"The reader's designated weakest assumption is the complex soft-thresholding step, and that is a genuine gap: the S operator in Sec. 3.1.2 is defined for real scalars only, yet Eqs. (9) and (14) apply it to complex STFT vectors. However, this is readily repaired by using the standard complex proximal operator S_kappa(v) = (v/|v|) max(|v|-kappa, 0) from the cited reference [23], and a corrected implementation could preserve the empirical results. The order-selection issue is harder to repair post hoc: if the Fig. 2 orders were chosen from PESQ curves on the test configurations, the entire performance comparison is confounded. Section 5 states the per-T60 K_t and K_f values flatly, with no statement that Eq. (18) produced them; Section 4 gives no numerical threshold values; and Fig. 1 uses PESQ to identify the optimal orders, creating the appearance of circular selection. The evaluation also omits error bars, code, and baseline parameter settings, but the order-selection opacity directly undercuts the abstract's 'particularly in high reverberation' claim. I therefore keep the reader's CONDITIONAL verdict: the method is plausible, but the performance claim needs a blind order-selection run and equivalently tuned baselines before it can be accepted.","tokens_in":8308,"tokens_out":19397,"duration_ms":187877,"concrete_test":"Rerun the Sec. 5 T60 sweep with K_t and K_f chosen blind by the Sec. 4 Pearson-correlation rule: pre-specify the threshold delta on a development set, apply Eq. (18), and report the resulting orders alongside the PESQ and SI-SNR curves. Also run GWPE, GWPE+MVDR, and WPD with the same development-set tuning procedure for their own parameters. If the proposed method no longer dominates at T60 = 0.6-1.0 s, the central claim is not supported; if it still dominates, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim (abstract; Fig. 2) is that the dual-path MCLP plus multi-norm beamformer outperforms GWPE, GWPE+MVDR, and WPD, especially at long T60. For this to hold, the comparison must use a valid, reproducible procedure for setting K_t and K_f. Sec. 4 proposes Eq. (18), K_t = (K_{delta1}+K_{delta2})/2, based on Pearson correlation thresholds, but the thresholds delta1 and delta2 are never specified and the method is never applied in Sec. 5. Instead, the experiments simply list K_t={10,14,18,22,24} and K_f={2,4,6,8,10} for T60={0.2,...,1.0}s with no statement that Eq. (18) produced these values. Because Fig. 1 first generates PESQ-versus-K_t curves and marks the resulting orders from those curves, the per-T60 orders used in Fig. 2 may have been selected with knowledge of the test PESQ, while the baselines are not reported as being tuned in the same way. If so, the reported advantage is an artifact of favorable hyperparameter selection rather than of the proposed method itself. The optimization ambiguity in Eqs. (9) and (14) (a real-valued soft-threshold operator applied to complex vectors) is a secondary but real reproducibility issue; the order-selection opacity is the more direct threat to the headline claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a speech enhancement algorithm that combines dual-path multi-channel linear prediction (MCLP) filters, operating in both time and frequency dimensions, with a minimum-power distortionless beamformer that also penalizes the l1 norm of the output. An auxiliary contribution is a method for selecting MCLP prediction orders based on Pearson correlation thresholds. The method is evaluated in simulated reverberant environments using TIMIT speech, an 8-microphone uniform linear array, and compared against GWPE, GWPE+MVDR, and WPD using PESQ and SI-SNR. The reported results show advantages over the baselines especially at long T60 values and are claimed to support the proposed order-selection rule.","tokens_in":8674,"tokens_out":7282,"duration_ms":71183,"significance":"If the technical issues are resolved, the dual-path extension of MCLP to frequency-domain prediction is a plausible and interesting direction, and a validated data-driven order-selection rule would be practically useful for MCLP-based dereverberation methods. However, the current manuscript contains two load-bearing problems: the complex soft-thresholding update is incorrectly specified, and the order-selection rule is not demonstrably the source of the experimental settings. These must be addressed before the reported performance gains can be attributed to the proposed method.","major_comments":[{"comment":"The soft-thresholding operator S_{λ_z/μ_z}(v) is defined for real scalars via two one-sided inequalities, but it is applied to complex STFT vectors z(n,ω). The correct proximal map for the complex ℓ1 norm is element-wise magnitude shrinkage, i.e., S_τ(v) = max(|v|−τ,0) · v/|v|, which is neither stated nor implied by the displayed definition. As written, the z-update in Eq. (9) does not implement the proximal step for the ℓ1 term in Eq. (5), and the same issue affects the complex scalar update in Eq. (14). Please correct the operator definition and confirm that the implementation uses the complex shrinkage form.","section":"Sec. 3.1.2, Eq. (9)"},{"comment":"The prediction-order selection method is not connected to the experiments. Equation (18) depends on thresholds δ1 and δ2, but no numerical values are given for these thresholds. In Sec. 5, the paper lists K_t = {10,14,18,22,24} and K_f = {2,4,6,8,10} for T60 = {0.2,...,1.0} s but does not state that these values were obtained from Eq. (18). Since Fig. 1 displays PESQ curves for K_t under each T60 and marks the selected orders, it is unclear whether the orders used in Fig. 2 were chosen from the PESQ curves themselves, which would bias the comparison in favor of the proposed method and would not demonstrate the advertised order-selection contribution. Please specify δ1 and δ2, present the K_t and K_f values predicted by Eq. (18) for each T60, and compare them with the values used in the simulations.","section":"Secs. 4 and 5"},{"comment":"The paper reports mean PESQ and SI-SNR over 100 Monte Carlo runs but gives no measure of variability. Several reported differences are small; for example, at T60 = 0.2 s the proposed method is below WPD in PESQ, and at higher SNRs some gaps are within a few hundredths of a point. Without standard deviations, confidence intervals, or significance tests, the reader cannot assess whether the observed ordering of methods is statistically meaningful. Please add error bars or significance tests, and clarify whether the same noise and reverberation realizations are used for all methods in each Monte Carlo run.","section":"Sec. 5, Fig. 2"},{"comment":"Equation (17) defines the Pearson correlation coefficient across Monte Carlo realizations i between values at time indices 0 and t. This is an ensemble correlation, not the temporal autocorrelation that is conventionally used for prediction-order selection. If the intended measure is the sample autocorrelation of a single recording, the equation and the surrounding description need to be corrected. If the ensemble correlation is truly intended, its relationship to the optimal MCLP prediction order should be justified. Please clarify this point, as it is load-bearing for the order-selection method.","section":"Sec. 4, Eq. (17)"}],"minor_comments":[{"comment":"There are stray spacing issues in the abstract, e.g., 'us ing' and 'thel 1' should be 'using' and 'the l1'.","section":"Abstract"},{"comment":"The phrase 'Korder convolution' should be 'K-th order convolution'.","section":"Sec. 2.1"},{"comment":"The sentence 'The prediction order selection method is present' should be 'is presented'.","section":"Sec. 4"},{"comment":"The caption 'Pearson correlation coefficients and PESQ with different T60 values' is vague; please label the axes and panel subcaptions to make clear which quantity is plotted and which T60 applies to each panel.","section":"Fig. 1"},{"comment":"The text 'the additive noises are Gaussian white' should be 'the additive noise is white Gaussian noise'.","section":"Sec. 5"},{"comment":"The experimental setup does not specify the number of sources Q, the locations of the target source and noise sources, or whether the noise is diffuse or a point source. Please state these details, as they affect the interpretation of the beamforming results.","section":"Sec. 5"},{"comment":"The hyperparameters λ_z, λ_w, ρ_G, ρ_w, μ_z, μ_w, γ, γ_w, and γ_1 are not given values in the text. A table with the selected values and the tuning procedure would improve reproducibility.","section":"Secs. 3 and 5"},{"comment":"The claim that the proposed order-selection method 'can also be applied to other MCLP-based methods' is not supported by any experiment; please either add such an experiment or temper the claim.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely suited for a signal-processing venue, but the current version does not substantiate the headline claims. The complex soft-thresholding issue is easily fixable, but the order-selection opacity is more serious: if the orders in Sec. 5 were chosen with knowledge of the test PESQ, the comparison is not fair. I would advise the editor to require a revised version where the order-selection rule is applied prospectively and statistical variability is reported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nQuick take: this paper has a genuinely new combination and a useful heuristic, but the evaluation has a tuning problem that directly affects the headline claim.\n\nThe new bit: applying MCLP filters along both time and frequency dimensions (dual-path) and coupling that with an l1-regularized distortionless beamformer. The Pearson-correlation order selection (Eq. 18) is a fresh idea, and if it works, it would be useful for other MCLP methods. The paper is honestly written: it shows the proposed method loses to WPD at T60=0.2s, which is a good sign.\n\nWhat's well done: the problem formulations are standard, the optimization via PALM and ADMM is clearly laid out, and the experiments cover several T60 and SNR conditions. The baseline comparisons are reasonable, though a bit old.\n\nWhere it gets soft. One mathematical ambiguity: Eq. (9) and (14) apply the real soft-threshold operator S_{λ/μ} to complex STFT vectors. The operator as defined only makes sense for real scalars; for complex coefficients the correct proximal map for the l1 norm should shrink magnitudes. If the code literally follows the paper, the objective being minimized isn't Eq. (5). That needs to be fixed or explained.\n\nMore important, the order selection in Sec. 5 may not be the one in Sec. 4. The text gives K_t and K_f per T60 without showing that Eq. (18) produced them. The thresholds δ1 and δ2 are never specified, and Fig. 1 shows PESQ curves used to mark the \"optimal\" orders. So it looks like the per-T60 orders were chosen with knowledge of the test-set PESQ, while the baselines were not tuned the same way. That would make the high-reverb advantage an artifact. The stress-test note is on point.\n\nOther minor things: no error bars despite 100 MC runs, no code, only Gaussian white noise and simulated RIRs. The l1 term is justified by sparsity, but the paper doesn't analyze when that sparsity actually holds.\n\nWho is this for? People working on classical (non-DL) multichannel dereverberation and beamforming. It deserves a serious referee because the core idea is plausible and the issues are fixable: clarify the complex soft-threshold, release code, and show that the order-selection rule is actually used in the experiments.\n\nMy recommendation: send it to review, but with a strong request for a corrected optimization step and an honest description of how K_t and K_f were set.\n\nYours,\n[Name]","headline":"A plausible dual-path MCLP + l1-beamforming combination with a nice order-selection idea, but the complex soft-thresholding is ambiguous and the per-T60 order choices look tuned to the test PESQ curves.","tokens_in":9204,"tokens_out":2627,"would_cite":false,"duration_ms":26352,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that combining temporal and frequential MCLP filters with an l1-regularized beamformer outperforms standard cascades in reverberant speech enhancement.","keywords":["speech enhancement","multichannel linear prediction","dereverberation","beamforming","l1 sparsity","prediction order selection","microphone array","dual-path filtering"],"falsifier":"Re-solve Eq. (5) with a proximal update that shrinks the magnitude of each complex STFT coefficient while keeping its phase, then rerun the reported TIMIT experiments; if PESQ and SI-SNR no longer beat the baselines, the claimed gains depend on the incorrectly specified thresholding step rather than on the dual-path or multi-norm idea.","tokens_in":8101,"feed_emoji":"🎙️","tokens_out":6801,"duration_ms":67604,"temperature":0.7,"pith_summary":"The paper sets out to show that microphone-array speech enhancement in noisy, reverberant rooms improves when late reflections are predicted along both the time axis and the frequency axis of the STFT domain, rather than time alone, and when the subsequent beamformer minimizes both output power and the $\\ell^1$ norm of its output. It combines a dual-path multichannel linear prediction (MCLP) filter with a multi-norm beamformer and evaluates the pair on simulated rooms. The reported result is higher PESQ and SI-SNR than the GWPE, GWPE+MVDR, and WPD baselines, with the largest advantage at high reverberation (T60 of 0.4 to 1.0 s). The paper also offers a correlation-threshold rule for choosing prediction orders that is meant to transfer to other MCLP methods. If correct, the method is a training-free signal-processing route to better far-field speech quality, and a simple knob for tuning dereverberation filters.","feed_headline":"Dual-path prediction plus sparse beamforming wins in heavy echo","feed_subtitle":"A training-free pipeline beats GWPE and MVDR cascades on PESQ and SI-SNR as T60 rises from 0.4 to 1.0 s.","key_machinery":"The load-bearing mechanism is a pair of filter matrices: G_t, a temporal MCLP filter estimated per frequency bin, and G_f, a frequential filter estimated per time frame, which together predict late reverberation from stacked observations in both directions. The filters are found by minimizing a summed $\\ell^2$ plus $\\ell^1$ cost over dereverberated STFT coefficients via Proximal Alternating Linearized Minimization (PALM), with soft thresholding supplying the $\\ell^1$ proximal step. A second stage applies a multi-norm beamformer—output power plus $\\ell^1$ penalty under a distortionless constraint—solved by ADMM. Prediction orders K_t and K_f are set by thresholding Pearson correlation coefficients between reference-microphone samples at increasing time or frequency lags, replacing grid search. The $\\ell^1$ terms are the reason the method is called 'multi-norm': both stages mix power ($\\ell^2$) minimization with sparsity ($\\ell^1$) regularization.","core_discovery":"The central claim is that jointly estimating two MCLP filter matrices—one that operates across time frames at each frequency and one that operates across frequency bins at each time frame—removes late reverberation more completely than temporal-only prediction, and that adding an $\\ell^1$ sparsity penalty to a distortionless beamformer's $\\ell^2$ power cost improves denoising. On 8-microphone simulated arrays with TIMIT speech and image-method room responses, the paper reports that this dual-path, multi-norm system outperforms GWPE, GWPE+MVDR, and WPD on PESQ and SI-SNR for T60 from 0.4 to 1.0 s and across all tested SNRs, with the gains concentrated in heavy reverberation. The authors attribute the improvement to more comprehensive modeling of late reverberation by the frequential filter path and to the sparsity prior on the enhanced output.","pith_inferences":["The complex soft-thresholding inconsistency in Eq. (9) and Eq. (14) means the reported gains should be rechecked with a magnitude-based complex proximal operator; the l1 objective may not be what is actually minimized.","Because the experiments use Gaussian white noise and simulated RIRs, the method's advantage in real rooms with babble or diffuse noise is unverified; a test with measured RIRs and nonstationary noise would be the natural next check.","The order-selection recipe could be plugged into standard WPE or MCLP pipelines as a cheap heuristic for choosing filter lengths, which would make the paper's contribution useful even if the dual-path gains do not replicate.","The sparsity penalty on beamformer output may trade off intelligibility against perceived quality; reporting STOI or word-error-rate on a downstream recognizer would clarify whether the PESQ gains are practically useful."],"forward_implications":["At T60 values from 0.4 to 1.0 s, the proposed method reports higher PESQ and SI-SNR than GWPE, GWPE+MVDR, and WPD, and the margin grows with reverberation.","The l1 norm on beamformer output adds denoising power beyond power minimization, yielding gains over WPD across all tested SNR levels at moderate T60.","The Pearson-correlation threshold method selects temporal and frequential prediction orders without grid search and carries over to other MCLP-based systems.","At very low T60 (0.2 s) the temporal-only WPD baseline is still slightly better, implying a T60-aware switch between dual-path and temporal-only filtering would be useful."],"supporting_citations":[{"why":"Supplies the GWPE baseline and the temporal MCLP estimation framework that the proposed dual-path filters extend.","marker":"[9]"},{"why":"Provides the GWPE+MVDR cascade baseline used for comparison.","marker":"[12]"},{"why":"Provides the WPD unified convolution beamformer baseline that jointly performs denoising and dereverberation.","marker":"[14]"},{"why":"Supplies the PALM optimization scheme used to solve the dual-path MCLP cost in Eq. (5).","marker":"[21]"},{"why":"Supplies the soft-thresholding operator used for the l1 proximal update.","marker":"[23]"},{"why":"Provides the ADMM framework used to solve the multi-norm beamforming problem.","marker":"[24]"},{"why":"Supplies the TIMIT speech corpus used as source material in all experiments.","marker":"[25]"},{"why":"Supplies the image-method room impulse response generator used to create simulated test environments.","marker":"[26]"}],"fun_headline_variants":["Dual-path MCLP + sparse beamforming beat heavy reverb","Time-frequency MCLP and l1 norm beat baselines in echo","Training-free dual-path filters plus sparse beamforming win","Joint dual-path MCLP and multi-norm beamforming outperform"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the soft-thresholding step used to enforce sparsity is valid for the complex STFT coefficients it is applied to, but the paper defines the operator only for real scalars; if an implementation applies that formula directly to complex values, the algorithm no longer solves the stated $\\ell^2$+$\\ell^1$ problem.","fun_headline_variants_meta":{"raw":{"variants":["Dual-path MCLP + sparse beamforming beat heavy reverb","Time-frequency MCLP and l1 norm beat baselines in echo","Training-free dual-path filters plus sparse beamforming win","Joint dual-path MCLP and multi-norm beamforming outperform"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001389,"raw_usage":{"total_tokens":5599,"prompt_tokens":897,"completion_tokens":4702,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":4628}},"tokens_in":513,"tokens_out":4702,"duration_ms":41732,"temperature":1.0,"reasoning_tokens":4628,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:13:54.094269+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-solve Eq. (5) with a proximal update that shrinks the magnitude of each complex STFT coefficient while keeping its phase, then rerun the reported TIMIT experiments; if PESQ and SI-SNR no longer beat the baselines, the claimed gains depend on the incorrectly specified thresholding step rather than on the dual-path or multi-norm idea.","supporting_citations":[{"cited_title":"Prediction of speech intelligibility in spatial noise and reverberation for normal-hearing and hearing- impaired listeners,","cited_arxiv_id":null,"evidence_quote":"Supplies the GWPE baseline and the temporal MCLP estimation framework that the proposed dual-path filters extend."},{"cited_title":"Ephraim and I","cited_arxiv_id":null,"evidence_quote":"Provides the GWPE+MVDR cascade baseline used for comparison."},{"cited_title":"Improved subspace-based single- channel speech enhancement using generalized super-gaussian priors,","cited_arxiv_id":null,"evidence_quote":"Provides the WPD unified convolution beamformer baseline that jointly performs denoising and dereverberation."},{"cited_title":"Multichannel online speech dereverberation under noisy environments,","cited_arxiv_id":null,"evidence_quote":"Supplies the PALM optimization scheme used to solve the dual-path MCLP cost in Eq. (5)."},{"cited_title":"Dual-path rnn: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,","cited_arxiv_id":null,"evidence_quote":"Supplies the soft-thresholding operator used for the l1 proximal update."},{"cited_title":"Dpt-fsnet:dual-path trans- former based full-band and sub-band fusion network for speech enhancement,","cited_arxiv_id":null,"evidence_quote":"Provides the ADMM framework used to solve the multi-norm beamforming problem."}],"review_version":2}