{"id":"4cb72ca3-bd63-49cc-b1c1-ff6cc70720bc","arxiv_id":"2507.05399","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A two-channel Kalman echo canceller with coherence-drift SRO estimation and resampling matches no-SRO performance for uncorrelated playback in simulated two-device tests.","lead":"Two wireless devices playing speech at slightly different clock speeds can break acoustic echo cancellation. A system that estimates the clock offset, resamples the far-end signal, and runs a two-channel Kalman filter restores echo cancellation in simulations, though identical playback still leaves a gap.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SRO estimator's reliance on an energy-based ideal VAD is the load-bearing assumption; without it, the double-talk results in Figs. 4-5 may not hold.","rationale":"The reader's conditional verdict already flags the ideal VAD as a limitation, but their weakest_assumption is the primary device's access to X2 before transmission. I consider the ideal VAD more directly load-bearing: the abstract explicitly claims performance during double-talk, and the SRO estimator uses an oracle VAD to gate the coherence computation. Without that oracle, near-end speech would contaminate the frames selected for SRO estimation, likely degrading the GCC-based estimate and slowing convergence. The paper gives no evidence that the estimator is robust to VAD errors or practical VAD replacement. The reference-access assumption is weaker because the described host-client architecture naturally places the primary device as the source of X2; it is a scoping condition, not an internal flaw. My concern does not invalidate the paper's proof-of-concept under ideal VAD; it strengthens the case for the reader's CONDITIONAL verdict, so no verdict adjustment is needed.","tokens_in":8221,"tokens_out":13576,"duration_ms":161928,"concrete_test":"Re-run the double-talk experiments in Figs. 4 and 5 with the ideal VAD replaced by a practical VAD (e.g., a simple far-end-only energy VAD on X2, or a standard single-channel VAD on the microphone), keeping all AEC settings identical. Compare Variant 2's PESQ for SRO = +100 ppm against the reported ideal-VAD result and the oracle-SRO baseline. If the PESQ drops by more than 0.3, or the SRO estimate converges more than 2x slower, the double-talk claim depends on the oracle VAD.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the SRO estimator in Sec. 3.1, which computes the coherence function Γ(k,m) only when the 'energy-based ideal VAD' detects speech in both the reference X2 and the input I(k,m). This VAD is an oracle: it knows true speech activity in the near-end/microphone signal. The paper's headline claim covers double-talk, but in double-talk the estimator is exactly evaluated under this oracle condition. If a practical VAD were used, frames containing near-end speech but weak X2 echo would still be selected (or valid X2-echo frames discarded), corrupting the GCC phase in Eq. (9) and slowing SRO convergence. The paper provides no ablation or sensitivity analysis for VAD errors; Fig. 5 already shows Variant 2 below oracle in double-talk, so the oracle VAD is not a benign simplification. Thus the central claim 'mitigates divergence ... during double-talk' is not established for non-oracle conditions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses acoustic echo cancellation in a two-device scenario where an auxiliary loudspeaker runs on an independent clock, introducing sample rate offset (SRO) between the far-end reference and the primary device's microphone. The authors propose a synchronous compensation scheme: estimate the SRO using the DWACD algorithm and resample the auxiliary signal before feeding it to a two-channel frequency-domain Kalman AEC. Two variants are presented: Variant 1 uses the two-channel AEC error for SRO estimation, while Variant 2 uses the error of an independent single-channel AEC to decouple SRO estimation from filter convergence. Experiments with 50 simulated two-device settings compare these variants against no-SRO, no-compensation, and oracle-SRO baselines for both uncorrelated and correlated playback, in echo-only and double-talk scenarios. The results show that for uncorrelated playback both variants reach the no-SRO baseline, whereas for correlated playback Variant 2 matches the oracle in echo-only but leaves a performance gap in double-talk.","tokens_in":8434,"tokens_out":5014,"duration_ms":53030,"significance":"If the results are reproducible, the paper offers a practical step toward handling clock drift between consumer devices in spatial teleconferencing, and the Variant 2 decoupling idea is a useful design insight. The authors are honest about the correlated double-talk gap and the limitations at higher SRO values, and they link to an online resource for reproducibility. The treatment is empirical rather than analytical, but the experimental design is reasonable for the scope. The main qualification is that the SRO estimator relies on an energy-based ideal VAD, which makes the double-talk claims weaker than they first appear for real systems.","major_comments":[{"comment":"The SRO estimator is gated by an 'energy-based ideal VAD' that knows the ground-truth speech activity in I(k,m). In the double-talk experiments this oracle makes the estimator immune to near-end speech interference, because frames containing near-end speech are excluded from the coherence updates. The abstract claims the system 'mitigates the divergence ... during ... double-talk', but Fig. 5 shows that Variant 2 already falls short of the oracle-SRO baseline in double-talk for correlated playback; with a practical VAD, misclassified speech-active frames would corrupt the GCC phase in Eq. (9) and likely enlarge the gap. Please add an ablation that replaces the ideal VAD with a realistic speech-activity detector (or with a VAD at controlled error rates) and report whether the double-talk results are preserved. Without this, the double-talk claim is not established for non-oracle conditions.","section":"Sec. 3.1, Eqs. (5)-(9)"},{"comment":"All ERLE and PESQ values are reported as averages over 50 test files with no measure of variance. The claim in Sec. 4.2 that Variant 2 and the oracle have 'almost identical performance' in the echo-only correlated case is a comparison of two point estimates; without error bars or paired significance testing, the reader cannot judge whether the small differences around ±75-150 ppm are meaningful. Report per-SRO standard deviations or confidence intervals, or provide a statistical test for the key comparisons.","section":"Sec. 4.2, Figs. 4-5"}],"minor_comments":[{"comment":"The passage 'we use ˜X2 (k, m) instead X2 (k, m) of as the reference signal' is ungrammatical; rewrite as 'we use ~X2(k,m), instead of X2(k,m), as the reference signal.'","section":"Sec. 3.2, sentence after Eq. (12)"},{"comment":"The sentence 'the filter bH1,1 and bH1,2 do not converge to the true AIR H1,1 and H2,2 respectively' contains an index typo: H2,2 should be H1,2.","section":"Sec. 4.2, correlated playback paragraph"},{"comment":"The statement 'For the estimation of the SRO, two previous segments, each of length 8192 samples corresponding to 0.512 s, are used' is unclear: are the two segments overlapping or disjoint, and is the SRO estimate updated every 8192 samples? Please clarify the update schedule.","section":"Sec. 4.2, SRO estimation description"},{"comment":"Reference [19] is cited for the WSJ0 corpus but is actually a speech-separation paper; please cite the WSJ0 corpus directly (e.g., Garofolo et al., 1993).","section":"References"},{"comment":"The condition m Nh epsilon_1,2 / f1 << Nw is stated but never checked for the experimental parameters (e.g., SRO ±150 ppm, hop 256, 36 s signals). Please confirm that the condition is satisfied for the farthest evaluated frame.","section":"Eq. (4), condition following it"}],"recommendation":"major_revision","confidential_remarks":"The oracle-VAD concern is the principal reason for major revision; it is load-bearing for the double-talk claim. If the authors provide a practical VAD ablation and show that the performance gap remains acceptable, I would be willing to accept. The manuscript also appears compact for a journal submission; the experimental methodology would benefit from more detail."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is an integration of two known components—DWACD SRO estimation and a partitioned-block Kalman AEC—into a two-device echo cancellation setting. That integration is the contribution, and it is a legitimate one. The authors show that for uncorrelated playback the SRO-compensated two-channel Kalman filter recovers the no-SRO baseline, and they identify a genuinely useful empirical fact: when the two devices play correlated content, the SRO estimator should be fed from an independent single-channel AEC error rather than from the multichannel error. That decoupling is the paper's real finding, and the SRO-versus-time plots make the point convincingly.\n\nThe paper is also honest. The abstract's claim about double-talk is stronger than what the body reports: for correlated playback in double-talk, Variant 2 still leaves a gap relative to the oracle, and the authors say so. That kind of reporting makes the rest of the results easier to trust.\n\nSoft spots, in order of importance. First, the SRO estimator relies on an energy-based ideal VAD. The coherence function in Eq. (5) is computed only when speech is detected in both the reference and the input. The VAD knows ground truth. In echo-only that is a non-issue, but the headline claim covers double-talk, and there the oracle VAD is doing real work: it is selecting exactly the frames where the echo-to-near-end ratio is usable. The paper gives no sensitivity analysis or ablation with a realistic VAD. The stress-test note is right that this is the load-bearing assumption; it doesn't sink the paper, but it caps the strength of the double-talk claim.\n\nSecond, all results are averages over 50 simulated files with no error bars, and the SRO is constant-valued per file. Those are presentation choices rather than errors, but they make it harder to judge how robust the advantage of Variant 2 is. Third, the assumption that the primary device has the auxiliary reference before transmission is stated clearly, but it does restrict practical deployment to setups with that coordination.\n\nThe math in the signal model and the Kalman-filter update is standard and consistent. The citation pattern is appropriate: the DWACD and Kalman sources are the right prior art, and the paper positions itself correctly against the single-device SRO-AEC literature.\n\nWho this is for: AEC researchers and engineers working on spatial teleconferencing with multiple Bluetooth or WiFi speakers. A serious referee should spend time on it, mainly to push for a non-oracle VAD ablation and error bars. I would accept it for peer review.","headline":"A solid empirical integration paper whose real contribution is the decoupling finding for correlated playback; the ideal-VAD dependence and lack of error bars are the soft spots, not the math.","tokens_in":8982,"tokens_out":2629,"would_cite":true,"duration_ms":31712,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sample rate offsets between devices no longer have to break acoustic echo cancellation.","keywords":["acoustic echo cancellation","sample rate offset","Kalman filtering","multi-device teleconferencing","coherence drift","resampling","clock drift","double-talk"],"falsifier":"Use the same setup but replace the clean pre-transmission $X_2$ with the delayed or mixed signal a primary device could actually observe after transmission; if ERLE and PESQ no longer recover to the no-SRO baseline, the pre-transmission-access assumption, not the SRO-estimation and resampling method, is what carries the result.","tokens_in":8014,"feed_emoji":"🎧","tokens_out":6839,"duration_ms":73090,"temperature":0.7,"pith_summary":"This paper claims that acoustic echo cancellation across multiple loudspeaker devices can be made to work despite sample rate offsets by treating the setup as a two-channel filtering problem, estimating the offset between device clocks, and resampling the auxiliary far-end reference to match the primary device's clock. In a simulated two-device room with uncorrelated playback, the proposed system recovers the same echo return loss enhancement and speech quality as a system with no offset, and it matches an oracle that knows the true offset. With correlated playback, the paper shows that an independent single-channel AEC error must drive the SRO estimator, otherwise the estimator and multi-channel filter interfere and performance stays near the uncompensated baseline. This matters because multi-device teleconferencing over Bluetooth or WiFi currently loses echo cancellation quality to clock drift, and a synchronous reference-based fix would avoid changing the devices or the network.","feed_headline":"Clock drift no longer breaks multi-device echo cancellation","feed_subtitle":"Resampling the auxiliary reference lets a two-channel Kalman filter match no-offset performance.","key_machinery":"The load-bearing object is the frequency-domain phase-rotation model $\\Lambda(k,m)$, which turns a drifting auxiliary signal into a stationary reference at the primary device's clock. The DWACD algorithm estimates $\\epsilon_{1,2}$ from the phase slope of the complex coherence between the auxiliary reference and a chosen error signal, using golden-search refinement for the fractional time lag; the estimated rotation then resamples $X_2$ to produce $\\tilde X_2$ for the two-channel Kalman filter. The second key piece is Variant 2's independent single-channel AEC, whose error $E_0$ removes the primary device's own echo before coherence estimation, decoupling SRO estimation from the slow-converging multi-channel filter.","core_discovery":"The central claim is that the divergence of a multi-channel Kalman AEC filter under sample rate offset can be mitigated by estimating the offset with the dynamic weighted average coherence drift algorithm and resampling the auxiliary far-end signal before filtering. Concretely, the SRO appears in the frequency domain as the phase rotation $\\Lambda(k,m) = e^{-j2\\pi k/N_w \\cdot mN_h \\epsilon_{1,2}/f_1}$ on $X_2(k,m)$, so resampling means multiplying the auxiliary reference by this rotation and using $\\tilde X_2 = X_2\\Lambda$ as the second channel's input to the partitioned-block Kalman filter. In two-device experiments, this restores no-offset ERLE and PESQ for uncorrelated playback, and for correlated playback the decoupled Variant 2 matches oracle performance in echo-only; a residual gap in double-talk is attributed to less robust SRO estimation when near-end speech is present.","pith_inferences":["The same phase-rotation and resampling recipe should generalize to $Q$ devices by estimating each auxiliary SRO independently, but the paper only validates the two-device case, so behavior across many devices is untested.","The energy-based ideal VAD used for coherence gating means real-world noise and double-talk may weaken SRO estimates; a robust voice activity detector or a non-speech reference would be a natural next test.","A fully blind setup, where the primary device only hears the auxiliary loudspeaker through the microphone instead of receiving a clean reference, escapes the paper's access assumption and is the main scenario the method cannot yet serve.","If SRO estimation accuracy is the double-talk bottleneck, improving the coherence estimator, for example by longer temporal smoothing or a multi-tap lag search, could close the correlated-playback gap without changing the two-channel Kalman structure."],"forward_implications":["For uncorrelated playback, SRO compensation makes the two-device Kalman AEC reach the no-SRO baseline for offsets within at least ±75 ppm and keeps most of that gain up to ±150 ppm.","For correlated playback, the SRO estimator must be fed by an independent single-channel AEC; without that decoupling, the proposed system performs no better than doing no compensation.","SRO estimates remain stable across echo path changes in echo-only conditions, so moving a device or switching a microphone does not force re-convergence of the offset estimate.","The resampling-based compensation does not add algorithmic delay, since the SRO estimate uses two previous 0.512-second segments.","In double-talk with correlated playback, even the oracle-SRO system leaves a performance gap, indicating that SRO estimation robustness, not the Kalman structure alone, is the remaining bottleneck."],"supporting_citations":[{"why":"Supplies the partitioned-block frequency-domain Kalman filter architecture that the two-channel AEC is built on.","marker":"[4]"},{"why":"Provides the dynamic weighted average coherence drift algorithm used for SRO estimation and buffer management.","marker":"[13]"},{"why":"Supplies the Taylor-series/frequency-domain approximation that turns the SRO into the phase-rotation term $\\Lambda(k,m)$.","marker":"[16]"},{"why":"Gives the validity condition $mN_h\\epsilon_{1,2}/f_1 \\ll N_w$ under which the phase-rotation model holds.","marker":"[17]"},{"why":"Source of the ±150 ppm SRO range used to simulate clock offsets between devices.","marker":"[21]"},{"why":"Provides the overlap-save STFT method used to simulate sample-rate-offset echo signals in the experiments.","marker":"[22]"},{"why":"Earlier synchronous method that models SRO as a time-scaling parameter and synchronizes before AEC; the paper extends this idea from single-device to two-device AEC.","marker":"[9]"},{"why":"Related clock-skew-robust AEC baseline that the paper contrasts with the explicit SRO-estimation-and-resampling approach.","marker":"[12]"}],"fun_headline_variants":["Resampling tames clock drift in multi-device echo cancellation","Sample-rate offset fix for multi-device echo cancellation","Multi-device echo cancellation without clock-drift divergence","Resampling the far-end signal keeps Kalman AEC on track","SRO estimation rescues multichannel echo cancellation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The primary device must have access to the auxiliary far-end signal $X_2$ before it is transmitted, so the SRO estimator and resampler always see a clean reference; if only a delayed, mixed, or post-transmission copy is available, the synchronous compensation chain cannot be set up.","fun_headline_variants_meta":{"raw":{"variants":["Resampling tames clock drift in multi-device echo cancellation","Sample-rate offset fix for multi-device echo cancellation","Multi-device echo cancellation without clock-drift divergence","Resampling the far-end signal keeps Kalman AEC on track","SRO estimation rescues multichannel echo cancellation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000864,"raw_usage":{"total_tokens":3707,"prompt_tokens":868,"completion_tokens":2839,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":2758}},"tokens_in":484,"tokens_out":2839,"duration_ms":23168,"temperature":1.0,"reasoning_tokens":2758,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:28:54.668774+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the same setup but replace the clean pre-transmission $X_2$ with the delayed or mixed signal a primary device could actually observe after transmission; if ERLE and PESQ no longer recover to the no-SRO baseline, the pre-transmission-access assumption, not the SRO-estimation and resampling method, is what carries the result.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the partitioned-block frequency-domain Kalman filter architecture that the two-channel AEC is built on."},{"cited_title":"On Deal- ing with Sampling Rate Mismatches in Blind Source Separa- tion and Acoustic Echo Cancellation,","cited_arxiv_id":null,"evidence_quote":"Provides the dynamic weighted average coherence drift algorithm used for SRO estimation and buffer management."},{"cited_title":"Asynchronous Acoustic Echo Cancellation Over Wireless Channels,","cited_arxiv_id":null,"evidence_quote":"Supplies the Taylor-series/frequency-domain approximation that turns the SRO into the phase-rotation term $\\Lambda(k,m)$."},{"cited_title":"Clock Skew Robust Acoustic Echo Cancella- tion,","cited_arxiv_id":null,"evidence_quote":"Gives the validity condition $mN_h\\epsilon_{1,2}/f_1 \\ll N_w$ under which the phase-rotation model holds."},{"cited_title":"Blind sam- pling rate offset estimation and compensation in wireless acoustic sensor networks with application to beamforming,","cited_arxiv_id":null,"evidence_quote":"Source of the ±150 ppm SRO range used to simulate clock offsets between devices."},{"cited_title":"Correlation maximization-based sam- pling rate offset estimation for distributed microphone arrays,","cited_arxiv_id":null,"evidence_quote":"Provides the overlap-save STFT method used to simulate sample-rate-offset echo signals in the experiments."},{"cited_title":"State-space architec- ture of the partitioned-block-based acoustic echo controller,","cited_arxiv_id":null,"evidence_quote":"Earlier synchronous method that models SRO as a time-scaling parameter and synchronizes before AEC; the paper extends this idea from single-device to two-device AEC."},{"cited_title":"An Analysis of Time Drift in Hand-Held Recording Devices,","cited_arxiv_id":null,"evidence_quote":"Related clock-skew-robust AEC baseline that the paper contrasts with the explicit SRO-estimation-and-resampling approach."}],"review_version":1}