{"id":"488e1a0f-1aa7-418f-9e42-ad54b47d8287","arxiv_id":"2507.05402","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A spatial-filter-based sample rate offset compensation method preserves binaural cues in stereo reproduction, shown via objective metrics and a MUSHRA listening test.","lead":"This paper proposes an audio-domain method to fix clock skew between wireless loudspeakers by estimating and compensating sample rate offsets using spatial filtering and a microphone array. It shows with simulations and a listening test that the compensation preserves binaural spatial cues that clock skew would otherwise degrade.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Oracle RTF (Eq. 19) is an unresolved component: the central validation assumes the true PSD in a single-source initialization, so the claim that the proposed method preserves binaural cues is not yet supported under realistic RTF estimation errors.","rationale":"I read the paper as a proof-of-concept for audio-domain SRO compensation in wireless stereo reproduction. The signal model (Section 2) and the compensation pipeline (Section 3) are coherent, and the objective/subjective results support the claim that, under the stated assumptions, the oracle-RTF-based compensation preserves binaural cues relative to the uncompensated case. The paper's own text flags the key simplification: 'Since the RTF estimation is not the main focus of the paper, in the current study, we use an oracle RTF' (Section 3.2, Eq. 19). This is exactly the load-bearing assumption for the central claim, because the DWACD estimator receives its input from an LCMV beamformer whose weights are oracle-dependent. A practical RTF estimate would introduce leakage between the two loudspeakers, and for correlated stereo content this leakage biases the coherence phase that drives the SRO estimate. The paper does not quantify this sensitivity. The secondary assumption that ε0 is known (Section 3.1) is also a practical limitation, but it is a stated condition that could plausibly be calibrated. I therefore agree with the reader's identification of the oracle RTF as the weakest assumption and with the CONDITIONAL verdict. The concrete test I propose replaces the oracle RTF with a finite-sample estimate and reruns the objective evaluation; this would settle whether the oracle assumption is genuinely load-bearing or whether the method degrades gracefully.","tokens_in":9423,"tokens_out":11976,"duration_ms":134539,"concrete_test":"Replace the oracle RTF in Eq. (19) by a practical estimate: compute the cross-PSD in Eq. (20) from a finite initialization segment (e.g., 10 s of single-source activity) using the same STFT parameters as the main evaluation, form Â[k,l] from the normalized cross-PSD, and rerun the full pipeline on the three SRO configurations with correlated stereo material (not identical noise in both channels). If the estimated SRO in Fig. 4 deviates by more than a few ppm from the ground truth or the residual ITD/IC difference in Fig. 3 worsens measurably at low/mid frequencies relative to the oracle-RTF condition, the oracle assumption is load-bearing; if performance is unchanged, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proposed pipeline cascades an LCMV beamformer (Eq. 18) into the DWACD SRO estimator (Eqs. 22-26). The beamformer weights are computed from the RTF matrix A[k,l]; in the experiments, A is taken as the oracle RTF of Eq. (19), formed from the true PSD E{z̄_q X_q^*} during a single-source initialization. In any real deployment this PSD is estimated from finite data, and the resulting RTF error propagates directly into the beamformer output Z_q. Imperfect interference suppression leaves a leakage term h0,r Λ_r X_r in Z_q. For correlated stereo content (as used in the MUSHRA test), this leakage biases the cross-PSD Φ_{Z_q X_q} and hence the phase of the complex coherence that DWACD converts into the SRO estimate Eq. (26). The paper neither analyzes this sensitivity nor reports how RTF estimation error degrades the ITD/IC preservation in Fig. 3 or the SRO accuracy in Fig. 4. Since the central claim is that the method mitigates SRO-induced perceptual degradation, and the method as evaluated contains an oracle component, the practical claim is conditionally supported at best.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper models the effect of sample rate offsets (SROs) in a two-loudspeaker wireless stereo reproduction system: the binaural signal is written with a phase drift term Λ_q (Eq. 3), and the microphone array at the primary device captures the same SRO-affected playback (Eq. 8). The proposed pipeline applies an LCMV beamformer (Eq. 18) using an oracle relative transfer function matrix (Eq. 19) to extract each loudspeaker contribution, estimates the SRO with the DWACD algorithm (Eqs. 22-26), and resamples the playback signal before transmission to compensate. The method is evaluated in a simulated room with three SRO configurations using objective binaural-cue difference plots (ITD and IC, Fig. 3), SRO tracking plots (Fig. 4), and a MUSHRA test with 11 listeners (Fig. 5).","tokens_in":9715,"tokens_out":8732,"duration_ms":102718,"significance":"If the results hold, this is a useful demonstration that audio-domain SRO compensation can preserve ITD and IC and reduce perceptual degradation without explicit clock synchronization, an underexplored problem in spatial audio reproduction. The paper's strengths are its clear formulation, use of a published DWACD estimator as a black box, and evaluation against ground truth, with no parameter fitting to force the outcome. However, the practical significance is currently limited by two idealized assumptions—oracle RTFs and a known primary-device SRO—and by the absence of error bars and statistical tests; the central claim is therefore defensible only as a controlled proof-of-concept.","major_comments":[{"comment":"The central validation uses an oracle RTF computed from the true PSD during a single-source initialization. In a real deployment, the RTF would be estimated from finite data, and any error propagates into the beamformer output \\hat{Z}_q, leaving interference leakage that biases the coherence phase used by DWACD (Eq. 22) and hence the SRO estimate (Eq. 26). The paper presents no sensitivity analysis and no experiment with a non-oracle RTF estimator (e.g., the methods of [21] or [22]); consequently, Fig. 3 and Fig. 5 support the compensation concept only under idealized RTFs. Please add an experiment with estimated RTFs, or at minimum a perturbation analysis of Eq. (19), and report its effect on SRO accuracy and on the ITD/IC metrics.","section":"Sec. 3.2, Eq. (19)"},{"comment":"The method requires the primary-device SRO ε0 to be known in order to recover the loudspeaker SRO ε_q from the estimated \\bar{ε}_q, but the experiments never state the value of ε0 or test robustness to its mismatch. The SRO configurations in Section 4 are described as being 'on the microphone signal', so it is unclear whether the simulated values are ε_q or \\bar{ε}_q. Please state the assumed ε0, validate the recovery ε_q = \\bar{ε}_q − ε0, and test at least one nonzero ε0 setting.","section":"Sec. 3.1, Eq. (7)"},{"comment":"The perceptual claim that the method 'significantly reduces' degradation is not supported by statistical evidence. Figure 5 shows MUSHRA results for only 11 listeners with no confidence intervals or pairwise significance tests, and Figure 4 plots SRO estimates averaged over seven files without variance. Please add error bars and statistical tests (e.g., Wilcoxon signed-rank for the MUSHRA scores and standard-deviation bands in Fig. 4), or soften the claim to a qualitative demonstration.","section":"Sec. 4, Figs. 4 and 5"}],"minor_comments":[{"comment":"In the sentence after Eq. (15), 'Ideally, \\hat{Z}_q[k,l] ≈ h_{0,q} Λ_q[k,l] X_q[k,l] + v' is dimensionally inconsistent because h_{0,q} is an M×1 vector while \\hat{Z}_q is a scalar; this should be H_{0,q,0} (the reference-microphone transfer function) to be dimensionally correct.","section":"Sec. 3.2, Eq. (15)"},{"comment":"The notation in Eqs. (19)-(21) does not consistently show the STFT frame index l; \\bar{Φ}_q and \\bar{z}_q should be written as \\bar{Φ}_q[k,l] and \\bar{z}_q[k,l] to clarify that the expectation in Eq. (20) is taken over frames or segments.","section":"Eqs. (19)-(21)"},{"comment":"The window size N_w used in the STFT analysis is never stated in the experiments; please report N_w together with the already stated hop size N_h and smoothing factors.","section":"Sec. 4"},{"comment":"The color or line-style mapping of the three rows (no compensation, oracle compensation, oracle-RTF compensation) is described only in the caption, not in the figure itself; please add labels or a legend inside the figure.","section":"Fig. 3 caption"},{"comment":"The reference list contains several OCR artifacts (e.g., 'T ime-V arying' in [11] and 'V eering' in [16]); please harmonize the bibliographic formatting.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a well-written proof-of-concept, but the oracle-RTF and known-ε0 assumptions are load-bearing, and the statistics are thin. I would support publication after a revision that adds an estimated-RTF condition, states and varies ε0, and adds error bars or statistical tests to the objective and subjective results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, honestly-scoped proof-of-concept for audio-domain SRO compensation in stereo playback. The new bit is the combination—LCMV beamforming to separate loudspeaker signals, DWACD on those outputs to estimate the SROs, resampling before rendering—plus an evaluation that actually looks at binaural cues and listener perception. That fills a genuine gap: prior SRO work in WASN and AEC focused on signal-level estimation, not on spatial reproduction.\n\nWhat I like: the pipeline is simple and reproducible. The DWACD estimator is used as a black box; the paper is not fitting parameters to make the plots look good. The self-cite to Ref. [15] is a prior AEC paper and not load-bearing. The objective metrics (ITD/IC difference plots) and the MUSHRA test are appropriate for the claim. The authors are also upfront that oracle-RTF compensation does not fully restore cues at high frequencies.\n\nThe soft spot is exactly what the stress-test says: the RTF in Eq. (19) is oracle, computed from the true PSD in an initialization phase. In any real deployment that PSD is estimated from finite data, and RTF error will leak correlated stereo content through the beamformer, biasing the coherence phase that DWACD uses. The paper neither analyzes that sensitivity nor simulates a non-oracle RTF estimator. This is a scope limitation, not an internal contradiction—the authors explicitly say RTF estimation is not the focus—but it does mean the practical claim is only conditionally supported. The assumption that ϵ0 is known is real but milder; someone has to be the reference clock.\n\nSecond issue: the evidence would be stronger with error bars or significance tests. Figure 4 averages seven files but shows no spread. The MUSHRA has 11 listeners and no statistical test, yet the text says \"significantly reduces.\" That word needs support. These are fixable in a revision.\n\nWho is this for? People working on wireless multi-speaker systems, especially consumer audio. It is a conference-grade paper with a genuine proof-of-concept and a clear limitation. I would send it to review and ask the authors to add an RTF-sensitivity experiment and proper statistics. The core demonstration holds within its stated scope.","headline":"A clean proof-of-concept for audio-domain SRO compensation in stereo reproduction, with the practical claim limited by an oracle RTF and missing error bars.","tokens_in":10218,"tokens_out":3196,"would_cite":true,"duration_ms":37066,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Clock skew between wireless speakers can be fixed by audio-domain resampling rather than network synchronization, and the fix preserves binaural cues.","keywords":["sample rate offset","clock skew","stereo reproduction","spatial audio","binaural cues","beamforming","coherence drift","wireless loudspeakers"],"falsifier":"Run the identical pipeline with an estimated RTF (for example, obtained from a single-source initialization frame with the same PSD estimator) instead of the oracle RTF of Eq. (19), in the simulated 7 m by 7 m by 6 m room with RT60 = 0.3 s and SROs (10, -100) ppm. If the estimated SRO trace deviates from ground truth by more than the smoothing tolerance, or the MUSHRA score for the compensation condition falls to within statistical noise of the uncompensated condition, the central claim that audio-domain SRO compensation preserves binaural cues fails in realistic conditions.","tokens_in":9240,"feed_emoji":"🎧","tokens_out":3788,"duration_ms":41979,"temperature":0.7,"pith_summary":"The paper establishes that sample rate offsets between wirelessly connected loudspeakers degrade stereo reproduction by drifting the phase of each loudspeaker's contribution, and that this degradation can be largely removed in the audio domain without network-level clock synchronization. The proposed system uses a microphone array at the primary device to spatially isolate each loudspeaker's signal, estimates each offset with the DWACD coherence-drift algorithm, and resamples the playback stream before transmission. Objective ITD and IC difference plots, together with a MUSHRA listening test with 11 listeners, show that the compensation preserves binaural cues and significantly reduces perceived degradation, though it does not fully eliminate it.","feed_headline":"Resampling cancels clock skew that breaks stereo cues","feed_subtitle":"A microphone array separates each speaker, estimates its sample-rate offset, and corrects playback before it reaches the listener.","key_machinery":"The central object is the SRO phase term $\\Lambda_q[k,l]$, a time- and frequency-dependent complex exponential that multiplies each loudspeaker's playback signal in the short-time Fourier domain. Compensation works by estimating the SRO $\\epsilon_q$ and resampling so that the inverse phase term cancels $\\Lambda_q$; the estimation chain is an LCMV beamformer (with diagonal loading for numerical stability) that isolates each loudspeaker's contribution, followed by the DWACD algorithm, which computes the complex coherence between beamformer output and reference signal, takes the conjugate product over a temporal distance $L$, and finds the lag maximizing the generalized cross-correlation with a golden-section refinement.","core_discovery":"Clock skew between two wireless loudspeakers appears in the binaural signal as a per-source phase term $\\Lambda_q[k,l] = \\exp\\left(-j\\frac{2\\pi k}{N_w}\\frac{l N_h \\epsilon_q}{f_s}\\right)$ that decorrelates the two channels over time, destroying interaural coherence and shifting interaural time difference. The paper shows that if each loudspeaker's contribution is first separated by an LCMV spatial filter using an oracle relative transfer function, the DWACD algorithm can estimate the underlying SRO accurately, and resampling the playback signal by the inverse phase term $\\Lambda_q^{-1}$ before transmission restores the no-SRO binaural cues at low and mid frequencies. In the MUSHRA test, the compensation condition scores well above the uncompensated condition and close to the hidden reference, establishing that audio-domain resampling is a viable substitute for explicit clock synchronization in stereo reproduction.","pith_inferences":["A straightforward extension would apply the same estimate-and-resample loop to more than two loudspeakers, provided the RTF matrix has linearly independent columns so the LCMV beamformer can separate them.","Because the paper assumes the primary device's own SRO $\\epsilon_0$ is known, a practical deployment would need joint estimation of both $\\epsilon_0$ and $\\epsilon_q$, perhaps through alternating updates; this is my inference, not the paper's claim.","The reported frequency dependence — full compensation at low and mid frequencies but not at high frequencies — suggests that residual high-frequency cue error is the next target, and a sub-sample delay refinement or multi-band approach may close the gap.","An online, non-oracle RTF estimator could be tested on the same simulated room and SRO configurations, turning the method into a fully blind system that does not require a single-source initialization phase."],"forward_implications":["Wireless loudspeaker systems can retain spatial fidelity without relying on PTP/NTP-style clock alignment, because compensation happens on the audio signal itself.","Offsets up to at least $\\pm 100$ ppm can be estimated robustly within about two minutes of audio, as shown in the estimation plots.","Perceived quality loss from SRO is reduced to a level close to the hidden reference in the MUSHRA listening test.","Binaural cues (ITD and IC) are preserved at low and mid frequencies, with high-frequency residuals smaller but not perfectly removed.","The method is source-agnostic: evaluation used Gaussian noise for objective cues and musical items for the listening test."],"supporting_citations":[{"why":"Supplies the time-frequency SRO signal model used to derive the phase term and the binaural signal expression.","marker":"[7]"},{"why":"States the validity condition for the SRO model, which the paper relies on in both the binaural and microphone signal derivations.","marker":"[8]"},{"why":"Provides the DWACD algorithm that the paper uses to estimate the SRO from the beamformer output and the reference signal.","marker":"[11]"},{"why":"Introduces the LCMV beamformer used as the spatial filter to isolate individual loudspeaker contributions.","marker":"[16]"},{"why":"Gives the original LCMV optimization formulation that underlies the beamformer weight computation in Eq. (16).","marker":"[18]"},{"why":"Supplies the Pyroomacoustics room simulation used to generate the microphone signals for the experimental evaluation.","marker":"[24]"},{"why":"Provides the STFT-based method used to apply the simulated SROs to the microphone signals.","marker":"[25]"},{"why":"Defines the MUSHRA methodology used for the subjective listening test.","marker":"[28]"}],"fun_headline_variants":["Resampling cancels clock skew that breaks stereo","Audio-domain resampling beats clock skew in wireless audio","Spatial filtering estimates SRO to preserve stereo image","Fix wireless speaker skew by resampling playback signal","Sample rate offset handled in audio domain for stereo"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire chain assumes the relative transfer function of each loudspeaker is known exactly (an oracle RTF computed from the true PSD matrix during a single-source initialization); if that RTF is imperfect in a real deployment, the beamformer will not isolate the loudspeaker contributions and the SRO estimate will degrade.","fun_headline_variants_meta":{"raw":{"variants":["Resampling cancels clock skew that breaks stereo","Audio-domain resampling beats clock skew in wireless audio","Spatial filtering estimates SRO to preserve stereo image","Fix wireless speaker skew by resampling playback signal","Sample rate offset handled in audio domain for stereo"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000494,"raw_usage":{"total_tokens":2399,"prompt_tokens":893,"completion_tokens":1506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":1432}},"tokens_in":509,"tokens_out":1506,"duration_ms":13573,"temperature":1.0,"reasoning_tokens":1432,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:26:41.335078+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical pipeline with an estimated RTF (for example, obtained from a single-source initialization frame with the same PSD estimator) instead of the oracle RTF of Eq. (19), in the simulated 7 m by 7 m by 6 m room with RT60 = 0.3 s and SROs (10, -100) ppm. If the estimated SRO trace deviates from ground truth by more than the smoothing tolerance, or the MUSHRA score for the compensation condition falls to within statistical noise of the uncompensated condition, the central claim that audio-domain SRO compensation preserves binaural cues fails in realistic conditions.","supporting_citations":[{"cited_title":"Blind samp ling rate offset estimation and compensation in wireless acoustic se nsor networks with application to beamforming,","cited_arxiv_id":null,"evidence_quote":"Supplies the time-frequency SRO signal model used to derive the phase term and the binaural signal expression."},{"cited_title":"Correlation maximization-based s ampling rate offset estimation for distributed microphone arrays,","cited_arxiv_id":null,"evidence_quote":"States the validity condition for the SRO model, which the paper relies on in both the binaural and microphone signal derivations."},{"cited_title":"O n Synchronization of Wireless Acoustic Sensor Networks in th e Presence of Time-V arying Sampling Rate Offsets and Speaker Changes,","cited_arxiv_id":null,"evidence_quote":"Provides the DWACD algorithm that the paper uses to estimate the SRO from the beamformer output and the reference signal."},{"cited_title":"Beamforming: A versatil e approach to spatial ﬁltering,","cited_arxiv_id":null,"evidence_quote":"Introduces the LCMV beamformer used as the spatial filter to isolate individual loudspeaker contributions."},{"cited_title":"An algorithm for linearly constrained ada ptive array processing,","cited_arxiv_id":null,"evidence_quote":"Gives the original LCMV optimization formulation that underlies the beamformer weight computation in Eq. (16)."},{"cited_title":"Pyroomacou stics: A python package for audio room simulation and array processing algo rithms,","cited_arxiv_id":null,"evidence_quote":"Supplies the Pyroomacoustics room simulation used to generate the microphone signals for the experimental evaluation."},{"cited_title":"Efﬁcient samp ling rate offset compensation - an overlap-save based approach,","cited_arxiv_id":null,"evidence_quote":"Provides the STFT-based method used to apply the simulated SROs to the microphone signals."},{"cited_title":"ITU-R BS.1534-3, October 2015","cited_arxiv_id":null,"evidence_quote":"Defines the MUSHRA methodology used for the subjective listening test."}],"review_version":1}