{"id":"f9286625-1d98-4a61-9541-6c906d68a929","arxiv_id":"1908.04672","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On 590 simulated room impulse responses, spectral subtraction with an energy-decay RT60 estimator raised Chirp decode rates from 54.53% to 79.78% for audible signals and from 86.15% to 92.11% for ultrasonic signals.","lead":"Reverberation reduces how often audible and ultrasonic machine-to-machine audio packets decode correctly. The authors report that a single-channel spectral subtraction method, paired with an RT60 estimator adapted to Chirp subband decays, raises simulated decode rates from 54.5% to 79.8% for audible and from 86.2% to 92.1% for ultrasonic signals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RT60 estimation may be biased by the Chirp signal's own spectral leakage within the -5 to -35 dB regression window; the average error in §IV-A does not rule out contamination that would mis-estimate Eq. (5) and reduce the claimed 25 pp decode-rate gain.","rationale":"The paper's headline result is the 25 pp (audible) and 6 pp (inaudible) decode-rate improvement from spectral subtraction combined with the proposed RT60 estimator. This claim requires the RT60 estimate to be accurate enough for Eq. (5) to model the late-reverberation PSD. The reader's weakest_assumption identifies exactly this dependence: the -5 to -35 dB regression window may not isolate the RIR tail, since overlapping tones, noise, or codec spectral structure can contaminate the decay. I agree with that assessment. The paper presents plausible independent support: a large simulated corpus (118,000 signals) and a real Chirp decoder, plus a comparison with two other dereverberation algorithms. Those elements strengthen the case that some form of spectral subtraction helps, but they do not test the specific RT60 estimator's internal validity. The reported validation of RT60 (mean error 0.11 s for RT60 < 2 s, §IV-A) is aggregate and does not show whether the error is concentrated in the exact subbands and time frames where the decoder depends on suppression. An anechoic control experiment is the cheapest decisive check: if the estimator reports nonzero RT60 on signals with no reverberation, the regression is not measuring room decay. If that test passes, the concern is resolved; if it fails, the reported decode-rate gain may be partly due to a generic attenuator rather than the proposed parameter estimation. Because this concern is already embedded in the reader's conditional verdict and does not by itself force a rejection, the appropriate verdict remains CONDITIONAL.","tokens_in":6715,"tokens_out":9009,"duration_ms":96188,"concrete_test":"Compute the STFT subband energy envelopes of anechoic (unconvolved) Chirp signals and apply the same peak-picking, offset, and -5 to -35 dB linear regression used in §III-A. If the median fitted 'RT60' on anechoic signals exceeds 0.05 s, the regression window is contaminated by the signal's own tone transitions and window sidelobes, so the RT60 estimates in §IV-A are biased; conversely, if the anechoic fit is near zero, the assumption of a clean RIR tail in this window survives for the tested signals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central decode-rate claim depends on the RT60 estimate being accurate enough for Eq. (5)'s exponential attenuation. Section III-A assumes that, after the peak and a fixed offset, the STFT subband energy envelope between -5 dB and -35 dB follows the room's RIR decay. However, Chirp packets are sequences of FSK tones that can start while a previous tone's subband is still decaying, and the analysis window's finite sidelobes allow adjacent tones to leak into that subband. Additionally, FSK may reuse the same frequency at later times, causing a second tone to interrupt the decay before it reaches -35 dB. The paper never checks the anechoic (non-reverberant) subband envelope for such self-contamination. The reported RT60 validation (mean error 0.11 s for RT60 < 2 s) is averaged over 590 RIRs and does not reveal failure rates in the specific time-frequency regions used for decoding. If the regression window is contaminated, RT60 is biased, the late-reverberation PSD is mis-estimated, and the gain function may over-suppress the direct path or under-suppress reverberation, reducing the actual decode-rate benefit.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses reverberation-induced decoding failures in machine-to-machine (M2M) Data-over-Sound signals, specifically Chirp's FSK-based codec. It proposes a single-channel dereverberation pipeline: (i) estimate the room's RT60 from the energy decay of individual STFT subbands; (ii) estimate the late-reverberation PSD from a delayed and attenuated copy of the observed PSD; and (iii) apply spectral subtraction with a gain floor. The method is evaluated on 118,000 simulated reverberant signals (59,000 audible and 59,000 ultrasonic) generated by convolving 100 Chirp packets with 590 AcouSP room impulse responses. The authors report an increase in decode rate from 54.53% to 79.78% for audible signals and from 86.15% to 92.11% for inaudible signals, together with log-spectral-distortion and reverberation-reduction metrics, and a comparison against LP residual cepstrum and source enhancement.","tokens_in":6926,"tokens_out":5448,"duration_ms":59099,"significance":"If the results hold, the paper offers a practically useful single-channel preprocessing method for data-over-sound in reverberant environments, a domain that is less studied than speech dereverberation. The evaluation scale is a strength: using 118,000 reverberant signals and a real Chirp decoder gives the result ecological validity, and the comparison with two established dereverberation baselines is useful. The reported gains are also broadly consistent with the improvement in LSD and RR. However, the technical presentation is incomplete in several load-bearing places, and the statistical evaluation lacks confidence intervals or significance tests, so the current manuscript does not yet support the strength of the claims made.","major_comments":[{"comment":"Equation (5) uses the decay-rate constant Δ (or Delta) in the exponential factor e^{-2ΔT}, but Δ is never defined in the manuscript. Without a definition relating Δ to RT60 (e.g., Δ = 3 ln(10)/RT60), the core dereverberation equation cannot be implemented or checked, and the link between the proposed RT60 estimator and the attenuation applied in Eq. (6) is not established. In addition, the sentence '80ms was found to be the best resulting value of T' gives no tuning procedure. If T was selected on the same 118,000-signal evaluation set used in Table I, the reported decode-rate improvements may be optimistic; the authors should describe the tuning protocol and, ideally, a validation split.","section":"§III-B, Eq. (5)"},{"comment":"The RT60 estimator assumes that, after the energy peak and a fixed offset, the STFT subband decay between -5 dB and -35 dB is governed by the room's impulse response. This assumption is load-bearing because a biased RT60 directly biases the late-reverberation PSD in Eq. (5) and hence the gain function. The paper does not check the anechoic subband envelopes for self-contamination from FSK tones: overlapping tones, spectral leakage from adjacent subbands, or repeated frequencies re-entering the decay window before -35 dB could all corrupt the regression. The validation in §IV-A reports only the overall mean error (0.11 s for RT60 < 2 s), averaged over 590 RIRs, and does not report failure rates or errors in the specific time-frequency regions used by the decoder. The authors should add an anechoic-control experiment and a per-RIR error analysis.","section":"§III-A, RT60 estimation"},{"comment":"The abstract and the introduction state that the dereverberation method was shortlisted through a pilot test, but no pilot test is described anywhere in the manuscript. The absence of the pilot protocol and selection criterion makes the method-selection step non-reproducible. Please add a description of the pilot test, including which methods were compared, the data used, and the metric that led to selecting spectral subtraction.","section":"Abstract, §I, and §VI"}],"minor_comments":[{"comment":"The decode-rate change in Table I is reported without confidence intervals or significance tests. Given the large sample size the effect is likely real, but because the 100 packets are reused across all 590 RIRs, a clustered analysis or a paired confidence interval would strengthen the claim.","section":"Table I and §V-B2"},{"comment":"The 'After' decode rates in Table II (77.79% audible, 87.79% inaudible) differ from the corresponding values in Table I (79.78% and 92.11%). The text should explicitly state that Table II uses a much smaller set (one packet per protocol, 590 RIRs) and explain why the absolute numbers differ.","section":"§IV-B2 and Table II"},{"comment":"Equation (13) uses both X(k,l) and S(k,l), while the surrounding text refers to c(n) and x(n). The clean-signal symbol should be defined consistently, and the notation for the STFT of the clean signal should be introduced before the LSD formula.","section":"§IV-B1, Eq. (13)"},{"comment":"The text says a positive correlation is noticed between reverberation reduction and RT60, but no correlation coefficient or fit is reported. Please include the numeric correlation and, if possible, a scatter plot with a trend line.","section":"Figure 8 and §V-B1"},{"comment":"There are several typos and formatting issues: 'dereveberation' in §IV-B, 'the the' in the introduction, and 'signiﬁes' in §V-B1. These should be corrected in a final pass.","section":"Throughout"},{"comment":"Reference [13] is cited as the AcouSP RIR database, but the listed reference is a recommendation document for annotation of acoustic data collections. Please cite the actual RIR database with a URL and access date, or clarify how the 590 RIRs were obtained.","section":"Reference [13]"}],"recommendation":"major_revision","confidential_remarks":"The paper currently reads more like a conference manuscript than a journal article. The undefined decay constant in Eq. (5) and the missing tuning details are the main blockers; once these are supplied, the empirical evaluation, though in need of standard error reporting, is potentially publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does what it says: it adapts Habets' spectral subtraction to M2M Chirp signals, adds an RT60 estimator based on subband EDC regression, and reports a 25 percentage-point decode-rate gain on 59,000 audible signals. That is a real, useful engineering result for the data-over-sound niche.\n\nWhat's genuinely new is the adaptation to non-speech FSK tones, the per-band RT60 estimator, and the evaluation on 590 AcouSP RIRs with a real decoder. The evaluation is substantial: 118,000 reverberant signals, and the decode-rate gains are consistent with the reported LSD and RR improvements. The comparison against LP residual cepstrum and source enhancement is fair, and the citation pattern is honest—Habets, Polack, and AcouSP are the right anchors.\n\nThe soft spots are in the reproducibility details. T=80ms is picked without a described tuning procedure; the pilot test that shortlisted the method is not reported; there are no confidence intervals or significance tests on any of the decode-rate tables; and no code is supplied. The stress-test concern about RT60 bias from spectral leakage is reasonable: the paper never checks the anechoic subband envelope to see whether overlapping tones or FSK reuse contaminate the -5 to -35 dB regression window. If that bias is real, Eq. (5) misestimates the late-reverb PSD, which could change the gain function. But the algorithm's smoothing (beta=0.9) and gain floor (lambda=0.1) may make it robust in practice; I'd want the anechoic checks or a failure-rate analysis to know whether the 25 pp number is even understated. The two-signal comparison with other algorithms is a minor weakness, and the paper acknowledges the computational constraint.\n\nWho should read this: anyone working on data-over-sound, IoT acoustic signalling, or applying dereverberation to non-speech signals. It is a modest extension of known techniques, not a breakthrough, but it is a solid application with a clear evaluation. A serious referee could fix the missing details; the central claim is plausible and the method is not circular.\n\nRecommendation: send it to peer review, but demand the tuning procedure, confidence intervals, and ideally code or a public dataset before acceptance.","headline":"A plainly written, well-scoped extension of single-channel spectral subtraction to data-over-sound M2M Chirp signals, with a large simulated evaluation and a genuinely useful decode-rate gain—held back by missing tuning details, absent significance tests, and no code.","tokens_in":7526,"tokens_out":1550,"would_cite":false,"duration_ms":16787,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A simple echo-removal preprocessor raises data-over-sound decode rates by 25 percentage points.","keywords":["data-over-sound","machine-to-machine signalling","dereverberation","spectral subtraction","reverberation time estimation","Chirp codec","room impulse response","ultrasonic communication"],"falsifier":"Take a Chirp packet with a known RT60 synthetic tail and add a second overlapping monophonic tone that starts inside the -5 dB to -35 dB regression window; if the estimated RT60 shifts by more than the reported roughly 0.2 second mean absolute error, or the post-dereverberation decode rate stops improving, the clean-decay assumption is the failure point.","tokens_in":6449,"feed_emoji":"🔊","tokens_out":5632,"duration_ms":52807,"temperature":0.7,"pith_summary":"The paper argues that reverberation, not noise or codec weakness, is the main correctable barrier to machine-to-machine data-over-sound, and that a single-channel spectral subtraction preprocessor can remove most of that barrier. The authors propose a reverberation-time estimator that reads the decay slope of each frequency band of a received Chirp packet, then uses that estimate to subtract the predicted late-reverberation power before the decoder sees the signal. On 118,000 simulated room-transmission tests they report decode-rate increases of 25 percentage points for audible and 6 percentage points for ultrasonic Chirp signals, along with improved log-spectral distance for every tested room impulse response. The claim matters because it offers a decoder-agnostic preprocessing stage for IoT and device-to-device audio links in ordinary reverberant rooms.","feed_headline":"Echo-removal preprocessor lifts data-over-sound success by 25 points","feed_subtitle":"The single-channel spectral subtraction step also boosts ultrasonic signals by 6 points and beats slower rivals in tests.","key_machinery":"The load-bearing object is the per-frequency-band energy decay curve (EDC) of the received Chirp packet, combined with the Polack-style assumption that this decay follows the room impulse response's exponential envelope. The RT60 estimator locates each subband's energy peak, starts the decay curve at a fixed offset after the tone, regresses log-energy from -5 dB to -35 dB, and maps the regression's x-intercept to a 60 dB decay time; thresholded subbands are excluded and the survivors are averaged. That RT60 estimate is then inserted into the late-reverberation PSD model $\\gamma_{x_r x_r}(k,l)=e^{-2\\Delta T}\\gamma_{xx}(k,l-T)$, whose output drives the spectral subtraction gain $G(k,l)=1-1/\\sqrt{\\text{SNR}_{\\text{post}}(k,l)}$ with smoothing and a non-zero floor. This machinery is what lets a single microphone, with no known impulse response, turn room acoustics into a predictable quantity.","core_discovery":"On its own terms, the paper's central discovery is that a classic speech-dereverberation technique, spectral subtraction, transfers to non-speech data-carrier signals when its input parameter, the reverberation time RT60, is estimated from the received signal itself rather than assumed or measured in advance. The authors model each STFT subband's post-tone energy decay as a straight line in log energy, fit it by least squares between -5 dB and -35 dB below the peak, read off RT60 from the 60 dB intercept, and feed that estimate into the standard delayed and attenuated late-reverberation PSD prediction of spectral subtraction. Evaluated against a real Frequency-Shift-Keying Chirp decoder, the complete preprocessor raises the audible decode rate from 54.53% to 79.78% and the ultrasonic rate from 86.15% to 92.11% across 59,000 convolved signals per band. In a small head-to-head comparison it also decodes better and runs faster than LP residual cepstrum and source-enhancement alternatives.","pith_inferences":["If the decay-slope estimator is as robust as reported, it could serve as a standalone single-channel RT60 meter for any tonal or pulsed signal, not only Chirp packets, since it needs only a post-burst decay tail.","The same pipeline should transfer to other frequency-shift-keying data-over-sound codecs; the argument depends on tone-like stationarity and a quiet tail, not on Chirp's specific packet structure.","A harder test the paper does not run is live-room playback: the 118,000 signals are all RIR convolutions, and real microphone noise, movement, and non-stationary interferers could break the clean-decay assumption.","One could extend the estimator to jointly estimate the noise floor and early-to-late energy ratio, which would let the preprocessor adapt its -35 dB regression limit and non-zero gain floor per room."],"forward_implications":["A decoder-agnostic preprocessor can sit in front of an existing Chirp receiver: no changes to the modulation, error-correction, or decoding logic are needed to get the reported gains.","Audible data-over-sound links become usable in ordinary rooms, where RT60 below 2 seconds shows mean RT60 estimation error around 0.11 seconds.","Ultrasonic links gain only modestly because they already decode well; the method's value there is small but positive, at 6 percentage points.","Because the gain computation costs about 1.76 seconds per signal versus 3.2 to 17.8 seconds for the compared methods, the preprocessor is viable for real-time or embedded use.","For very dry rooms with RT60 below 0.6 seconds the paper observes a slight decode-rate decrease, so the preprocessor is best applied when reverberation is actually present."],"supporting_citations":[{"why":"Supplies the spectral subtraction gain function and the delayed and attenuated late-reverberation PSD estimate (Equations 5-9) that the proposed algorithm builds on.","marker":"[11]"},{"why":"Provides the exponentially decaying statistical room-impulse-response model that justifies reading RT60 from the slope of each subband's energy decay.","marker":"[12]"},{"why":"Supplies the 590 room impulse responses used to convolve 100 audible and 100 ultrasonic Chirp signals, yielding the 118,000 test signals behind the decode-rate tables.","marker":"[13]"},{"why":"Defines the log spectral distortion metric used to show the dereverberated output is closer to the anechoic signal than the reverberant input.","marker":"[10]"},{"why":"Supplies the near-ultrasound communication context used to explain why ultrasonic Chirp signals are less affected by reverberation and gain less from dereverberation.","marker":"[9]"}],"fun_headline_variants":["Self-estimated RT60 enables echo removal, boosting data-over-sound by 25 points","Auto-estimating acoustic parameters makes echo removal work for M2M data","Dereverberation with in-situ RT60 estimation halves data-over-sound errors","Speech-grade echo cancellation moves to machine-to-machine acoustics","Echo-adaptive preprocessor cuts M2M acoustic errors by over half"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that after each Chirp tone the energy in each frequency band decays as one clean exponential tied to the room's reverberation time, so a straight-line fit over the 30 dB tail recovers the true RT60; overlapping tones or noise that contaminate that tail bias the estimate and the whole dereverberation.","fun_headline_variants_meta":{"raw":{"variants":["Self-estimated RT60 enables echo removal, boosting data-over-sound by 25 points","Auto-estimating acoustic parameters makes echo removal work for M2M data","Dereverberation with in-situ RT60 estimation halves data-over-sound errors","Speech-grade echo cancellation moves to machine-to-machine acoustics","Echo-adaptive preprocessor cuts M2M acoustic errors by over half"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001018,"raw_usage":{"total_tokens":4288,"prompt_tokens":929,"completion_tokens":3359,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":3257}},"tokens_in":545,"tokens_out":3359,"duration_ms":24792,"temperature":1.0,"reasoning_tokens":3257,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:35:51.861628+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a Chirp packet with a known RT60 synthetic tail and add a second overlapping monophonic tone that starts inside the -5 dB to -35 dB regression window; if the estimated RT60 shifts by more than the reported roughly 0.2 second mean absolute error, or the post-dereverberation decode rate stops improving, the clean-decay assumption is the failure point.","supporting_citations":[{"cited_title":"Single-channel speech dereverberation based on spectral subtraction,","cited_arxiv_id":null,"evidence_quote":"Supplies the spectral subtraction gain function and the delayed and attenuated late-reverberation PSD estimate (Equations 5-9) that the proposed algorithm builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the exponentially decaying statistical room-impulse-response model that justifies reading RT60 from the slope of each subband's energy decay."},{"cited_title":"The acousp recommendation for annotation of acoustic data collections. (version 1.0),","cited_arxiv_id":null,"evidence_quote":"Supplies the 590 room impulse responses used to convolve 100 audible and 100 ultrasonic Chirp signals, yielding the 118,000 test signals behind the decode-rate tables."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the log spectral distortion metric used to show the dereverberated output is closer to the anechoic signal than the reverberant input."},{"cited_title":"Near-ultrasound communication for tv’s 2nd screen services,","cited_arxiv_id":null,"evidence_quote":"Supplies the near-ultrasound communication context used to explain why ultrasonic Chirp signals are less affected by reverberation and gain less from dereverberation."}],"review_version":1}