{"id":"c1012a83-2998-4bc5-9b0f-ad339074b373","arxiv_id":"2507.12122","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Soft-constrained spatially selective active noise control uses one parameter to continuously interpolate between full noise cancellation and perfect speech preservation, improving SNR, PESQ, and ESTOI over the hard-constrained design in simulations.","lead":"A new active noise control design for open-fitting earbuds lets a single dial trade off how much speech gets preserved against how much noise gets canceled, instead of insisting on zero speech distortion. A smart generalist might read it because it promises hearing aid and earbud users better speech quality without giving up noise suppression.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Secondary-path mismatch is the load-bearing assumption: if \\hat g \\neq g, the soft-constrained filter no longer implements the claimed trade-off, so the practical claim needs a robustness test.","rationale":"I verified Eqs. (21) and (27) by differentiating the respective objectives; the limiting relationships in Eq. (28) follow cleanly when the secondary path is known exactly. The mathematics is not the weak point. The load-bearing assumption is the perfect secondary-path estimate, exactly as the reader identified: with \\hat g\\neq g, the leakage estimate is contaminated by the control signal, so the soft-constrained filter no longer minimizes the intended objective and the speech-distortion trade-off is not guaranteed. This is a practical robustness issue rather than an internal inconsistency, and it justifies the conditional verdict. The lack of error bars, single anechoic scenario, and missing comparison to [14] are additional reasons for caution, but they are secondary to the secondary-path assumption.","tokens_in":9162,"tokens_out":9511,"duration_ms":115314,"concrete_test":"Re-run the Section 6 \\mu sweep with a perturbed secondary-path estimate, e.g., \\hat g = g + \\delta where \\delta is 10% RMS white noise plus a 5% gain error, or a measured g from a different insertion depth if available. Compare NR, \\Delta SNR_{intellig}, \\Delta PESQ, and \\Delta ESTOI at \\mu=10^{-2} and across the full sweep. If the soft-constrained operating point no longer dominates hard-constrained SSANC, or if the \\mu\\to\\infty limit no longer yields the hard-constrained speech-distortion level, the central practical claim fails under secondary-path mismatch. Additionally, re-derive the frequency-domain solution with \\hat G \\neq G to check whether Eq. (28) still holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The derivation of Eq. (21) and Eq. (27) is algebraically correct, but it relies on the identity p(n)=q^T x(n), which requires a perfect secondary-path estimate \\hat g=g in Eq. (6). With mismatch, the estimated leakage becomes \\hat p(n)=p(n)+(g-\\hat g)^T y(n), so the input vector x(n) contains a control-dependent component. The clean expression e(n)=(q+Gw)^T x(n) and its frequency-domain counterpart no longer hold; the soft-constraint penalty \\mu\\|H(q+Gw)-\\delta_\\Delta\\|^2 penalizes the wrong response, and the \\mu\\to\\infty limit need not preserve the desired speech. Because open-fitting hearables have secondary paths that vary with insertion, fit, and anatomy, the claimed practical benefit for a broad range of \\mu is not established outside the idealized exactly-known-g case simulated in Section 6. The paper explicitly assumes this after Eq. (6), but the conclusion frames the system as practical, making this the weakest load-bearing point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a soft-constrained spatially selective active noise control (SSANC) formulation for open-fitting hearables. A single scalar parameter mu is added as a quadratic penalty on speech distortion, yielding a closed-form time-domain filter (Eq. (21)) and a frequency-domain filter (Eq. (27)). The authors show analytically that conventional ANC and hard-constrained SSANC are recovered as mu approaches 0 and infinity, respectively (Eq. (28)), and they report a simulation study with one anechoic scenario in which an intermediate mu yields improved SNR, PESQ, and ESTOI relative to the hard-constrained design.","tokens_in":9352,"tokens_out":12021,"duration_ms":135832,"significance":"The derivation is self-contained and the algebra checks out: the proposed filters are valid least-squares solutions of a well-defined objective, and the limiting cases follow directly from the equations rather than from an assumed outcome. The simulation uses realistic measured impulse responses for a KEMAR-based open-fitting hearable, and the data provenance (impulse-response database, VCTK, NOISEX-92) is described well enough to support reproducibility. The main value is a simple, interpretable trade-off parameter that unifies two existing designs. The practical claims, however, currently rest on an ideal secondary-path estimate and on a single simulation scenario, which limits the strength of the conclusions until those points are addressed.","major_comments":[{"comment":"The identity p(n)=q^T x(n), on which the entire derivation rests, assumes a perfect secondary-path estimate, i.e., \\hat g = g, as stated after Eq. (6). When \\hat g ≠ g, the estimated leakage becomes \\hat p(n)=p(n)+(g-\\hat g)^T y(n), so the stacked input vector x(n) is control-dependent, Eq. (8) no longer holds, and the soft-constraint penalty in Eq. (20) penalizes a biased response. In particular, the mu→∞ limit in Eq. (28) need not preserve the desired speech. Section 6.1 uses the measured impulse response as the secondary-path estimate, so the simulation does not probe this assumption, which is critical for open-fitting hearables whose secondary paths vary with insertion and anatomy. Please add a robustness study, e.g., re-running Figure 3 with perturbed \\hat g (gain error, delay error, or re-insertion impulse responses) and reporting the resulting NR, SD, SNR, PESQ, and ESTOI curves, or explicitly restrict the practical claims to perfectly known secondary paths.","section":"Section 2, Eq. (8)"},{"comment":"The conclusion that a 'broad range' of mu provides substantial improvements over the hard-constrained design is supported by only a single anechoic scenario: one speech source, two noise sources, one input SNR (−5 dB), one set of relative impulse responses, and no error bars or multiple realizations. This evidence is insufficient to establish the generalizability of the practical claim. Please add variations in source positions, input SNR, reverberation, and repeated hearable insertions, or temper the conclusions to the specific configuration tested.","section":"Section 6.3, Fig. 3"}],"minor_comments":[{"comment":"The sentence 'As the secondary path estimate we used the measured impulse response between the outer receiver and the inner error microphone' needs a comma after 'estimate'; also, 'the outer receiver at the right ear as the secondary source' could be phrased more precisely as 'the right-ear outer receiver served as the secondary source.'","section":"Section 6.1"},{"comment":"The notation is inconsistent: h appears without its frequency argument in the denominator terms h^H Φ_x^{-1}(ω) h, while h(ω) is used elsewhere. Please harmonize the notation throughout Section 5.","section":"Section 5, Eqs. (27)-(28)"},{"comment":"The definition of δ_Δ uses underbraced length expressions that are difficult to parse; rewriting with explicit index positions (e.g., 1 at index L_a+Δ) would improve clarity.","section":"Eq. (16c)"},{"comment":"The intelligibility-weighted spectral distortion SD_intellig is reported in negative dB, with more negative values meaning less distortion; this sign convention should be stated explicitly in Section 6.2 or in the Figure 3 caption to avoid ambiguity.","section":"Section 6.2 and Fig. 3"},{"comment":"No confidence intervals or sensitivity analyses are provided for the reported metrics; even in a single-scenario study, rerunning with a different speech utterance or noise realization would strengthen the claim that the observed improvements are not sample-specific.","section":"Section 6.3"}],"recommendation":"major_revision","confidential_remarks":"The technical core of the paper is sound and within the scope of the journal; the derivations are correct and the limiting cases are genuine. My main concern is that the paper's practical framing is not yet supported without a secondary-path mismatch analysis and a broader simulation study. I would support acceptance after those additions, and I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this paper does exactly what it says, quietly and correctly. It adds a quadratic penalty for speech distortion to the SSANC objective, derives the time- and frequency-domain closed forms, and shows that conventional ANC and hard-constrained SSANC are the mu=0 and mu->infinity limits. The algebra in Eqs. (21), (27), and (28) checks out, and the simulations match the theory. That is a legitimate, if modest, contribution to the hearable ANC subfield.\n\nWhat is new: the explicit soft-constrained formulation for spatially selective ANC, with a single frequency-independent parameter. That specific packaging is not in the cited literature. The paper also does a decent job of positioning itself against prior work, including the authors' own hard-constrained SSANC, and the limiting-case relationship is stated cleanly.\n\nWhere the soft spots are, in proportion: First, the perfect secondary-path assumption (bg = g, used right after Eq. (6)) is load-bearing. When the estimate is wrong, the leakage estimate bp(n) is contaminated by the anti-noise, the neat expression e(n) = (q + Gw)^T x(n) breaks down, and the soft-constrained penalty penalizes the wrong response. The paper is upfront about the assumption, but then the abstract and conclusion call the method \"practical\" without testing this. Second, the experimental section is a single anechoic scenario, one speech source, two noise sources, no error bars or repeated trials, and no code released. Third, the closest related work—Serizel et al. [14], which already used speech-distortion weighting in integrated ANC—is cited but never compared against. That comparison would sharpen the novelty claim.\n\nThese are fixable rather than fatal. The core derivation is sound, the idea is clearly useful, and the paper does not overclaim technically, only in its practical framing. A serious referee could reasonably accept this after adding a robustness test for secondary-path mismatch, a comparison to [14], and softer conclusions. If an editor wanted to desk-reject on novelty, I would push back; the extension is small but real and will be cited.\n\nFor a reading group, it is worth a skim if you work on ANC or spatial filtering. For my own work, I would cite it as the reference for soft-constrained SSANC. Recommendation: yes, send it to peer review, with expectations of a revision.","headline":"A clean, correctly derived extension of SSANC with a tunable trade-off; the math is solid, but the practical claims outrun the single idealized simulation.","tokens_in":9900,"tokens_out":1447,"would_cite":true,"duration_ms":19870,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single trade-off parameter connects full noise cancellation to distortion-free speech preservation in hearable ANC.","keywords":["active noise control","spatially selective","open-fitting hearables","soft-constrained optimization","speech distortion","noise reduction","SSANC","trade-off parameter"],"falsifier":"Run the same open-fitting hearable simulation while deliberately corrupting the secondary-path estimate, for example by scaling $\\hat{g}$ by 1.1 or delaying it by one sample, and check whether the intermediate operating point $\\log_{10}\\mu=-2$ still beats the hard-constrained baseline on SNR improvement. If it does not, the claimed practical advantage depends on exact secondary-path knowledge. Alternatively, replace the anechoic impulse responses with reverberant ones and test whether the reported PESQ and ESTOI gains at that $\\mu$ persist.","tokens_in":8941,"feed_emoji":"🎧","tokens_out":19009,"duration_ms":186596,"temperature":0.7,"pith_summary":"Open-fitting hearables leave the ear partly open, so both unwanted noise and the desired frontal speech leak to the eardrum; conventional active noise control cancels all of it, while hard-constrained spatially selective ANC preserves a chosen speech direction exactly but at the cost of limited noise reduction. The paper tries to establish that both are the two endpoints of one soft-constrained design: instead of imposing zero speech distortion as a hard equality, it adds a quadratic penalty with a single frequency-independent parameter $\\mu$. The resulting filter is derived in the time and frequency domains (Eqs. (21) and (27)), with $\\mu=0$ recovering conventional ANC and $\\mu\\to\\infty$ recovering hard-constrained SSANC. Simulations on a pair of open-fitting hearables with one frontal speech source and two noise sources show that intermediate $\\mu$ (e.g., $\\log_{10}\\mu=-2$) gives more noise reduction than the hard-constrained design while accepting modest speech distortion, improving intelligibility-weighted SNR by 17.2 dB, perceived speech quality (PESQ) by 0.54, and objective intelligibility (ESTOI) by 0.39 over the no-control leakage. The reason to care is practical: a single knob, requiring no hardware change, lets a wearer choose how much hearing-through speech to keep versus how much ambient noise to silence.","feed_headline":"One knob spans full noise-canceling and speech-preserving ANC","feed_subtitle":"A wide middle range of the trade-off setting improves SNR, PESQ, and ESTOI over the hard-constrained design.","key_machinery":"The load-bearing object is the soft-constrained cost $\\min_w E\\{e^2(n)\\}+\\beta\\|w\\|^2+\\mu\\|H(q+Gw)-\\delta_\\Delta\\|^2$, whose normal equations produce the time-domain filter (21). In the frequency domain the identical relaxation produces a regularized projection: the scalar $1/\\mu$ sits inside the denominator of Eq. (27), so $\\mu$ controls how much of the distortionless speech-preservation term is admixed into the pure noise-cancellation solution. The matrix $H$ is the convolution matrix of relative impulse responses (ReIRs), i.e., the transfer paths from each outer microphone to the reference microphone; $q+Gw$ represents the total path from input to the inner error microphone, so $H(q+Gw)-\\delta_\\Delta$ measures the speech deviation from the desired delayed reference, and $\\mu$ converts that deviation from a hard equality constraint into a quadratic penalty.","core_discovery":"The central claim is that the optimal soft-constrained SSANC filter is an interpolation between two known filters. In the frequency domain the solution is $w_{\\mathrm{soft}}(\\omega)=\\frac{1}{G^{*}(\\omega)}\\left[-q_{\\omega}+\\frac{\\Phi_x^{-1}(\\omega)h(\\omega)e^{i\\omega\\Delta}}{1/\\mu+h^{H}(\\omega)\\Phi_x^{-1}(\\omega)h(\\omega)}\\right]$, so that $\\lim_{\\mu\\to 0}w_{\\mathrm{soft}}(\\omega)=w_{\\mathrm{ANC}}(\\omega)$ and $\\lim_{\\mu\\to\\infty}w_{\\mathrm{soft}}(\\omega)=w_{\\mathrm{hard}}(\\omega)$. The same interpolation is expressed in the time domain by Eq. (21), where the regularized covariance $\\Phi_{rr}+\\mu G^{T}H^{T}HG$ appears. Because the distortionless constraint is relaxed rather than enforced, the optimizer is free to remove more of the leakage noise; simulations on measured anechoic impulse responses from an open-fitting earpiece confirm that for intermediate $\\mu$ (e.g., $\\log_{10}\\mu=-2$) the system reaches 20.7 dB noise reduction at the cost of $-14.7$ dB intelligibility-weighted spectral distortion, yielding an SNR improvement of 17.2 dB, a PESQ gain of 0.54, and an ESTOI gain of 0.39, each larger than the hard-constrained baseline.","pith_inferences":["The paper fixes $\\mu$ to a single frequency-independent value; a natural extension is to let $\\mu$ vary across frequency bands, tuning the trade-off where speech intelligibility matters most and full cancellation elsewhere.","The derivation assumes a perfectly known secondary path; because the soft penalty is not an equality constraint, the design may tolerate secondary-path mismatch better than the hard-constrained one, a robustness that could be tested by perturbing the path estimate.","The same soft-constraint cost transfers directly to closed-fitting and open-ear ANC devices and to feedback controllers, since it needs only an inner error signal, a leakage estimate, and relative impulse responses.","The reported gains come from an anechoic, stationary scenario; a reverberant or moving-talker test would show whether a single fixed $\\mu$ generalizes across scenes or whether an adaptive $\\mu$ scheduler is needed."],"forward_implications":["Conventional ANC and hard-constrained SSANC are not rival designs but limiting cases of a single filter parameterized by $\\mu$, so one implementation covers both operating modes.","There is a broad intermediate range of $\\mu$ in which SNR improvement, PESQ, and ESTOI all exceed the hard-constrained baseline, meaning zero speech distortion is not the best operating point for perceived quality in this setup.","The soft constraint requires no extra microphones or hardware; it changes only the optimization criterion used to compute the control filter.","The frequency-domain identity holds for any positive-definite input covariance and known secondary path, so the interpolation between conventional ANC and hard-constrained SSANC is not tied to the particular simulation setup."],"supporting_citations":[{"why":"Introduces the hard-constrained SSANC optimization problem that this paper relaxes, and supplies the baseline filter and results it compares against.","marker":"[8]"},{"why":"Establishes the delay needed to preserve the desired speech component at the inner error microphone, motivating the causal constraint structure used here.","marker":"[20]"},{"why":"Introduces acausal relative impulse responses into the SSANC optimization, providing the H-matrix formulation the soft constraint penalizes.","marker":"[21]"},{"why":"Provides the measured anechoic impulse responses of the open-fitting earpiece used for the simulations.","marker":"[29]"},{"why":"Supplies the speech corpus used as the desired frontal speech source in the evaluation.","marker":"[30]"},{"why":"Supplies the babble-noise signals used as the two undesired noise sources in the evaluation.","marker":"[31]"},{"why":"Defines the PESQ metric used to measure speech quality improvement with and without control.","marker":"[33]"},{"why":"Defines the ESTOI metric used to measure intelligibility improvement with and without control.","marker":"[34]"},{"why":"Shows how ANC and noise reduction can be integrated for hearing aids, the problem class this paper extends to a soft-constrained SSANC design.","marker":"[6]"},{"why":"Introduces the speech-distortion-weighted multichannel Wiener filter whose soft-constraint idea is carried over to ANC here.","marker":"[22]"}],"fun_headline_variants":["Soft constraint lets hearables dial between ANC and speech focus","Tunable soft-constrained ANC broadens noise-cancel to speech-preserve","Hearable ANC: soft constraint interpolates from full cancel to full speech","Soft-constrained SSANC: one parameter spans noise and speech control","Beyond hard constraint: soft SSANC improves SNR and speech quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the acoustic path from the loudspeaker to the inner error microphone (the secondary path) is known exactly, $\\hat{g}=g$; if that estimate is wrong, the extracted leakage $\\hat{b}_p(n)$ is biased, the equality $p(n)=q^T x(n)$ breaks down, and the optimized soft-constrained filter, like the hard-constrained one, no longer implements the intended trade-off between speech distortion and noise reduction.","fun_headline_variants_meta":{"raw":{"variants":["Soft constraint lets hearables dial between ANC and speech focus","Tunable soft-constrained ANC broadens noise-cancel to speech-preserve","Hearable ANC: soft constraint interpolates from full cancel to full speech","Soft-constrained SSANC: one parameter spans noise and speech control","Beyond hard constraint: soft SSANC improves SNR and speech quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1622,"prompt_tokens":1031,"completion_tokens":591,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":498}},"tokens_in":647,"tokens_out":591,"duration_ms":6405,"temperature":1.0,"reasoning_tokens":498,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:53:30.159014+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same open-fitting hearable simulation while deliberately corrupting the secondary-path estimate, for example by scaling $\\hat{g}$ by 1.1 or delaying it by one sample, and check whether the intermediate operating point $\\log_{10}\\mu=-2$ still beats the hard-constrained baseline on SNR improvement. If it does not, the claimed practical advantage depends on exact secondary-path knowledge. Alternatively, replace the anechoic impulse responses with reverberant ones and test whether the reported PESQ and ESTOI gains at that $\\mu$ persist.","supporting_citations":[{"cited_title":"Spatially selective active noise control systems,","cited_arxiv_id":null,"evidence_quote":"Introduces the hard-constrained SSANC optimization problem that this paper relaxes, and supplies the baseline filter and results it compares against."},{"cited_title":"Effect of target signals and delays on spatially selective active noise control for open-fitting hearables,","cited_arxiv_id":null,"evidence_quote":"Establishes the delay needed to preserve the desired speech component at the inner error microphone, motivating the causal constraint structure used here."},{"cited_title":"Spatially selective active noise control for open-fitting hearables with acausal optimization,","cited_arxiv_id":null,"evidence_quote":"Introduces acausal relative impulse responses into the SSANC optimization, providing the H-matrix formulation the soft constraint penalizes."},{"cited_title":"The hearpiece database of individual transfer functions of an in-the-ear earpiece for hearing device research,","cited_arxiv_id":null,"evidence_quote":"Provides the measured anechoic impulse responses of the open-fitting earpiece used for the simulations."},{"cited_title":"Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,","cited_arxiv_id":null,"evidence_quote":"Supplies the babble-noise signals used as the two undesired noise sources in the evaluation."},{"cited_title":"Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,","cited_arxiv_id":null,"evidence_quote":"Defines the PESQ metric used to measure speech quality improvement with and without control."},{"cited_title":"Integrated active noise control and noise reduction in hearing aids,","cited_arxiv_id":null,"evidence_quote":"Shows how ANC and noise reduction can be integrated for hearing aids, the problem class this paper extends to a soft-constrained SSANC design."},{"cited_title":"GSVD-based optimal filtering for single and multimicrophone speech enhancement,","cited_arxiv_id":null,"evidence_quote":"Introduces the speech-distortion-weighted multichannel Wiener filter whose soft-constraint idea is carried over to ANC here."}],"review_version":1}