{"id":"c416c679-8e2b-4caa-8528-362c89a11cc7","arxiv_id":"2501.18227","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"BSM-iMagLS optimizes binaural reproduction filters from small head-mounted microphone arrays by jointly matching magnitude, magnitude slope, and interaural level difference, cutting ILD errors roughly in half to near-JND levels with only a small magnitude error penalty.","lead":"Binaural listening for augmented and virtual reality needs the sound at each ear to match what a listener would naturally hear. This paper introduces a new method that tunes the audio filters from head-mounted microphones to improve the left-right loudness differences that help us locate sounds, while keeping the overall tone quality nearly unchanged.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'comparable magnitude accuracy' claim is not supported under reverberation: Table IV shows BSM-iMagLS LSD error up to ~3.2 dB worse than BSM-MagLS (5-mic glasses), so the central claim needs scoping or correction.","rationale":"The reader's verdict is CONDITIONAL and already flags the reverberation evidence in its rationale, so my stress-test does not move the verdict. My primary concern differs from the reader's stated weakest_assumption: rather than lambda sensitivity, I focus on the direct contradiction between the abstract's 'comparable magnitude accuracy' claim and Table IV. The paper's own data show that under reverberation the proposed method has materially larger LSD error than BSM-MagLS for three of the four arrays, with the largest gap (~3.2 dB) for the head-mounted glasses array. Because LSD is described in the paper as a coloration/magnitude-spectral measure, this is a magnitude-accuracy failure in exactly the class of wearable arrays the method targets. The ILD improvements and listening-test results are genuine support for the method's spatial benefit, and the anechoic magnitude results do support a narrower claim, so the appropriate remedy is to scope the central claim to anechoic/normalized-magnitude conditions or to present the reverberant LSD trade-off explicitly in the abstract. A paired statistical re-analysis of Table IV is the decisive check; if the differences are not robust, the original claim can stand.","tokens_in":20601,"tokens_out":4629,"duration_ms":43890,"concrete_test":"Recompute the reverberant LSD comparison of Table IV per array and room with paired statistics over the 100 source positions, testing BSM-iMagLS against BSM-MagLS. If the 5-mic glasses and 4-mic EasyCom mean differences remain above 1 dB and are significant after correction (p < 0.05), the 'comparable magnitude accuracy' claim should be amended to specify anechoic, free-field normalized-magnitude conditions; if the differences vanish or drop below an audibility threshold, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim asserts that BSM-iMagLS maintains 'comparable magnitude accuracy' to state-of-the-art solutions. The anechoic normalized magnitude-error results in Table II (gap roughly 0.6–2.0 dB) support this reading. However, the reverberant-room LSD results in Table IV show substantially larger magnitude/coloration errors for BSM-iMagLS than for BSM-MagLS: 12-mic circular 4.14 vs 3.76 dB; 6-mic semi-circular 5.22 vs 3.82 dB; 5-mic glasses 6.75 vs 3.56 dB; 4-mic EasyCom 6.51 vs 5.08 dB. The 5-mic glasses difference is ~3.2 dB and the 4-mic EasyCom difference is ~1.4 dB, which is difficult to call 'comparable.' Section V-D4 states that BSM-iMagLS 'remains closely aligned' with BSM-MagLS, but the displayed Fig. 7 only covers the EasyCom array in the medium room; Table IV contradicts this for the smaller arrays. Since the abstract does not restrict the magnitude claim to anechoic conditions, the headline claim overreaches. The method still shows robust ILD improvements and listening-test support, so the issue is one of scope and overclaiming rather than a failed method.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes BSM-iMagLS, an extension of binaural signal matching (BSM) with magnitude least-squares (MagLS) that adds a Gammatone-band interaural level difference (ILD) loss and a magnitude-derivative loss to the optimization objective (Eq. 14). The objective is minimized per HRTF/array instance by a feed-forward DNN-based solver initialized from BSM-MagLS coefficients (Section III-B/C). The method is evaluated against BSM-LS, BSM-MagLS, and eMagLS on four head-mounted microphone arrays, using the KU100 and HUTUBS HRTF databases, under anechoic conditions, head-rotation compensation, and simulated reverberant rooms, followed by a MUSHRA listening test. The reported anechoic results show ILD errors reduced from roughly 6-7 dB to 0.7-3.1 dB with a modest magnitude-error penalty, and the listening test shows a significant spatial-quality improvement over BSM-MagLS.","tokens_in":20853,"tokens_out":3175,"duration_ms":33110,"significance":"If properly scoped, the work is a useful contribution to signal-independent binaural rendering for small head-mounted arrays: it directly addresses ILD fidelity, which is a known weakness of BSM and Ambisonics-based rendering at high frequencies. The paper is strong on breadth of validation: four array geometries including measured head-mounted arrays, two HRTF databases, head-rotation compensation, reverberant conditions, and a listening test with non-parametric statistical analysis. The DNN-based per-instance optimization is a practical alternative to classical iterative solvers. The main weakness is that the headline claim of 'comparable magnitude accuracy' is not supported under reverberation, and the ILD gain is to a large extent the direct minimization of the same metric used for evaluation, so the contribution should be framed more carefully.","major_comments":[{"comment":"The abstract claims that BSM-iMagLS maintains 'comparable magnitude accuracy to state-of-the-art solutions' without restricting this to anechoic conditions. This is contradicted by Table IV: under reverberant BRIRs, BSM-iMagLS LSD errors are substantially higher than BSM-MagLS for three of the four arrays, e.g. 6.75 vs 3.56 dB for the 5-mic glasses array and 6.51 vs 5.08 dB for the 4-mic EasyCom array (medium room). Section V-D4 states that BSM-iMagLS 'remains closely aligned' with BSM-MagLS, but Figure 7 shows only the EasyCom array in the medium room, so it does not substantiate that general statement. The magnitude claim should be scoped to the anechoic/NMSE-type metrics, or the reverberant LSD penalty should be acknowledged in the abstract and conclusions.","section":"Abstract / Table IV"},{"comment":"The loss weights lambda=[0.4, 10] are chosen empirically 'based on preliminary experiments' with no sensitivity analysis. Because the loss couples magnitude, derivative, and ILD terms, the reported ILD/magnitude trade-off may depend on the array, HRTF, head-rotation angle, and room condition. Table IV shows that under reverberation the magnitude penalty varies strongly across arrays (e.g., about 0.4 dB LSD increase for the 12-mic array but about 3.2 dB for the 5-mic glasses array). Without reporting how the results vary as a function of lambda, or at least a justification for why one setting generalizes, the robustness claim is not fully supported.","section":"Section V-B5 / Eq. (14)"},{"comment":"The headline ILD improvement is substantially a fitting outcome: the DNN minimizes exactly the D_ILD metric of Eq. (15), and the same metric (Eq. (15), Section V-C.3) is used to report ILD errors. It is therefore not surprising that Table II shows large ILD reductions relative to methods that do not minimize this objective. The manuscript should state this explicitly and rely more heavily on the listening test and on any metric that is not part of the training loss as independent evidence that the ILD gain transfers to perception. The theoretical analysis in Section IV explains conditions for zero ILD error but does not quantitatively predict the observed 4-6 dB improvements, so it does not resolve this circularity concern.","section":"Section III-B and Section V-C"}],"minor_comments":[{"comment":"The sentence 'previous work [10], [10] demonstrated' contains a duplicated citation; one of the bracketed references should be removed or replaced.","section":"Section II-C"},{"comment":"The sentence introducing the narrow-band ILD error says 'given by (15)', but the displayed equation is numbered (27); the cross-reference should be corrected.","section":"Section IV"},{"comment":"The caption reads 'WITH REVERBERANT BRIRS COMPENSATION', which appears to be a typo; 'conditions' or 'analysis' is likely intended.","section":"Table IV caption"},{"comment":"The DNN is trained for 200 iterations with a single run and no reported variance over initializations or optimizer seeds; given that the method is a per-instance nonlinear solver, a brief note on convergence variability or multiple restarts would strengthen reproducibility.","section":"Section V-B5"},{"comment":"The DNN operates with complex-valued weights and activations, but the initialization scheme and the treatment of complex arithmetic in PyTorch are not described; a sentence specifying how complex parameters are initialized and differentiated would help implementation.","section":"Section III-C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid extension of the authors' IWAENC 2024 feasibility study, with a much broader evaluation and a listening test. The main issue for the editor is scope: the abstract and conclusions overclaim magnitude fidelity under reverberation, and the ILD evaluation is partly circular since the same D_ILD is both loss and metric. These are fixable with rewording, a sensitivity analysis, and an explicit acknowledgment of the fitting nature of the ILD gain. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: solid, incremental extension of the same group's BSM-MagLS work. What's new: the derivative-magnitude term, a per-case DNN optimizer, theoretical ILD limits, and the first head-tracked MUSHRA listening test. The anechoic results are convincing: ILD error drops from 6-7 dB to 0.7-3.1 dB across four arrays and 96 HRTFs, with about a 1 dB magnitude penalty. The listening test shows a clear, significant spatial-quality preference over BSM-MagLS, with timbre statistically comparable. That's real evidence.\n\nThe soft spots are real but fixable. The abstract claims 'comparable magnitude accuracy' with no anechoic qualifier, and Table IV contradicts that under reverberation: BSM-iMagLS is about 3.2 dB worse than BSM-MagLS for the 5-mic glasses, and 1.5 dB worse for EasyCom. Section V-D4 says the method 'remains closely aligned' with BSM-MagLS, but that statement is backed only by the EasyCom medium-room figure; the table says otherwise for the smaller arrays. The claim needs scoping or the trade-off needs to be foregrounded.\n\nSecond, part of the ILD gain is simply the minimization of the same D_ILD metric used for evaluation. Not fatal—the listening test independently supports the perceptual benefit—but the numerical ILD numbers should not be treated as a novel prediction, and the Section IV bounds are too coarse to explain the observed 3-6 dB gains. Lambda was chosen without sensitivity analysis, and there is no variance reporting across DNN initializations. Given the consistent wins across HRTFs and arrays, I doubt lambda is knife-edge, but one sensitivity sweep would close the issue. No code is released; the auralizations are a nice touch but don't support full reproduction.\n\nCitation pattern is appropriate; this is a direct continuation of their own conference papers. Bottom line: the core method works, the subjective test is valuable, and the main problem is an overclaim in the abstract plus missing robustness checks. I'd send it to peer review and ask for a revised abstract, a lambda sensitivity check, and a more careful discussion of the reverberant magnitude penalty. Worth a reading-group slot.","headline":"Solid incremental extension of BSM-MagLS with real ILD gains and a useful listening test, but the abstract overclaims magnitude accuracy under reverberation.","tokens_in":21424,"tokens_out":2779,"would_cite":true,"duration_ms":23794,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adding interaural level difference optimization to binaural signal matching cuts ILD errors from roughly 6–7 dB to 0.74–3.10 dB on head-mounted arrays at a cost of about 1 dB in magnitude accuracy.","keywords":["binaural signal matching","interaural level difference","magnitude least squares","head-mounted microphone arrays","binaural reproduction","deep neural network optimization","spatial audio","head-related transfer function"],"falsifier":"Run BSM-iMagLS on the 4-microphone EasyCom array in a room with $T_{60}=0.6$ s and compare ILD error and LSD to BSM-MagLS: if the ILD error rises above the anechoic value of 3.1 dB by more than a few dB, or if the LSD penalty exceeds the reported ~3 dB to a degree that listeners detect timbre changes, the claim of robust ILD improvement with comparable magnitude accuracy would be falsified.","tokens_in":20383,"feed_emoji":"🎧","tokens_out":7865,"duration_ms":59321,"temperature":0.7,"pith_summary":"Binaural audio for AR and VR headsets is usually captured with only a few microphones mounted on glasses or headbands, and these sparse arrays struggle to reproduce the left–right level differences that the brain uses to localize sound. This paper argues that the errors in those level differences, which run 6–7 dB with current methods, can be cut to between 0.7 and 3.1 dB by adding an interaural level difference (ILD) term to the magnitude-least-squares cost function used in binaural signal matching. The joint cost is minimized by a small per-instance neural network that refines the filter coefficients for each head-related transfer function and array geometry. Simulations across four array types, dozens of HRTFs, head rotations, and reverberant rooms, together with a MUSHRA listening test, support the claim that the ILD gain comes with only a small, roughly 1 dB, increase in magnitude error and better perceived spatial quality. If correct, the method gives wearable devices a signal-independent path to accurate spatial audio without requiring many microphones.","feed_headline":"Binaural level errors drop from 6 dB to under 3 dB","feed_subtitle":"Adding a left-right level term to binaural matching sharpens sound localization for head-mounted arrays.","key_machinery":"The load-bearing object is the dissimilarity measure $D_{\\mathrm{iMLS}}$ of Eq. (14), which the paper calls iMagLS. It combines three terms: the standard magnitude-least-squares error for each ear, a first-derivative magnitude term that smooths the magnitude error curve and suppresses audible spectral artifacts, and an ILD error that compares reference and BSM-rendered signals through Gammatone filter bands over the horizontal plane. The ILD term is what forces the left and right ear filters to be optimized jointly rather than independently. The optimization is carried out by a feed-forward MLP that maps an initial coefficient tensor from BSM-MagLS to the refined iMagLS coefficients, using the ADAM optimizer and $D_{\\mathrm{iMLS}}$ as the loss, with the network weights discarded after each per-instance solve.","core_discovery":"On its own terms, the paper establishes that one can jointly optimize magnitude, magnitude derivatives, and interaural level difference in binaural signal matching, and that doing so removes most of the ILD error that plagues few-microphone head-mounted arrays. The proposed BSM-iMagLS minimizes the loss $D_{\\mathrm{iMLS}}$ of Eq. (14), which sums the per-ear MagLS magnitude error, a first-derivative magnitude matching term, and an ILD term computed over Gammatone filter bands from 1.5 to 20 kHz, with weights $\\lambda=[0.4,10]$. A feed-forward MLP solves this non-convex problem by starting from the BSM-MagLS coefficients and iterating on the loss, so the network is not trained across conditions but is fitted per HRTF and array. Across the 12-microphone circular, 6-microphone semi-circular, 5-microphone glasses, and 4-microphone EasyCom arrays, ILD error falls from roughly 6–7 dB with BSM-MagLS to 0.74–3.10 dB with BSM-iMagLS, while magnitude error worsens by about 1 dB and NMSE stays comparable. A theoretical analysis in Section IV shows that zero ILD error is achievable for a single direction when the array steering matrix has a non-empty null space, but becomes impossible to guarantee when the number of directions exceeds the number of microphones, which is why the method is evaluated as an optimization rather than a closed-form solution.","pith_inferences":["The fixed weight vector $\\lambda=[0.4,10]$ is chosen empirically; a sensitivity sweep across arrays and room conditions would reveal whether the ILD-versus-magnitude trade-off is a robust property of the loss or an artifact of the evaluated settings.","The Section IV null-space analysis implies that the benefit should grow with microphone count or array conditioning; a controlled sweep of microphone number on a fixed head-mounted geometry could test that prediction directly.","The reported three-minute per-instance neural solve is too slow for live head-tracked rendering; amortizing the solver over many HRTFs or replacing it with a fast classical optimizer would be needed for real-time use.","The loss optimizes ILD only above 1.5 kHz, leaving interaural time difference and interaural coherence untouched; adding those terms could extend the same joint-optimization approach to low-frequency localization and externalization."],"forward_implications":["Wearable AR/VR devices with four to six microphones can render binaural audio whose ILD error approaches the 1 dB just-noticeable difference, improving horizontal localization and externalization.","Because the DNN refines coefficients per HRTF and array rather than being trained on a corpus, the method transfers to new head-mounted arrays and head-related transfer functions without retraining.","Head-tracked playback, where HRTFs are counter-rotated to keep the scene world-locked, retains most of the ILD benefit, with ILD error remaining near the JND at smaller rotation angles.","The cost is a roughly 1 dB increase in magnitude error and, under reverberation, an LSD increase of up to about 3 dB relative to BSM-MagLS, so the method intentionally trades spectral accuracy for spatial cue fidelity.","The listening test, conducted under the MUSHRA protocol with 14 participants, shows that BSM-iMagLS is rated closer to the reference than BSM-MagLS on spatial quality and statistically indistinguishable from BSM-MagLS on timbre."],"supporting_citations":[{"why":"Establishes the BSM framework and the MSE and MagLS solutions (Eqs. (10)–(11)) that this paper extends.","marker":"[10]"},{"why":"Introduces iMagLS for first-order Ambisonics, the SH-domain predecessor of the proposed ILD-informed loss.","marker":"[14]"},{"why":"Preliminary feasibility study of iMagLS for BSM, which the current paper generalizes with formal optimization and listening tests.","marker":"[15]"},{"why":"Supplies the MagLS coefficient computation (Eq. (11)) used as initialization for the DNN solver.","marker":"[24]"},{"why":"Provides the ILD definition and Gammatone-band evaluation used in the loss and metrics.","marker":"[27]"},{"why":"Defines eMagLS, the end-to-end magnitude least-squares baseline compared in the simulations.","marker":"[36]"},{"why":"Provides the 1 dB ILD just-noticeable-difference threshold used to interpret the error reductions.","marker":"[40]"}],"fun_headline_variants":["ILD errors drop from ~7 dB to under 3 dB with new binaural matching","Head-worn mic arrays get sharper spatial audio via ILD-optimized BSM","Binaural reproduction improved by joint magnitude and ILD optimization","Few-microphone headset arrays achieve precise binaural sound with BSM-iMagLS","DNN solver cuts interaural level errors in binaural signal matching"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The single empirically chosen weighting $\\lambda=[0.4,10]$ in Eq. (14) is assumed to balance magnitude, derivative, and ILD terms equally well for all array geometries, HRTFs, head rotations, and room conditions; if the optimal balance shifts across settings, the reported ILD-versus-magnitude trade-off may not generalize.","fun_headline_variants_meta":{"raw":{"variants":["ILD errors drop from ~7 dB to under 3 dB with new binaural matching","Head-worn mic arrays get sharper spatial audio via ILD-optimized BSM","Binaural reproduction improved by joint magnitude and ILD optimization","Few-microphone headset arrays achieve precise binaural sound with BSM-iMagLS","DNN solver cuts interaural level errors in binaural signal matching"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000266,"raw_usage":{"total_tokens":1687,"prompt_tokens":1098,"completion_tokens":589,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":714,"completion_tokens_details":{"reasoning_tokens":483}},"tokens_in":714,"tokens_out":589,"duration_ms":6103,"temperature":1.0,"reasoning_tokens":483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T00:14:40.556523+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run BSM-iMagLS on the 4-microphone EasyCom array in a room with $T_{60}=0.6$ s and compare ILD error and LSD to BSM-MagLS: if the ILD error rises above the anechoic value of 3.1 dB by more than a few dB, or if the LSD penalty exceeds the reported ~3 dB to a degree that listeners detect timbre changes, the claim of robust ILD improvement with comparable magnitude accuracy would be falsified.","supporting_citations":[{"cited_title":"De- sign and analysis of binaural signal matching with arbitrary microphone arrays and listener head rotations,","cited_arxiv_id":null,"evidence_quote":"Establishes the BSM framework and the MSE and MagLS solutions (Eqs. (10)–(11)) that this paper extends."},{"cited_title":"Imagls: Interaural level difference with magnitude least-squares loss for optimized first- order head-related transfer function,","cited_arxiv_id":null,"evidence_quote":"Introduces iMagLS for first-order Ambisonics, the SH-domain predecessor of the proposed ILD-informed loss."},{"cited_title":"Feasibility of iMagLS-BSM-ILD informed binaural signal matching with arbitrary microphone arrays,","cited_arxiv_id":null,"evidence_quote":"Preliminary feasibility study of iMagLS for BSM, which the current paper generalizes with formal optimization and listening tests."},{"cited_title":"Six-degrees-of-freedom binaural reproduction of head-worn microphone array capture,","cited_arxiv_id":null,"evidence_quote":"Supplies the MagLS coefficient computation (Eq. (11)) used as initialization for the DNN solver."},{"cited_title":"Xie, Head-related transfer function and virtual auditory display","cited_arxiv_id":null,"evidence_quote":"Provides the ILD definition and Gammatone-band evaluation used in the loss and metrics."},{"cited_title":"End-to-end magnitude least squares binaural rendering of spherical microphone array signals,","cited_arxiv_id":null,"evidence_quote":"Defines eMagLS, the end-to-end magnitude least-squares baseline compared in the simulations."},{"cited_title":"Discrimination of interaural differences of level as a function of frequency,","cited_arxiv_id":null,"evidence_quote":"Provides the 1 dB ILD just-noticeable-difference threshold used to interpret the error reductions."}],"review_version":1}