{"id":"f2640e0a-8696-480c-b49e-0d452e14d82c","arxiv_id":"2509.00400","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of deep learning for personalized binaural audio, covering explicit HRTF prediction and end-to-end synthesis, datasets, metrics, and open challenges.","lead":"This survey reviews deep learning methods for personalized binaural audio reproduction, grouped into explicit HRTF modeling and direct end-to-end synthesis. It maps the current landscape of datasets, metrics, and applications, useful for researchers entering this fast-moving field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coverage claim unverifiable: no literature search protocol or date cutoff underlies the asserted gap-filling overview.","rationale":"The reader's weakest assumption identifies the undocumented literature selection as the central weakness. I agree with that assessment: for a survey, the most load-bearing requirement is that the reviewed set fairly represents the field, because all downstream claims about open challenges, trends, and 'filling a gap' depend on it. The paper provides no methodology section, no search protocol, and no inclusion criteria, making the coverage claim unfalsifiable as written. I also note a secondary concern about scope mismatch—Section III covers generic binaural synthesis that is not evidently listener-personalized, despite the survey's title—but that is less central than the selection issue. My recommendation remains UNCHANGED because the reader already accepted the survey with this limitation. The concrete test (a systematic literature comparison) would settle whether the selection is actually biased or incomplete; if the test passes, the survey's value as an organizing reference stands; if it fails, the central claim would need to be revised. I do not see grounds to reject or make acceptance conditional beyond what the reader already stated, as the survey is internally consistent and its taxonomy is reasonable.","tokens_in":39987,"tokens_out":6867,"duration_ms":83535,"concrete_test":"Conduct a PRISMA-style systematic search over arXiv, IEEE Xplore, ACM DL, and ISCA archive for 2018–2025 using keywords: 'binaural synthesis', 'binaural rendering', 'spatial audio deep learning', 'HRTF personalization'. Compare the resulting reference set against the survey's bibliography. Criteria for a weakened claim: (a) identify any prior survey covering both explicit HRTF personalization and end-to-end binaural synthesis, or (b) find that >10% of papers from top venues (ICASSP, Interspeech, WASPAA, NeurIPS, CVPR) in the relevant area are absent from the survey. If either criterion holds, the 'fills a gap' and 'first comprehensive overview' assertions require explicit qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central claim, 'fills a gap... one of the first comprehensive overview of DL-based end-to-end spatial audio synthesis' (Section I-C), rests on an undocumented literature selection. No search strategy, database list, inclusion/exclusion criteria, or date range is stated. The reference list is a convenience sample: it heavily cites works by the authors' own research network (e.g., [50,55,76,89] from the ECNU/IOA group) and a large number of 2024–2025 arXiv preprints, but it does not substantiate that this set fairly represents the field. Without a protocol, the 'recent advances' boundary is arbitrary, and the asserted gap may already be filled by a prior survey (the paper only cites four related surveys, none from 2023–2025 that focus specifically on binaural synthesis). If a systematic search reveals a missing prior survey or a significant omission of core papers from top venues (Interspeech, ICASSP, WASPAA), the survey's contribution as the 'first comprehensive' overview weakens materially, and the 'open challenges' discussion could mislead by reflecting only a selected subset of the literature.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of deep learning for personalized binaural audio reproduction. It organizes recent work into two paradigms: explicit HRTF-based personalization/filtering and end-to-end binaural synthesis. The survey covers HRTF data representations, personalization from morphological and environmental cues, spatial interpolation, dataset fusion, and evaluation metrics; it also reviews end-to-end synthesis from audio, visual, text, and joint multimodal inputs, along with relevant datasets, applications, and open challenges. The paper claims to fill a gap in the literature by being one of the first comprehensive overviews covering both pathways in a unified structure.","tokens_in":40230,"tokens_out":5272,"duration_ms":60410,"significance":"If the coverage is accepted, the survey provides a useful organizing reference for a rapidly growing field. Its strengths include a clear taxonomy separating explicit and end-to-end paradigms, extensive summary tables for methods, datasets, and metrics, and a number of useful code links. Spot-checks of the metric formulas in Tables IV and VIII indicate they are consistent with standard usage, and the textual descriptions of methods align with the cited literature. The paper does not introduce new derivations, so internal consistency is high. However, the central claim of being a comprehensive, gap-filling overview rests on an undocumented literature selection, which limits confidence in the representativeness of the reviewed set and in the conclusions derived from it.","major_comments":[{"comment":"The claimed contribution, 'one of the first comprehensive overview of DL-based end-to-end spatial audio synthesis', is not supported by a reproducible literature selection. No search strategy, database list, inclusion/exclusion criteria, date cutoff, or screening process is reported. The reference set appears to be a convenience sample with many 2024-2025 arXiv preprints and a notable number of works from the authors' own network (e.g., [55], [76], [89]); those works are summarized like all others, but without a protocol the reader cannot verify that the selected set fairly represents the field or that a prior survey focused on binaural synthesis does not already fill part of the claimed gap. This is load-bearing for the survey's contribution. Please add a methodology subsection describing the search and selection protocol, and use it to substantiate the gap assertion by comparing with t","section":"Section I-C"},{"comment":"The notation fw : (θ, φ) → H(θ, φ, f) is ambiguous: if f is not an input, then H(θ, φ, f) must be a vector-valued output over frequency, but this is not stated. Clarify whether the network outputs a full spectrum or whether frequency is also an input coordinate. This is a presentation issue, but it affects the interpretation of the continuous-domain interpolation methods summarized in Table II.","section":"Section II-D, Eq. (1)"}],"minor_comments":[{"comment":"The phrase 'one of the first comprehensive overview' mixes singular and plural; suggest 'one of the first comprehensive overviews'.","section":"Section I-C"},{"comment":"The text refers to 'SonicSet [202]' in the paragraph after Table VII, but the table row and reference [202] are for 'SonicSim'. Please align the naming.","section":"Section III-D"},{"comment":"In Table IV, the LMD formula is rendered as '20 log10 |Hhat/H|' and described as 'mean absolute log-magnitude difference'; if the metric is the mean absolute difference, the outer absolute value should appear explicitly around the log term. Clarify the typesetting.","section":"Section II-F"},{"comment":"The open challenges in Section V-A through V-E are reasonable but somewhat generic; tying each challenge back to specific gaps visible in Tables I-VIII would strengthen the survey's contribution.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The survey is well organized and likely useful, but the absence of a literature-search methodology is a substantive weakness for a paper whose main claim is comprehensiveness. I recommend requiring the authors to add a transparent selection protocol and to temper or defend the 'first comprehensive overview' claim. No issues of scholarly integrity are evident; the self-citations are summarized normally, though the selection's network bias should be addressed through the methodology."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis survey is worth your time. It does something real: it organizes the field into two paradigms—explicit HRTF personalization and end-to-end synthesis—and the end-to-end half covers ground that prior HRTF-focused surveys left open, especially multimodal synthesis from visual, text, and parametric inputs. The tables of datasets and metrics are handy; I spot-checked the formulas in Tables IV and VIII and they're correct. For a newcomer this is one of the best maps of the area I've seen, and it's precisely the kind of reference that saves hours of literature hunting.\n\nThe main soft spot is the one the stress-test flags, and it's legitimate: there is no documented search strategy, inclusion criteria, or date cutoff anywhere in Section I-C. The reference list leans on 2024–2025 arXiv preprints and on a few works from the authors' own network (e.g., [55], [76], [89]). That doesn't make the survey wrong, but the claim to be 'one of the first comprehensive overview[s]' of end-to-end binaural synthesis is unverifiable as stated. A skeptical reader can't tell whether the selection fairly represents the field or whether a systematically searched review would have looked different. This is a real weakness, though not a fatal one: most surveys of this kind are narrative rather than systematic, and the value here is in the organizing framework and the dense summaries, not in bibliometric exhaustiveness.\n\nThe other thing I'd push back on is the conclusion's 'fundamentally reshaping' language. The body supports 'significant transformation,' not 'fundamental reshaping.' It's a sweeping tone that the evidence doesn't need.\n\nIn proportion: the central argument holds up. The taxonomy is legitimate, the summaries match standard usage, and the open-challenges section is reasonable. None of the soft spots break the survey's usefulness. If I ran a journal, I would send this to review—it deserves referee time—with a request for a short paragraph on how references were collected and a softer framing of the novelty claims. That's a revision, not a rejection.\n\nRecommendation: engage with it. Cite it when you need a broad pointer into the field, and push for the coverage protocol to be made explicit. It's a good paper that would become a better one with a little methodological transparency.","headline":"A solid, genuinely useful survey that organizes DL binaural audio into explicit HRTF personalization and end-to-end synthesis; the gap claim is plausible but rests on an undocumented literature selection, so treat the 'first comprehensive' language with caution.","tokens_in":40694,"tokens_out":1916,"would_cite":true,"duration_ms":26311,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that deep learning is transforming personalized binaural audio through two pathways—explicit HRTF filtering and end-to-end synthesis—and organizes the field around that split.","keywords":["binaural audio","head-related transfer function","HRTF personalization","end-to-end binaural synthesis","multi-modal spatial audio","implicit neural representations","spatial audio evaluation","deep learning survey"],"falsifier":"A reader could run a structured bibliographic search for deep-learning binaural audio papers from 2019-2025 with explicit inclusion criteria. If a substantial number of well-cited methods fall outside both paradigms, or if important datasets and metrics are absent from the tables, the coverage claim would be falsified.","tokens_in":39922,"feed_emoji":"🎧","tokens_out":4677,"duration_ms":50552,"temperature":0.7,"pith_summary":"This survey claims that deep learning now drives personalized binaural audio through two complementary paradigms: explicit personalized filtering, where a model predicts a listener-specific head-related transfer function (HRTF) and then renders with it, and end-to-end rendering, where a model maps source audio plus visual, text, or parametric guidance directly to binaural signals. The paper organizes recent methods into this split, covering HRTF prediction from anthropometry, photos, and 3D scans; interpolation and dataset fusion; and single- or multi-modal binaural synthesis. It also catalogs public datasets, evaluation metrics, applications, and open challenges. A sympathetic reader would take the central contribution to be the organizing structure itself: a field map that lets new work be positioned and compared, rather than any single technical result.","feed_headline":"Deep learning is remaking binaural audio through two pathways","feed_subtitle":"An overview that pairs HRTF personalization with end-to-end synthesis, plus the datasets and metrics that make them comparable.","key_machinery":"The organizing device is a two-paradigm taxonomy. Paradigm one, explicit personalized filtering, keeps the classic rendering pipeline: a deep model predicts the head-related transfer function (the filter that encodes how a listener's head, torso, and pinna shape sound directionally) and then convolution renders binaural audio. Paradigm two, end-to-end rendering, discards the explicit filter and learns the whole mapping from source audio plus guidance (visual, textual, or parametric) to binaural output. Within this split, the survey uses data representation (time-domain, frequency-domain, sparse, and implicit neural representations), datasets, and metrics as secondary organizing axes.","core_discovery":"The central claim is that deep learning is not just improving but reshaping both fundamental pathways to binaural audio. On the explicit side, models predict personalized HRTFs from sparse measurements, morphological features, 3D geometry, or even ambient listening cues, replacing costly lab measurement. On the end-to-end side, models synthesize binaural audio directly from mono audio, video, text, or scene geometry, learning personalization implicitly. The survey's thesis is that these two paradigms, plus shared datasets and metrics, now define the field, and that this is the first structured overview to treat end-to-end synthesis as a first-class area alongside HRTF modeling.","pith_inferences":["The authors do not say this, but the taxonomy suggests the two paradigms are converging: end-to-end models learn implicit HRTFs, while explicit methods are adopting neural fields, so the split may soon be a matter of rendering control rather than architecture.","An inference beyond the survey: if in-the-wild and text-guided personalization mature, personalized binaural audio could be generated on demand from a photo of a pinna or a sentence describing a scene, without any lab measurement.","The survey's emphasis on dataset fusion implies that the next large gains may come from data, not new architectures; readers should expect cross-database training to become standard.","Because the review is narrative rather than systematic, quantitative claims about trends should be read as illustrative until a formal literature search confirms the coverage."],"forward_implications":["Researchers can position any new method as either explicit HRTF personalization or end-to-end synthesis, and compare it against the tables of representative approaches.","End-to-end binaural synthesis is a recognized first-class research area, with text- and image-guided generation as active directions.","Implicit neural representations appear to be the leading mechanism for continuous HRTF interpolation and cross-dataset fusion, easing the data heterogeneity bottleneck.","Perceptual validity, not raw signal error, is the field's main evaluation gap; objective metrics need stronger correlation with listening tests.","Applications in VR/AR, hearing aids, telepresence, audio-visual navigation, and scene understanding follow directly from the two paradigms' progress."],"supporting_citations":[{"why":"Foundational end-to-end mono-to-binaural speech synthesis model; anchors the single-modal synthesis paradigm.","marker":"[38]"},{"why":"Earlier survey on machine learning for HRTF personalization; the paper distinguishes its broader scope from this one.","marker":"[39]"},{"why":"A review of HRTF generation methods; marks the existing coverage gap the paper aims to fill.","marker":"[40]"},{"why":"Survey of machine learning techniques for HRTF individualization; positioning reference for the explicit personalization section.","marker":"[41]"},{"why":"Broader overview of data-based spatial audio capture and reproduction; boundary against which the paper defines its narrower focus.","marker":"[42]"},{"why":"Implicit neural representation with periodic activations; underpins the continuous-domain HRTF interpolation and dataset fusion discussion.","marker":"[91]"},{"why":"Self-supervised spatial audio generation from 360-degree video; early basis for visual-guided synthesis.","marker":"[147]"},{"why":"Visual sound method that combines video with audio; core reference for visual-guided binaural synthesis.","marker":"[148]"},{"why":"Introduces text-guided audio spatialization and its benchmark; defines the text-guided synthesis category.","marker":"[155]"},{"why":"Large-scale language-driven spatial audio dataset and model; anchors the joint multi-modal and text-guided directions.","marker":"[160]"}],"fun_headline_variants":["Two deep-learning paths now define binaural audio","Deep learning personalizes binaural audio via two routes","Binaural audio’s deep-learning split: HRTF vs end-to-end","AI picks your ears: two routes to binaural sound","Two deep-learning approaches reshape binaural audio"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The survey's usefulness depends on its reviewed set of methods, datasets, and metrics fairly representing the whole field, but the paper gives no systematic search strategy or inclusion criteria.","fun_headline_variants_meta":{"raw":{"variants":["Two deep-learning paths now define binaural audio","Deep learning personalizes binaural audio via two routes","Binaural audio’s deep-learning split: HRTF vs end-to-end","AI picks your ears: two routes to binaural sound","Two deep-learning approaches reshape binaural audio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001117,"raw_usage":{"total_tokens":4451,"prompt_tokens":669,"completion_tokens":3782,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":413,"completion_tokens_details":{"reasoning_tokens":3698}},"tokens_in":413,"tokens_out":3782,"duration_ms":29743,"temperature":1.0,"reasoning_tokens":3698,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:36:08.621210+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could run a structured bibliographic search for deep-learning binaural audio papers from 2019-2025 with explicit inclusion criteria. If a substantial number of well-cited methods fall outside both paradigms, or if important datasets and metrics are absent from the tables, the coverage claim would be falsified.","supporting_citations":[{"cited_title":"AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis","cited_arxiv_id":"2503.12806","evidence_quote":"Introduces text-guided audio spatialization and its benchmark; defines the text-guided synthesis category."}],"review_version":1}