{"id":"97d6e3fb-5af7-4131-975b-9a6723e7bf46","arxiv_id":"2508.03724","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":0.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive literature review of audio-visual segmentation, covering problem setup, datasets, methods, training paradigms, and future directions.","lead":"This paper surveys audio-visual segmentation, the task of finding and outlining the object in a video that makes the sound we hear. A smart generalist could read it as a map of the field's datasets, methods, benchmarks, and open problems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No specific technical objection identified: the survey's central claim cannot be checked because the submitted body text is empty, so UNVERDICTED remains the appropriate verdict.","rationale":"The reader's assessment was based only on the abstract because no full text was provided, and the same limitation applies to this stress-test pass. I agree that the central claim's validity depends entirely on the accuracy and completeness of the survey's literature selection and benchmark comparisons. Since the full text is empty, there is no basis to accept or reject the survey's claims; UNVERDICTED is the only epistemically honest stance. I found no self-contained inconsistency in the abstract itself, and there are no equations, tables, or citations to scrutinize. The proposed concrete test (a citation-to-table cross-check) would settle the concern once the manuscript is available, because it directly tests the 'extensive comparison across standard benchmarks' assertion. No further technical concern can be responsibly raised without fabricating content to critique, which would violate the good-faith review posture.","tokens_in":655,"tokens_out":1727,"duration_ms":22945,"concrete_test":"Retrieve the full manuscript and run a cross-check of the main benchmark comparison table: for every row, locate the cited source paper and confirm that the reported metric uses the same dataset split, evaluation protocol, and metric definition as the table header; flag any row whose value cannot be reproduced from the cited source. If any row is inconsistent, the 'extensive comparison' claim weakens proportionally.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim is that the paper presents a comprehensive overview with an extensive, accurate comparison of AVS methods across standard benchmarks. For that claim to hold, the full text must contain representative literature coverage and benchmark tables computed under consistent evaluation protocols. The provided manuscript contains no body text whatsoever, so no internal consistency check, citation check, or numerical verification is possible. This is a verification gap rather than a demonstrated error: I cannot identify a specific mathematical or empirical flaw in the argument because the argument is not present in the available material. The honest assessment is that no significant objection can be raised on the available evidence, but the survey's value is entirely contingent on the accuracy and completeness of its retrieved full text.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript, as submitted, consists solely of an abstract and an empty full-text body. The abstract claims to present a comprehensive survey of audio-visual segmentation (AVS), covering problem formulation, benchmark datasets, evaluation metrics, methodological progression, training paradigms, and benchmark comparisons, and to close with challenges and future directions. Because no body text, figures, tables, or references are provided, the claims in the abstract cannot be checked or assessed.","tokens_in":762,"tokens_out":1919,"duration_ms":23511,"significance":"If the full survey were present and the content matched the abstract's claims, the paper would serve as a potentially valuable reference map for the growing AVS community, consolidating method taxonomies, benchmarking protocols, and open problems. The abstract indicates a clear scope and an organized structure, and the promised extensive comparison across standard benchmarks would be a practical contribution. However, with the full text absent, the current submission has no verifiable content, and thus its significance is entirely contingent on the completion and accuracy of the missing material.","major_comments":[{"comment":"The submitted manuscript contains no body text whatsoever; it is only an abstract. The central claim of the paper—that it provides a 'comprehensive overview' and an 'extensive comparison' of AVS methods—cannot be verified, replicated, or even located in the submission. There are no sections describing problem formulation, datasets, evaluation metrics, methodology taxonomy, benchmark tables, or discussions of training paradigms, despite each being promised in the abstract. This is a load-bearing omission: a survey paper's value resides entirely in its detailed exposition and evidence, none of which is present. The authors must provide the full manuscript before any substantive review can proceed.","section":"Full Text (entire body)"},{"comment":"The abstract asserts that the paper includes 'an extensive comparison of AVS methods across standard benchmarks, highlighting the impact of different architectural choices, fusion strategies, and training paradigms.' In the absence of the corresponding tables, experimental protocols, and cited sources, this assertion is unsupported. No benchmark names, metric definitions, or numerical results are given in the available material, so the reader cannot judge whether the comparison is fair, consistent, or representative.","section":"Abstract, 'extensive comparison'"}],"minor_comments":[{"comment":"The phrase 'selfand weakly supervised learning' contains a typographical error; it should be 'self-supervised and weakly supervised learning' or the like.","section":"Abstract, line on learning paradigms"},{"comment":"The title promises a journey 'From Waveforms to Pixels,' but the abstract does not discuss waveform-level audio representations or how they are processed prior to visual alignment. If the full text covers audio feature extraction from waveforms, the abstract should mention it for coherence.","section":"Title and abstract"}],"recommendation":"major_revision","confidential_remarks":"To the editor: This submission is an arXiv stub containing only an abstract. It is not yet a reviewable manuscript. I recommend returning it to the authors for completion of the full text before sending it out for further review. If this is a pipeline artifact, please disregard this report and process the actual manuscript. Otherwise, the major revision should be conditioned on submission of the complete paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is simple: the submitted manuscript is an abstract and nothing else. The full-text section is empty, so there is no survey to evaluate. That determines my verdict.\n\nWhat the abstract does well: it sketches a sensible map of the audio-visual segmentation field. The categories it lists—problem formulation, datasets, metrics, encoding and fusion architectures, decoder designs, training paradigms, and open challenges—are the right ones for a survey, and the mention of weakly supervised and training-free methods shows the authors know where the field is moving. If the actual paper delivers on that outline, it could be a useful reference for newcomers.\n\nBut the body text is missing, and that is a load-bearing problem, not a minor one. A survey's value lives in its literature coverage and benchmark comparisons. Here I cannot check whether the selected methods are representative, whether the benchmark tables are computed under consistent evaluation protocols, or whether the characterizations of prior work are accurate. The stress-test note is correct that no specific technical objection can be raised against an empty argument, but the flip side is that the central claim of a \"comprehensive overview\" is entirely unsupported. The abstract promises an \"extensive comparison\" of methods across standard benchmarks; none of that comparison is present.\n\nI also cannot assess the citation pattern, because there are no citations to inspect. That is not a small thing for a survey; fairness to prior work is the whole game, and there is no evidence of it here.\n\nBottom line: this submission should be desk rejected, not because the topic is unworthy but because there is no paper. If the authors accidentally submitted only the abstract, the fix is obvious. I would welcome a resubmission with the full text; a well-done AVS survey deserves referee time. But I cannot send an empty manuscript to referees, and I would not cite it in its current state.","headline":"As submitted, this is an abstract with no body text, so the survey's central claims cannot be checked and the paper is not ready for peer review.","tokens_in":1172,"tokens_out":2051,"would_cite":false,"duration_ms":25371,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Locating and outlining every object that is making the sound in a video has become a research field large enough for a single systematic survey, and this paper is that survey.","keywords":["audio-visual segmentation","multimodal perception","video object segmentation","audio-visual fusion","sound source localization","weakly supervised learning","training-free methods","benchmark comparison"],"falsifier":"Collect the reported results for one standard benchmark and check each method's evaluation protocol—exact split, metric definition, post-processing, and whether audio is used at test time. If the numbers mix incompatible protocols, or if a spot re-run of the top methods changes their ordering, the survey's comparative conclusions fail.","tokens_in":500,"feed_emoji":"🎵","tokens_out":5662,"duration_ms":63147,"temperature":0.7,"pith_summary":"Audio-visual segmentation (AVS) is the task of outlining, frame by frame, the objects in a video that are producing the sound being heard. This paper argues that the field has grown enough to be surveyed as a whole, and it offers a single reference that connects the problem's definition, its benchmark datasets and evaluation metrics, and the progression of methods built to solve it. The survey's value is organizational: a researcher can see how encoding, fusion, and decoder designs have evolved, how training has moved from fully supervised to weakly supervised and training-free regimes, and which choices currently win on standard benchmarks. It also names the field's bottlenecks—limited temporal modeling, a bias toward visual evidence, weak robustness in complex scenes, and high computation—and points to future directions such as better temporal reasoning, foundation-model generalization, and less dependence on labeled data. A sympathetic reader would take the paper as a claim that AVS is now ripe for systematic comparison and that its next advances will come from the directions the survey highlights.","feed_headline":"One survey maps how machines find what makes a sound","feed_subtitle":"It organizes datasets, metrics, fusion designs, and training paradigms, then compares methods on standard benchmarks.","key_machinery":"The central object is the audio-visual segmentation task itself, defined as the pixel-level delineation of every object in a video that is producing the sound in the accompanying audio track. The survey's machinery is a three-axis taxonomy of methods: the choice of unimodal versus multimodal encoding, the strategy used to fuse audio and visual features, and the design of the decoder that turns fused features into segmentation masks. A fourth axis crosses these with training paradigms, ranging from fully supervised through weakly supervised to training-free setups. Benchmark datasets and evaluation metrics do the measuring work: they are what allow the survey to compare methods and to attribute performance differences to architectural choices, fusion strategies, and training regimes.","core_discovery":"The central claim of this survey is that audio-visual segmentation has become a distinct, mature research area whose entire arc—from problem formulation through datasets, metrics, architectures, fusion strategies, decoder designs, and training paradigms—can be laid out in a single coherent narrative. On the paper's own terms, the discovery is taxonomic: methods are not a scattered set of tricks but a progression along identifiable axes, and their relative standing can be judged through extensive comparison across standard benchmarks. The survey further asserts that the field's current limits are concentrated in specific, named weaknesses, and that the remedies are already visible in weakly supervised and training-free approaches, foundation-model generalization, and higher-level reasoning. This is a reference-map claim rather than a new algorithm, and its force lies in how completely and fairly it organizes what has been done.","pith_inferences":["A testable extension follows from the reported vision bias: design audio-first or audio-only probes that attempt segmentation without visual features; if they recover significant structure, current methods may be leaning on vision more than the fusion narrative suggests.","The survey's emphasis on fusion implies that a deliberately simple, parameter-free fusion baseline—such as broadcasting audio features to every pixel and concatenating—would be a cheap way to expose how much of the reported gains come from architecture sophistication rather than training or data.","The future direction of higher-level reasoning suggests that AVS will likely converge with audio-visual language grounding; one concrete prediction is that models trained with language or event-level supervision will outperform pixel-level fusion methods on rare or novel sound sources.","Because the survey covers training-free methods, an immediate benchmark extension would be to test foundation-model segmenters prompted by audio-derived text descriptions; the survey's comparison framework could be reused to score such zero-shot pipelines."],"forward_implications":["A newcomer to AVS can use the problem formulation, dataset descriptions, and metric discussion as a direct entry point for designing a first method or benchmark study.","Because the survey compares methods under standard benchmarks, future work can pick the strongest encoding, fusion, and decoder configurations as baselines rather than re-deriving them.","The survey's challenge list—temporal modeling, vision bias, robustness, computation—effectively defines a near-term research agenda for the field.","The progression from fully supervised to weakly supervised and training-free methods indicates that reducing reliance on labeled data is a live and promising direction.","If the identified limitations are accurate, approaches that strengthen temporal reasoning and audio-visual fusion in complex scenes should expect the largest performance gains."],"supporting_citations":[],"fun_headline_variants":["Survey charts audio-visual segmentation from inputs to benchmarks","How machines link sounds to pixels: a field survey","Audio-visual segmentation: one map of methods and gaps","The full arc of audio-visual segmentation, in one survey","From waveforms to pixels: a research map"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's value rests on its literature selection and benchmark numbers being accurate, representative, and compared under consistent conditions; if key methods are missing or scores come from different evaluation protocols, the field map and its rankings mislead.","fun_headline_variants_meta":{"raw":{"variants":["Survey charts audio-visual segmentation from inputs to benchmarks","How machines link sounds to pixels: a field survey","Audio-visual segmentation: one map of methods and gaps","The full arc of audio-visual segmentation, in one survey","From waveforms to pixels: a research map"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1513,"prompt_tokens":914,"completion_tokens":599,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":522}},"tokens_in":530,"tokens_out":599,"duration_ms":6503,"temperature":1.0,"reasoning_tokens":522,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:52:15.919639+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect the reported results for one standard benchmark and check each method's evaluation protocol—exact split, metric definition, post-processing, and whether audio is used at test time. If the numbers mix incompatible protocols, or if a spot re-run of the top methods changes their ordering, the survey's comparative conclusions fail.","supporting_citations":[],"review_version":1}