{"id":"6489875a-9812-4d3b-9200-82cbd0206a49","arxiv_id":"2506.02499","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"DnR-nonverbal moves non-verbal vocal sounds into the speech stem for cinematic audio source separation, fixing the misallocation of laughter and screams to the effect stem.","lead":"The authors introduce DnR-nonverbal, a dataset for cinematic audio source separation that adds non-verbal vocal sounds such as laughter and screams into the speech stem. Adding this data to training makes a state-of-the-art separator keep these voices in the speech stem, improving scores on synthetic mixes and listener preference on real movie clips.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training-label contamination is the load-bearing risk: the 395 eval clips are tag-filtered, not acoustically verified, so part of the Table 2 gains may reflect learning the same contaminated speech labels rather than clean non-verbal vocal separation.","rationale":"The reader's weakest_assumption already identifies curation quality as the key risk, and I agree that it is the most load-bearing issue. Tables 1 and 2 establish that the baseline model tends to put non-verbal sounds into the effect stem and that adding the dataset moves them into the speech stem, but both tables depend entirely on the correctness of the eval labels. If those labels are contaminated, the reported SDR deltas do not reliably measure real separation quality. The paper's own Section 4.4 observation that the model mistakes an animal's voice for screaming is direct evidence that some contamination has been internalized. I am not choosing the data-size confound or the small A/B test as the main attack, because the causal claim survives or fails primarily on label purity: even a perfectly sized and statistically polished experiment would be misleading if the speech-stem labels are systematically noisy. My concrete check is a manual audit plus recomputation, which would settle whether the improvement persists on cleanly labeled data. This is essentially the same condition the reader set, so I keep the verdict unchanged: the dataset is a reasonable contribution, but acceptance should require a clean-label audit and re-computed SDR results, along with reported contamination rates.","tokens_in":7834,"tokens_out":10755,"duration_ms":113724,"concrete_test":"Manually audit all 395 evaluation clips (or a stratified random sample with two annotators) for the presence of non-human sounds, music, or noise; remove or relabel every such clip. Recompute Table 2 on the clean remainder. If the speech SDR gain over DnR-v2 drops materially (e.g., below about 7 dB), contamination is inflating the result; if the gain persists, the concern does not land. As a secondary check, perform the same audit on a random sample of training clips to estimate the contamination rate and to verify whether train/eval contamination overlaps.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is causal: placing non-verbal vocal sounds in the speech stem fixes their misallocation to the effect stem. The load-bearing condition is that the clips labeled as non-verbal speech are actually human non-verbal vocalizations and not animal sounds, music, or noise. Section 3.3 admits this condition is not met: 'Despite LLM-based filtering, some clips still contained non-human sounds.' The LLM filter on FreeSound sees only tags and description, not audio, and FSD50K tags are weak, non-exhaustive labels. The evaluation set (395 FSD50K clips) passes the same tag-based filters, so it can contain the same contamination. Section 4.4 then confirms that the model trained with DnR-nonverbal mistakes an animal's voice for screaming in real movie audio, i.e., it has learned a contaminated mapping. If contaminated clips appear in both training and evaluation, the Table 2 speech SDR improvement (5.62 to 9.30 dB) and effect SDR improvement (2.54 to 5.23 dB) may be inflated by matching label noise rather than by genuinely separating human laughter, screams, and sighs from animal or environmental sounds. The dataset may still be useful, but the headline claim needs a clean-label evaluation to rule out this mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DnR-nonverbal, a new CASS dataset that extends the DnR-v2 speech stem with non-verbal human vocal sounds (laughter, screaming, whispering, crying, sighing, shouting). Clips are collected from FSD50K and new FreeSound crawls, filtered by rule-based tag checks and an LLM (GPT-4o) that inspects tags and descriptions rather than audio, then mixed into 60-second speech stems together with reading-style speech. The authors train BandIt on DnR-v2 and on DnR-v2 + DnR-nonverbal, report substantially higher speech and effect SDR on a synthetic evaluation set, and run a small A/B test on real movie clips showing 76.9% preference for the retrained model. The central claim is that placing non-verbal sounds in the speech stem fixes the misallocation of these sounds to the effect stem.","tokens_in":8138,"tokens_out":4265,"duration_ms":39738,"significance":"If the central claim holds, DnR-nonverbal is a useful, focused contribution to cinematic audio source separation: it targets a real domain gap that existing datasets (DnR-v2, DnR-v3, and speech-music datasets) ignore, and it provides a public dataset with a clear baseline comparison. The paper deserves credit for releasing the dataset on Zenodo, for including both objective and subjective evaluations on real movie audio, and for transparently acknowledging limitations in Sections 3.3 and 4.4. However, the main empirical conclusion rests on the assumption that the tag- and LLM-filtered clips are actually human non-verbal vocalizations, and the paper itself admits this assumption is imperfect; this makes the clean-label evaluation a load-bearing point that needs strengthening.","major_comments":[{"comment":"The central claim is causal: adding non-verbal human vocal sounds to the speech stem teaches the model to extract these sounds as speech rather than effect. The paper admits in §3.3 that 'Despite LLM-based filtering, some clips still contained non-human sounds', and in §4.4 that the model can mistake an animal's voice for screaming. The evaluation set (395 FSD50K clips) is selected by the same tag-based procedure (Human voice descendant tags minus Singing) and is not acoustically verified. Since both training and evaluation pass through the same contaminated pipeline, part of the measured improvement in Table 2 (speech SDR 5.62 to 9.30 dB, effect SDR 2.54 to 5.23 dB) may reflect the model learning a mapping that assigns non-human sounds to the speech stem, rather than genuinely isolating human laughter, screams, and sighs. I request a clean-label validation: manually verify the evaluation clips (or a representative subset), remove or re-label non-human sounds, and recompute Table 2 on the verified subset. The paper should also state how many eval clips were excluded and whether the SDR gains remain after this correction.","section":"§3.3, §4.4, Table 2"},{"comment":"The SDR metric in Eq. (3) is a simplified energy ratio without the usual signal-to-distortion decomposition or scale-invariant projection of BSSEval-style SDR. The paper reports single-run numbers with no confidence intervals, error bars, or significance tests across the 100 evaluation mixtures. Since the reported gains are large, the conclusions would likely survive a statistical test, but the paper should either report standard SDR with projection or provide significance tests (e.g., paired tests across mixtures, or variance across multiple training runs) to support the claim that the improvement is not due to a few outlier clips.","section":"§4.3, Eq. (3)"},{"comment":"The LLM-based filtering for FreeSound clips relies on tags and descriptions only; the prompt itself states the LLM is 'determin[ing] the availability by guessing the given tags and description', not by listening to the audio. This means the LLM cannot detect acoustic contamination such as animal sounds, music, or noise that are not reflected in the metadata. Since the paper acknowledges residual contamination, the evaluation set should not be presumed clean without audio-level verification, and the manuscript should explicitly state this limitation in the evaluation section rather than only in the filtering section.","section":"§3.3"}],"minor_comments":[{"comment":"The caption contains a typo: 'Comparision' should be 'Comparison'.","section":"Figure 1 caption"},{"comment":"The A/B test uses 20 clips and 13 raters; the paper reports only the aggregate preference percentages without per-clip breakdown or inter-rater agreement. Adding a confidence interval or a per-clip best-worst count would strengthen the external support.","section":"§4.4, Table 3"},{"comment":"Reference [15] misspells the first author's name as 'Warcharasupat' (should be 'Watcharasupat'), and reference [17] reads 'Proceedings of Proceedings of Interspeech'; these should be corrected.","section":"Bibliography"},{"comment":"The sentence 'The model may treat neither reading-speech nor music content as effects' is unclear and should be rephrased; the intended meaning appears to be that the model assigns an unseen sound class to the effect stem by default.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is squarely within the scope of a CASS/dataset-focused venue, and the dataset release is valuable. My main concern is the same one the authors themselves raise: the lack of acoustic verification of the curation pipeline and, consequently, of the evaluation set. This is not a fatal flaw, but it needs to be addressed with a clean-label evaluation before the headline claim can be considered established. I would be willing to look at a revised version that includes such validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The DnR-nonverbal paper earns its place: it identifies a genuine gap in cinematic audio source separation, builds a first-of-its-kind dataset that puts non-verbal vocal sounds in the speech stem, and shows that conventional CASS models dump laughter and screams into the effect stem. The construction is transparent, the dataset is released on Zenodo, and the A/B test on real movie clips is a nice touch. I'd cite this if I worked in CASS.\n\nThat said, the causal claim is softer than the Table 2 numbers suggest. The training and evaluation clips come through the same tag-based filtering pipeline, so the eval set can carry the same contamination the authors admit remains in training ('Despite LLM-based filtering, some clips still contained non-human sounds'). If both sides contain the same label noise, part of the 5.62->9.30 dB speech improvement could be matching that noise rather than cleanly separating human non-verbal sounds. The model confusing an animal's voice with screaming (Sec. 4.4) is direct evidence that the learned mapping is not purely verbal. The SDR gains are large enough that the main findings probably survive a cleaner test, but the paper should provide one.\n\nOther soft spots are minor: no significance tests on the SDR gaps, a small and subjective A/B test (20 clips, 13 raters), and the simplified SDR metric. None of these are fatal. The most convincing part is Table 1, which shows the same model behaves very differently depending on how you label the non-verbal sounds — that is a clean demonstration of the dataset bias the paper is about.\n\nWho benefits: people building practical CASS systems, film restoration tooling, and dataset-curation researchers. It is a solid resource paper that deserves a serious referee. I would send it to peer review and ask for a clean-label subset evaluation (e.g., manually checked eval clips or a separate corpus) and significance reporting. Those are addressable, and the resource itself is valuable enough to justify the extra work.","headline":"Useful new CASS dataset with a real domain-mismatch fix, but the evaluation shares a contamination risk between train and eval and needs a cleaner external check.","tokens_in":8601,"tokens_out":1633,"would_cite":true,"duration_ms":18660,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding non-verbal vocal sounds to the speech stem of a cinematic source separation dataset fixes the model's tendency to misroute laughter and screams into the effects stem.","keywords":["cinematic audio source separation","non-verbal sounds","speech stem","DnR-nonverbal","FSD50K","FreeSound","BandIt","dataset curation"],"falsifier":"Manually audit the 471 FreeSound-derived clips (and a sample of the FSD50K clips) in the final DnR-nonverbal training set: if a substantial fraction contain non-human sounds, or if retraining on a hand-verified clean subset fails to reproduce the speech-SDR gain (9.30 dB vs 5.62 dB), the claim that clean non-verbal vocal labels drive the improvement would be falsified.","tokens_in":7623,"feed_emoji":"🎬","tokens_out":5854,"duration_ms":50722,"temperature":0.7,"pith_summary":"The paper claims that the reason cinematic audio source separation models push laughter, screams, and other emotionally heightened voices into the effects stem is that existing training datasets only contain reading-style speech. To fix this, the authors build DnR-nonverbal, a dataset that places non-verbal human vocalizations in the speech stem, using clips from FSD50K and FreeSound filtered by rules and a large language model. Training a BandIt model on DnR-v2 plus DnR-nonverbal raises speech SDR from 5.62 to 9.30 dB and effect SDR from 2.54 to 5.23 dB on a synthetic evaluation set. In a listener A/B test on real movie clips, 76.9% preferred the model retrained with the new dataset, so the fix appears to transfer beyond synthetic mixtures. If this holds, the paper's contribution is a dataset that removes a domain mismatch that made CASS models unreliable for expressive dialogue.","feed_headline":"Training on screams and laughter fixes movie-audio voice separation","feed_subtitle":"A new dataset improves speech SDR from 5.62 to 9.30 dB and wins 76.9% of listener tests on real movie clips.","key_machinery":"The load-bearing object is the DnR-nonverbal dataset itself: a 60-second-track extension of DnR-v2 in which the speech stem is mixed from reading-style LibriSpeech clips interleaved with non-verbal vocal clips drawn from FSD50K (by AudioSet-ontology tags) and newly crawled from FreeSound, with rule-based and GPT-4o-based filtering to remove non-human sounds, and a mix algorithm (zero-truncated Poisson clip counts, skew-Gaussian silences, LUFS-based loudness sampling) that mirrors DnR-v2's mixing for the music and effects stems. The dataset is what changes the model's behavior; BandIt is used only as a fixed probe to measure the effect.","core_discovery":"On the paper's own terms, the central claim is that non-verbal vocal sounds (laughter, screaming, whispering, crying, sighing, shouting) belong in the speech stem of a cinematic audio source separation dataset, and that including them there retrains the model to treat expressive voice as speech rather than as an effect. The paper shows that a BandIt model trained only on the conventional DnR-v2 dataset scores 5.62 dB on the speech stem when non-verbal sounds are counted as speech, but 6.52 dB when it is allowed to route those same sounds to the effects stem, which is direct evidence that the model is systematically misallocating them. Adding DnR-nonverbal closes that gap: the same model reaches 9.30 dB on speech while holding non-verbal sounds in the speech stem, and the subjective test on real movie audio indicates the retrained model extracts actors' voices more naturally and consistently. The intended consequence is a practical CASS tool that can handle acted-out dialogue in filmmaking and post-production.","pith_inferences":["The same domain-mismatch argument likely applies to other 'non-speech' vocal events in sound separation, e.g., in music source separation where beatboxing or spoken ad-libs may be routed to the wrong stem; DnR-nonverbal provides a template for auditing that.","The paper's own admitted failure mode — the retrained model sometimes mistakes an animal's voice for screaming — suggests that purely acoustic training has limits; a vision-modality context (as the authors suggest) or a stricter human-voice classifier would be a natural next test.","One unstated consequence is practical: if a CASS model can reliably keep expressive voice in the speech stem, dubbing, ADR, and subtitle workflows could be automated for dialogue that is not read calmly, which is most movie dialogue.","The 76.9% preference on real movie clips comes from a single small A/B test (13 raters, 20 clips); scaling that evaluation with more raters and clips, or measuring downstream tasks, would tell whether the subjective gain is robust."],"forward_implications":["Models trained with DnR-nonverbal separate expressive voice (screams, laughter, whispers) into the speech stem instead of dumping it into the effects stem, closing the main failure mode of CASS models on acted-out dialogue.","Because the effects stem is no longer contaminated by voice, effect SDR also improves (from 2.54 to 5.23 dB), and the music stem improves slightly as a side effect.","The fix transfers to real movie audio: 76.9% of listeners preferred the retrained model's speech extraction as more natural and consistent.","The dataset pipeline (tag-based collection plus LLM filtering) can be extended to other non-verbal vocal categories, and the resulting datasets could support query-based source separation and audio captioning, as the paper notes in conclusion."],"supporting_citations":[{"why":"Supplies the DnR-v2 dataset, mixing formulation, and loudness settings that DnR-nonverbal extends by changing only the speech stem.","marker":"[9]"},{"why":"Provides the FSD50K corpus from which the majority of non-verbal clips are collected via AudioSet-ontology tags.","marker":"[14]"},{"why":"BandIt, the state-of-the-art CASS model used in both training conditions to measure the dataset's effect.","marker":"[10]"},{"why":"LibriSpeech, the reading-style ASR corpus that makes up the conventional speech stem and is the source of the domain mismatch.","marker":"[12]"},{"why":"AudioSet ontology, which defines the six voice-related child tags used to select non-verbal clips.","marker":"[21]"},{"why":"Dynamic mixing, the training-time mixing strategy used for both the baseline and the proposed dataset.","marker":"[26]"}],"fun_headline_variants":["New dataset teaches AI to separate laughter and screams as speech","Screams and laughs now count as voice in movie audio separation","New audio dataset treats screams and laughter as speech, not effects","Dataset puts screams and laughs in speech stem to fix audio separation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The curation pipeline reliably isolates non-verbal human vocal clips, excluding music, animal sounds, and other noise, so that the speech stem teaches the model the right association and the reported gains are not artifacts of contaminated labels.","fun_headline_variants_meta":{"raw":{"variants":["New dataset teaches AI to separate laughter and screams as speech","Screams and laughs now count as voice in movie audio separation","New audio dataset treats screams and laughter as speech, not effects","Dataset puts screams and laughs in speech stem to fix audio separation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000678,"raw_usage":{"total_tokens":3065,"prompt_tokens":912,"completion_tokens":2153,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":2082}},"tokens_in":528,"tokens_out":2153,"duration_ms":15128,"temperature":1.0,"reasoning_tokens":2082,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:21:49.081405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Manually audit the 471 FreeSound-derived clips (and a sample of the FSD50K clips) in the final DnR-nonverbal training set: if a substantial fraction contain non-human sounds, or if retraining on a hand-verified clean subset fails to reproduce the speech-SDR gain (9.30 dB vs 5.62 dB), the claim that clean non-verbal vocal labels drive the improvement would be falsified.","supporting_citations":[{"cited_title":"The cocktail fork problem: Three-stem audio separation for real- world soundtracks,","cited_arxiv_id":null,"evidence_quote":"Provides the FSD50K corpus from which the majority of non-verbal clips are collected via AudioSet-ontology tags."},{"cited_title":"The 2018 signal separation eval- uation campaign,","cited_arxiv_id":null,"evidence_quote":"BandIt, the state-of-the-art CASS model used in both training conditions to measure the dataset's effect."},{"cited_title":"Hybrid transformers for music source separation,","cited_arxiv_id":null,"evidence_quote":"LibriSpeech, the reading-style ASR corpus that makes up the conventional speech stem and is the source of the domain mismatch."},{"cited_title":"Hy- perbolic audio source separation,","cited_arxiv_id":null,"evidence_quote":"AudioSet ontology, which defines the six voice-related child tags used to select non-verbal clips."},{"cited_title":"Audio set: An ontology and human-labeled dataset for audio events,","cited_arxiv_id":null,"evidence_quote":"Dynamic mixing, the training-time mixing strategy used for both the baseline and the proposed dataset."}],"review_version":1}