{"id":"9048d88c-ebcf-4fac-81c8-d144e658d89c","arxiv_id":"1908.10055","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Environmental sound synthesis should be evaluated with intelligibility, realism, and naturalness tests together, since intelligibility alone can rate unnatural sounds highly.","lead":"Okamoto and colleagues review the field of environmental sound synthesis and conversion, and run listening tests on WaveNet-generated sounds from event labels. They find that evaluating synthesized sounds on intelligibility alone can miss poor naturalness, so multiple subjective tests should be used.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's only direct evidence for its central claim is the whistle's high intelligibility/low naturalness gap in Fig. 9, which rests on roughly two sound files and 24 ratings per condition with no significance test; if that gap is sampling noise, the proposal loses its empirical basis.","rationale":"The reader's weakest assumption is that the whistle result generalizes from one model, one dataset, and a small listener pool to environmental sound synthesis broadly. My concern is more specific and more internal: even within this experiment, the whistle dissociation is the sole direct support for the central claim, and it is backed by a very small per-class sample with no reported inferential statistics. The reader already asked for confidence intervals or significance tests and recommended a conditional accept, so my critique reinforces the same condition rather than changing the verdict. I do not see an internal inconsistency in the paper's logic: the conclusion that intelligibility alone is insufficient does follow if the whistle dissociation is real, and the paper is appropriately cautious in recommending 'distinguishability and/or naturalness' rather than prescribing a single fixed battery. However, because the empirical cornerstone is statistically fragile, the paper should add significance testing and ideally replicate the dissociation across more items and at least one additional synthesis model before the recommendation is treated as established. The verdict remains CONDITIONAL, which is the same as the reader's verdict, so no adjustment is needed.","tokens_in":7265,"tokens_out":3553,"duration_ms":43039,"concrete_test":"Obtain the raw per-class MOS ratings behind Fig. 9 and fit a mixed-effects model with listener and sound-item random effects for each class, testing the real-versus-synthesized difference for the whistle. If the 95% confidence interval for the whistle gap includes zero, or the effect is not significant after accounting for item-level clustering, the claimed dissociation is not established. As a robustness check, repeat Experiment III with 10 or more synthesized whistle clips (and 10 real clips) to rule out a single-file artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central proposal in Sec. 3.2 is that intelligibility alone is unsatisfactory for evaluating environmental sound synthesis, so methods should also be tested on distinguishability and/or naturalness. Logically, one credible dissociation between intelligibility and another quality dimension would be enough to show that a single intelligibility test is insufficient. The whistle result is that dissociation: listeners classify synthesized whistles almost as well as real ones, but rate their naturalness much lower. The load-bearing assumption is that this gap is real and not an artifact of the experiment.\n\nThat assumption is weakly supported. Experiment III used only 2 samples per label with 24 listeners, for 480 total ratings across 10 labels. For the whistle class, the real-versus-synthesized MOS comparison is therefore based on at most a handful of audio files, possibly one file per condition, and 24 ratings per cell. No significance test or confidence interval is reported for the whistle gap; the paper only says the differences are 'large.' If the whistle gap is within sampling variability, or is specific to the single synthesized whistle file used, then the paper contains no remaining direct evidence that intelligibility and naturalness dissociate. Experiment II reports distinguishability, but it is not connected to the intelligibility results in a way that supports the insufficiency claim. Thus the central recommendation depends on a single per-class comparison whose statistical reliability is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reviews the emerging area of environmental sound synthesis and conversion, organizing it into four task families: synthesis from sound event/scene labels, synthesis from onomatopoeic words, conversion between environmental sounds, and synthesis/conversion from multimedia inputs. It then reports a subjective evaluation study of a conditional WaveNet system for sound event synthesis using ten classes from RWCP-SSD. Three experiments were conducted with 24 listeners: a forced-choice intelligibility test (classification recall), an AB distinguishability test between real and synthesized sounds, and a five-scale MOS naturalness test. The intelligibility results show an average F-score drop from 86.22% (real) to 76.30% (synthesized); the AB test shows listeners identify real sounds with 82.71% accuracy; and the MOS results show generally lower naturalness for synthesized sounds, with particularly large gaps for the electric shaver, trash box, and whistle. On the basis of the whistle class, where intelligibility is high but naturalness is much lower for synthesized sounds, the paper proposes that environmental sound synthesis should be evaluated not only by intelligibility but also by distinguishability and/or naturalness.","tokens_in":7511,"tokens_out":4258,"duration_ms":42527,"significance":"If the central claim is accepted, the paper provides a useful practical guideline for a field that currently lacks standardized evaluation protocols. The review portion is valuable in itself: it gives the first compact taxonomy of environmental sound synthesis and conversion tasks, links them to applications such as data augmentation and media production, and identifies the absence of established subjective and objective evaluation methods. The paper also ships demo audio, which supports reproducibility of the perceptual judgments. The proposed three-dimensional evaluation (intelligibility, distinguishability, naturalness) is plausible and aligns with practices in speech and music synthesis. However, the empirical foundation for the core methodological recommendation is narrow: the dissociation between intelligibility and naturalness rests on a single sound class, a small number of stimuli, and no inferential statistics. The strength of the recommendation therefore currently exceeds the strength of the evidence.","major_comments":[{"comment":"The central recommendation is supported by exactly one empirical dissociation: the whistle class shows high intelligibility (93.3% recall for synthesized sounds in Fig. 6 and 95.8% for real sounds in Fig. 5) but a large apparent drop in naturalness MOS in Fig. 9. This comparison rests on only two whistle samples per condition (Table 2), 24 listeners per cell, and no significance test or confidence interval for the whistle MOS gap. If that gap is sampling noise or an artifact of the single synthesized whistle file used, the paper contains no remaining direct evidence that intelligibility alone is insufficient for evaluating environmental sound synthesis. Please report per-condition confidence intervals or a paired test for the whistle class, or explicitly reframe the proposal as a preliminary observation rather than a general methodological conclusion.","section":"Sec. 3.2, Fig. 9, Table 2"},{"comment":"Experiment II shows that listeners identify real versus synthesized sounds with 82.71% accuracy, which indicates that synthesized sounds are not indistinguishable from real sounds. However, this result is not connected to the intelligibility scores per class, so it does not by itself show that intelligibility testing is insufficient. To support the insufficiency claim, the paper needs a per-class comparison showing high intelligibility coexisting with poor distinguishability or naturalness; Experiment II as reported does not provide that. Please present the AB results alongside the intelligibility results, or analyze the joint per-class outcomes, so that the distinguishability dimension can actually be compared with intelligibility.","section":"Sec. 3.2, Experiment II"},{"comment":"The experiments use only 24 listeners, ten labels, and two to five samples per label per condition, and all comparisons are reported as point estimates without uncertainty quantification. With this sample size, the absence of a large difference for classes such as coffee grinder, clock, and maracas in Fig. 9 cannot be interpreted as equivalence between real and synthesized sounds. Please provide uncertainty quantification for at least the main comparisons, and ideally use more stimuli per class, before drawing general conclusions about which evaluation dimensions are needed for environmental sound synthesis.","section":"Sec. 3.1, Table 2"}],"minor_comments":[{"comment":"The sentence 'From the results of experiment I, it considered that this subjective test is particularly helpful' should read 'it is considered that'; the current phrasing appears to be a typo.","section":"Sec. 3.2"},{"comment":"Reference [4] lists the year as '20111' instead of '2011'.","section":"References"},{"comment":"The caption 'Recognition rate of real sounds' is ambiguous for a figure reporting the AB test; please clarify that it is the percentage of trials in which listeners correctly identified the real sample.","section":"Fig. 8"},{"comment":"The claim that 'there is no literature giving an overview of the problem definitions and evaluation methods' would be safer as 'to the best of our knowledge'; as written, it overstates certainty about the literature.","section":"Sec. 2.1"},{"comment":"For the objective evaluation methods listed (PESQ, POLQA, PEAQ), the paper does not mention whether any correlation with subjective naturalness or intelligibility has been shown for environmental sounds; one sentence of context would help readers judge their applicability.","section":"Sec. 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a workshop-scale study; the review portion is useful and the experimental question is appropriate, but the empirical support for the central methodological claim is too thin in its current form. I would encourage a revision that either strengthens the statistical evidence or narrows the claim, rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the taxonomy in Sec. 2: SES, SSS, SEC, SSC, and onomatopoeic synthesis laid out as a coherent set of tasks. That alone is worth having, since prior work (Kong, Ikawa, Grinstein) just proposes methods without organizing the field. The three subjective experiments are a reasonable first pass at evaluation methodology, and the whistle dissociation in Fig. 9 is a neat illustration of why intelligibility alone can mislead.\n\nThe soft spots are real but not fatal. Experiment III uses only two samples per label, so the whistle comparison rests on one or two synthesized files and 24 ratings per cell. The paper does show 95% confidence intervals in Fig. 9, which the stress-test note missed, but it never reports whether the real-vs-synthesized gap for whistle is significant or whether the intervals overlap. That is the load-bearing comparison. If that gap is sampling noise, the recommendation to always test naturalness or distinguishability loses its empirical basis. Experiment II measures distinguishability but is never connected to the intelligibility results, so it does not directly support the insufficiency claim either.\n\nThat said, the recommendation itself — test more than intelligibility — is sensible and likely correct even if this particular evidence is thin. The paper would be stronger with significance tests on the whistle gap, more samples per class, or a second model/dataset to show the dissociation is not file-specific. As it stands, it is a useful position paper with a modest empirical pilot, not a settled result.\n\nThe references look appropriate, and I do not see any citation problems. The writing is clear and the authors are honest about the limits of WaveNet's quality.\n\nWho should read this? Anyone working on environmental sound synthesis, data augmentation, or subjective evaluation of non-speech audio. It deserves a serious referee; I would send it out with the expectation that the empirical section needs tightening before publication.","headline":"Useful taxonomy of environmental sound synthesis tasks, but the empirical case for multi-dimensional evaluation rests on one class with very few samples and no direct significance test.","tokens_in":8062,"tokens_out":1793,"would_cite":false,"duration_ms":20447,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Intelligibility alone cannot judge synthesized environmental sounds: a whistle identified correctly can still sound unnatural, so the paper recommends always pairing intelligibility tests with distinguishability and naturalness tests.","keywords":["environmental sound synthesis","sound event synthesis","sound scene synthesis","environmental sound conversion","subjective evaluation","intelligibility","naturalness","WaveNet"],"falsifier":"Run the same three subjective tests across several generative models and a larger, more diverse set of sound-event classes. If every class that is correctly identified by listeners also receives a naturalness score close to that of the real recording, or if the whistle is the only class showing the dissociation, then the claim that intelligibility alone is unsatisfactory would not generalize. A cheap first check is spectral analysis: if the whistle's low naturalness comes from missing fine spectral structure, then classes with strong tonal identity but weak noise texture should show the same split, while broadband noisy classes should not.","tokens_in":7089,"feed_emoji":"🎧","tokens_out":6110,"duration_ms":64801,"temperature":0.7,"pith_summary":"This paper argues that a single subjective test, whether listeners can identify what sound they heard, is not enough to judge environmental sound synthesis; researchers should also measure how easily synthesized sounds can be told from real recordings and how natural they sound. The argument builds on a review of environmental sound synthesis and conversion tasks, followed by a WaveNet-based sound event synthesis experiment over ten everyday sound classes. When listeners tried to name synthesized sounds, most classes, including whistles, were identified about as well as real sounds, but whistles scored far below real whistles on naturalness. The paper concludes that intelligibility and naturalness come apart, so evaluation protocols should combine intelligibility with distinguishability and/or naturalness.","feed_headline":"Intelligibility alone can't judge synthesized sounds","feed_subtitle":"A whistle that listeners identify perfectly still sounds unnatural, so environmental sound synthesis needs extra subjective tests.","key_machinery":"The central object carrying the argument is a conditional WaveNet, an autoregressive generative model that produces raw audio waveforms conditioned on a one-hot sound-event label, because it supplies the synthesized sounds whose quality is under test. The measurement machinery is the three-experiment subjective protocol: Experiment I forces listeners to choose a sound-event label for each sound, Experiment II is an AB preference test asking which of a real and synthesized pair sounds more real, and Experiment III is a five-scale mean opinion score for naturalness. The argument works by comparing performance across these three lenses for the same sound classes; the whistle's high identification score alongside low naturalness is what makes intelligibility look insufficient on its own.","core_discovery":"The central claim is that intelligibility is not a sufficient subjective measure for environmental sound synthesis: a synthesized sound can be correctly labeled by listeners yet still be perceived as clearly less natural than the real thing. From three listener experiments on sound event synthesis with a conditional WaveNet, the paper reports that average recognition F-scores were 86.22 percent for real sounds and 76.30 percent for synthesized sounds, while listeners identified real sounds as real only 82.71 percent of the time in an AB test, and mean opinion scores for naturalness varied strongly by category. The decisive case is the whistle: listeners named real and synthesized whistles with comparably high accuracy, yet the synthesized whistle's naturalness score was much lower than the real whistle's. The authors infer that a synthesis method that gets the category right but sounds wrong can pass an intelligibility-only evaluation, and they recommend evaluating environmental sound synthesis with intelligibility plus distinguishability and/or naturalness.","pith_inferences":["Extending the paper's logic, an automatic classifier-based intelligibility score would likely show the same blind spot, over-rating tonal classes such as whistles; a robust objective metric for environmental sound would probably need two components, classifiability plus a distributional naturalness measure.","The spectrogram evidence suggests a testable hypothesis: the intelligibility-naturalness gap should be largest for sounds whose identity lives in narrow spectral bands while their naturalness depends on fine stochastic texture, and smallest for broadband noisy sounds.","The paper stops short of proposing a single combined subjective score; a practical next step would be to define a gated reporting convention, such as reporting intelligibility only for sounds that first pass a naturalness threshold, so results across papers become comparable."],"forward_implications":["Future comparisons of environmental sound synthesis methods should report at least one perceptual metric beyond label intelligibility, or a method that reproduces only category-level cues can appear artificially competitive.","For applications that use synthesized sound directly, such as film and game production and virtual reality, naturalness and indistinguishability from real recordings become the deciding quality measures, whereas data augmentation for detection may continue to rely primarily on intelligibility.","The same three-test protocol transfers naturally to sound scene synthesis and to sound event and scene conversion, where no standard subjective evaluation currently exists.","Until an objective metric for environmental sound quality is validated, subjective multi-test evaluation remains necessary; the speech and audio objective metrics cited in the paper do not cover this domain.","The failure case for intelligibility-only evaluation is not marginal: a category such as whistle can pass identification almost perfectly while failing naturalness by a large margin."],"supporting_citations":[{"why":"Provides the conditional autoregressive raw-audio generator whose synthesized sounds are used in the three subjective experiments.","marker":"[9]"},{"why":"Defines label-conditioned environmental sound synthesis and supplies the reviewed task formulation that the experiment evaluates.","marker":"[10]"},{"why":"Supplies the SampleRNN architecture on which the reviewed label-conditioned sound scene synthesis method is based.","marker":"[12]"},{"why":"Supplies the ten-class sound-event database from which training and test sounds are drawn for the subjective tests.","marker":"[20]"},{"why":"Represents objective speech-quality evaluation, cited as an example of objective metrics that do not yet exist for environmental sound.","marker":"[13]"},{"why":"Represents objective audio-quality evaluation, used to contrast with the absence of objective metrics for environmental sound.","marker":"[15]"}],"fun_headline_variants":["Whistle test: accurate labels still sound fake","Sound synthesis needs more than a correct label","Intelligibility isn't enough for sound synthesis","Why naturalness tests matter for synthetic audio","Synthetic sounds pass label tests, fail on naturalness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recommendation rests on the assumption that the split between good intelligibility and poor naturalness seen for one sound class, one generative model, one ten-class database, and 24 listeners is a general fact about environmental sound synthesis rather than a quirk of that particular setup.","fun_headline_variants_meta":{"raw":{"variants":["Whistle test: accurate labels still sound fake","Sound synthesis needs more than a correct label","Intelligibility isn't enough for sound synthesis","Why naturalness tests matter for synthetic audio","Synthetic sounds pass label tests, fail on naturalness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1269,"prompt_tokens":844,"completion_tokens":425,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":354}},"tokens_in":460,"tokens_out":425,"duration_ms":4314,"temperature":1.0,"reasoning_tokens":354,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:53:02.347788+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three subjective tests across several generative models and a larger, more diverse set of sound-event classes. If every class that is correctly identified by listeners also receives a naturalness score close to that of the real recording, or if the whistle is the only class showing the dissociation, then the claim that intelligibility alone is unsatisfactory would not generalize. A cheap first check is spectral analysis: if the whistle's low naturalness comes from missing fine spectral structure, then classes with strong tonal identity but weak noise texture should show the same split, while broadband noisy classes should not.","supporting_citations":[{"cited_title":"Visual to sound: Generating natural sound for videos in the wild,","cited_arxiv_id":null,"evidence_quote":"Provides the conditional autoregressive raw-audio generator whose synthesized sounds are used in the three subjective experiments."},{"cited_title":"Scaper: A library for soundscape synthesis and aug- mentation,","cited_arxiv_id":null,"evidence_quote":"Defines label-conditioned environmental sound synthesis and supplies the reviewed task formulation that the experiment evaluates."},{"cited_title":"Seed: Resynthesi z- ing environmental sounds from examples,","cited_arxiv_id":null,"evidence_quote":"Supplies the SampleRNN architecture on which the reviewed label-conditioned sound scene synthesis method is based."},{"cited_title":"Generating sound words from audio signals of acoustic events with sequence-to-sequenc e model,","cited_arxiv_id":null,"evidence_quote":"Supplies the ten-class sound-event database from which training and test sounds are drawn for the subjective tests."},{"cited_title":"Introduction to acoustic event and scene ana lysis,","cited_arxiv_id":null,"evidence_quote":"Represents objective audio-quality evaluation, used to contrast with the absence of objective metrics for environmental sound."}],"review_version":1}