Pith. sign in

REVIEW 3 major objections 5 minor 25 references

Overview of Tasks and Investigation of Subjective Evaluation Methods in Environmental Sound Synthesis and Conversion

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Intelligibility alone cannot judge synthesized environmental sounds: a whistle identified correctly can still sound unnatural, so the paper recommends always pairing intelligibility tests with distinguishability and naturalness tests.

desk verdict Useful taxonomy of environmental sound synthesis tasks, but the empirical case for multi-dimensional evaluation rests on one class with very few samples and no direct significance test. read the letter →

arxiv 1908.10055 v1 pith:DSU2OSZU submitted 2019-08-27 cs.SD eess.AS

classification cs.SDeess.AS
keywords environmentalsoundsynthesiseventsceneconversionsubjectiveevaluationintelligibilitynaturalnessWaveNet
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a single subjective test, whether listeners can identify what sound they heard, is not enough to judge environmental sound synthesis; researchers should also measure how easily synthesized sounds can be told from real recordings and how natural they sound. The argument builds on a review of environmental sound synthesis and conversion tasks, followed by a WaveNet-based sound event synthesis experiment over ten everyday sound classes. When listeners tried to name synthesized sounds, most classes, including whistles, were identified about as well as real sounds, but whistles scored far below real whistles on naturalness. The paper concludes that intelligibility and naturalness come apart, so evaluation protocols should combine intelligibility with distinguishability and/or naturalness.

What carries the argument

The central object carrying the argument is a conditional WaveNet, an autoregressive generative model that produces raw audio waveforms conditioned on a one-hot sound-event label, because it supplies the synthesized sounds whose quality is under test. The measurement machinery is the three-experiment subjective protocol: Experiment I forces listeners to choose a sound-event label for each sound, Experiment II is an AB preference test asking which of a real and synthesized pair sounds more real, and Experiment III is a five-scale mean opinion score for naturalness. The argument works by comparing performance across these three lenses for the same sound classes; the whistle's high identification score alongside low naturalness is what makes intelligibility look insufficient on its own.

What would settle it

Run the same three subjective tests across several generative models and a larger, more diverse set of sound-event classes. If every class that is correctly identified by listeners also receives a naturalness score close to that of the real recording, or if the whistle is the only class showing the dissociation, then the claim that intelligibility alone is unsatisfactory would not generalize. A cheap first check is spectral analysis: if the whistle's low naturalness comes from missing fine spectral structure, then classes with strong tonal identity but weak noise texture should show the same split, while broadband noisy classes should not.

Watch

Extended reading notes

Core claim

The central claim is that intelligibility is not a sufficient subjective measure for environmental sound synthesis: a synthesized sound can be correctly labeled by listeners yet still be perceived as clearly less natural than the real thing. From three listener experiments on sound event synthesis with a conditional WaveNet, the paper reports that average recognition F-scores were 86.22 percent for real sounds and 76.30 percent for synthesized sounds, while listeners identified real sounds as real only 82.71 percent of the time in an AB test, and mean opinion scores for naturalness varied strongly by category. The decisive case is the whistle: listeners named real and synthesized whistles with comparably high accuracy, yet the synthesized whistle's naturalness score was much lower than the real whistle's. The authors infer that a synthesis method that gets the category right but sounds wrong can pass an intelligibility-only evaluation, and they recommend evaluating environmental sound synthesis with intelligibility plus distinguishability and/or naturalness.

Load-bearing premise

The recommendation rests on the assumption that the split between good intelligibility and poor naturalness seen for one sound class, one generative model, one ten-class database, and 24 listeners is a general fact about environmental sound synthesis rather than a quirk of that particular setup.

Editorial extensions

If this is right

  • Future comparisons of environmental sound synthesis methods should report at least one perceptual metric beyond label intelligibility, or a method that reproduces only category-level cues can appear artificially competitive.
  • For applications that use synthesized sound directly, such as film and game production and virtual reality, naturalness and indistinguishability from real recordings become the deciding quality measures, whereas data augmentation for detection may continue to rely primarily on intelligibility.
  • The same three-test protocol transfers naturally to sound scene synthesis and to sound event and scene conversion, where no standard subjective evaluation currently exists.
  • Until an objective metric for environmental sound quality is validated, subjective multi-test evaluation remains necessary; the speech and audio objective metrics cited in the paper do not cover this domain.
  • The failure case for intelligibility-only evaluation is not marginal: a category such as whistle can pass identification almost perfectly while failing naturalness by a large margin.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extending the paper's logic, an automatic classifier-based intelligibility score would likely show the same blind spot, over-rating tonal classes such as whistles; a robust objective metric for environmental sound would probably need two components, classifiability plus a distributional naturalness measure.
  • The spectrogram evidence suggests a testable hypothesis: the intelligibility-naturalness gap should be largest for sounds whose identity lives in narrow spectral bands while their naturalness depends on fine stochastic texture, and smallest for broadband noisy sounds.
  • The paper stops short of proposing a single combined subjective score; a practical next step would be to define a gated reporting convention, such as reporting intelligibility only for sounds that first pass a naturalness threshold, so results across papers become comparable.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reviews the emerging area of environmental sound synthesis and conversion, organizing it into four task families: synthesis from sound event/scene labels, synthesis from onomatopoeic words, conversion between environmental sounds, and synthesis/conversion from multimedia inputs. It then reports a subjective evaluation study of a conditional WaveNet system for sound event synthesis using ten classes from RWCP-SSD. Three experiments were conducted with 24 listeners: a forced-choice intelligibility test (classification recall), an AB distinguishability test between real and synthesized sounds, and a five-scale MOS naturalness test. The intelligibility results show an average F-score drop from 86.22% (real) to 76.30% (synthesized); the AB test shows listeners identify real sounds with 82.71% accuracy; and the MOS results show generally lower naturalness for synthesized sounds, with particularly large gaps for the electric shaver, trash box, and whistle. On the basis of the whistle class, where intelligibility is high but naturalness is much lower for synthesized sounds, the paper proposes that environmental sound synthesis should be evaluated not only by intelligibility but also by distinguishability and/or naturalness.

Significance. If the central claim is accepted, the paper provides a useful practical guideline for a field that currently lacks standardized evaluation protocols. The review portion is valuable in itself: it gives the first compact taxonomy of environmental sound synthesis and conversion tasks, links them to applications such as data augmentation and media production, and identifies the absence of established subjective and objective evaluation methods. The paper also ships demo audio, which supports reproducibility of the perceptual judgments. The proposed three-dimensional evaluation (intelligibility, distinguishability, naturalness) is plausible and aligns with practices in speech and music synthesis. However, the empirical foundation for the core methodological recommendation is narrow: the dissociation between intelligibility and naturalness rests on a single sound class, a small number of stimuli, and no inferential statistics. The strength of the recommendation therefore currently exceeds the strength of the evidence.

major comments (3)
  1. [Sec. 3.2, Fig. 9, Table 2] The central recommendation is supported by exactly one empirical dissociation: the whistle class shows high intelligibility (93.3% recall for synthesized sounds in Fig. 6 and 95.8% for real sounds in Fig. 5) but a large apparent drop in naturalness MOS in Fig. 9. This comparison rests on only two whistle samples per condition (Table 2), 24 listeners per cell, and no significance test or confidence interval for the whistle MOS gap. If that gap is sampling noise or an artifact of the single synthesized whistle file used, the paper contains no remaining direct evidence that intelligibility alone is insufficient for evaluating environmental sound synthesis. Please report per-condition confidence intervals or a paired test for the whistle class, or explicitly reframe the proposal as a preliminary observation rather than a general methodological conclusion.
  2. [Sec. 3.2, Experiment II] Experiment II shows that listeners identify real versus synthesized sounds with 82.71% accuracy, which indicates that synthesized sounds are not indistinguishable from real sounds. However, this result is not connected to the intelligibility scores per class, so it does not by itself show that intelligibility testing is insufficient. To support the insufficiency claim, the paper needs a per-class comparison showing high intelligibility coexisting with poor distinguishability or naturalness; Experiment II as reported does not provide that. Please present the AB results alongside the intelligibility results, or analyze the joint per-class outcomes, so that the distinguishability dimension can actually be compared with intelligibility.
  3. [Sec. 3.1, Table 2] The experiments use only 24 listeners, ten labels, and two to five samples per label per condition, and all comparisons are reported as point estimates without uncertainty quantification. With this sample size, the absence of a large difference for classes such as coffee grinder, clock, and maracas in Fig. 9 cannot be interpreted as equivalence between real and synthesized sounds. Please provide uncertainty quantification for at least the main comparisons, and ideally use more stimuli per class, before drawing general conclusions about which evaluation dimensions are needed for environmental sound synthesis.
minor comments (5)
  1. [Sec. 3.2] The sentence 'From the results of experiment I, it considered that this subjective test is particularly helpful' should read 'it is considered that'; the current phrasing appears to be a typo.
  2. [References] Reference [4] lists the year as '20111' instead of '2011'.
  3. [Fig. 8] The caption 'Recognition rate of real sounds' is ambiguous for a figure reporting the AB test; please clarify that it is the percentage of trials in which listeners correctly identified the real sample.
  4. [Sec. 2.1] The claim that 'there is no literature giving an overview of the problem definitions and evaluation methods' would be safer as 'to the best of our knowledge'; as written, it overstates certainty about the literature.
  5. [Sec. 3.1] For the objective evaluation methods listed (PESQ, POLQA, PEAQ), the paper does not mention whether any correlation with subjective naturalness or intelligibility has been shown for environmental sounds; one sentence of context would help readers judge their applicability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the evaluation-method recommendation rests on external subjective measurements, not on fitted or self-referential quantities.

full rationale

The paper's central claim in Sec. 3.2 is that environmental sound synthesis should be evaluated by intelligibility plus distinguishability and/or naturalness. This claim is supported by independently measured listener responses (classification accuracy, AB preference, MOS) on real and WaveNet-synthesized sounds. No parameter is fitted and then renamed as a prediction; the intelligibility/naturalness dissociation (Figs. 5, 6, 9) is an empirical observation, not a quantity defined in terms of the model's training objective. The paper does cite one author's prior overview [11], but that citation is background context and is not load-bearing for the experimental conclusion. The WaveNet configuration is external tooling, and the subjective evaluation results are measured data rather than outputs of a derivation. Consequently, no circular step can be exhibited, and the paper is self-contained with respect to its evaluation-methods proposal.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters are fitted to the subjective scores. The WaveNet hyperparameters (Table 1) are conventional settings, not tuned to maximize the reported metrics. The central claim rests on four domain assumptions about data, listeners, scaling, and model representativeness.

assumptions (4)
  • domain assumption RWCP-SSD sound event labels are accurate and representative of the 10 chosen environmental sound classes.
    The experiments treat these labels as ground truth for the intelligibility test and for training the conditional WaveNet (Sec. 3.1).
  • domain assumption The 24 listeners (13 male, 11 female) are a reliable and generalizable panel for judging environmental sound naturalness.
    All subjective conclusions rest on this listener panel from a single university, with no inter-rater reliability analysis (Sec. 3.1).
  • domain assumption A 5-point MOS scale yields interval-level data suitable for averaging and comparison with a 95% confidence interval.
    The naturalness comparison in Fig. 9 averages ordinal ratings across listeners and items, an assumption common in speech synthesis evaluation but not tested here.
  • domain assumption Conditional WaveNet is a representative statistical generative model for environmental sound synthesis, so its performance indicates the current state of the art.
    The paper draws general conclusions about statistical environmental sound synthesis from a single model implementation (Sec. 3.1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Overview of Tasks and Investigation of Subjective Evaluation Methods in Environmental Sound Synthesis and Conversion." pith.science (2026). https://pith.science/paper/DSU2OSZU

@misc{pith2026190810055,
  author       = {Pith},
  title        = {Pith review of: Overview of Tasks and Investigation of Subjective Evaluation Methods in Environmental Sound Synthesis and Conversion},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DSU2OSZU}},
  note         = {Machine review of arXiv:1908.10055}
}
read the original abstract

Synthesizing and converting environmental sounds have the potential for many applications such as supporting movie and game production, data augmentation for sound event detection and scene classification. Conventional works on synthesizing and converting environmental sounds are based on a physical modeling or concatenative approach. However, there are a limited number of works that have addressed environmental sound synthesis and conversion with statistical generative models; thus, this research area is not yet well organized. In this paper, we review problem definitions, applications, and evaluation methods of environmental sound synthesis and conversion. We then report on environmental sound synthesis using sound event labels, in which we focus on the current performance of statistical environmental sound synthesis and investigate how we should conduct subjective experiments on environmental sound synthesis.

Figures

Figures reproduced from arXiv: 1908.10055 by the authors.

Figure 1
Figure 1. Problem definition of sound scene synthesis [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Problem definition of environmental sound synthes [PITH_FULL_IMAGE:figures/full_fig_p001_2.png] view at source ↗
Figure 3
Figure 3. Problem definition of environmental sound synthes [PITH_FULL_IMAGE:figures/full_fig_p002_3.png] view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: Confusion matrix of classification accuracy for or [PITH_FULL_IMAGE:figures/full_fig_p003_5.png]
Figure 4
Figure 4. Figure 4: We call this kind of task a sound event conversion (SEC) or sound scene conversion (SSC). When we have some background sounds or sound effects but they are not suitable for the movie or game, environmental sound conversion can also be applied to obtain desirable sounds…
Figure 7
Figure 7. Figure 7: Spectrograms of real and synthesized environment [PITH_FULL_IMAGE:figures/full_fig_p003_7.png]
Figure 8
Figure 8. Figure 8: Recognition rate of real sounds detail; thus, we here discuss how we should conduct subjective tests for environmental sound synthesis. For synthesized sounds, it is im￾portant that (I) they are distinguishable from other types of environ￾mental sound, (II) they are no…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

25 extracted references · 23 canonical work pages

  1. [1]

    Sound synthesis and conversion me thods with the aim of generating speech or music have been widely de - veloped [1, 2, 3]

    INTRODUCTION Sound synthesis and conversion are techniques for generating a nat- ural sound using a statistical model that associates input information with the generated sound. Sound synthesis and conversion me thods with the aim of generating speech or music have been widely de - veloped [1, 2, 3]. Recently, some researchers have also deve loped methods...

  2. [2]

    style” and “content

    PROBLEM DEFINITIONS OF ENVIRONMENTAL SOUND SYNTHESIS AND CONVERSION In this section, we review applications, problem definition s, and evaluation methods of environmental sound synthesis and co nver- sion, specifically environmental sound synthesis using eve nt or scene labels (Sec. 2.1), environmental sound synthesis usi ng ono- toNtoNtoNtoNtoN zaaaaaabror...

  3. [3]

    INVESTIGATION OF SUBJECTIVE EV ALUATION METHOD 3.1. Experimental Conditions In this section, by evaluating SES using sound event labels based on the conditional WaveNet [9], we discuss the current perform ance of environmental sound synthesis and how we should conduct a Predicted label Coffee grinder Cup Clock Whistle Maracas Drum Shaver Trash box Tearing...

  4. [4]

    We then discussed the current performance of sound event synthesis and subjective evaluation methods of environmental sound synthesis

    CONCLUSION In this paper, we presented the problem definitions of sound e vent synthesis, sound scene synthesis, and sound event and scene con- version. We then discussed the current performance of sound event synthesis and subjective evaluation methods of environmental sound synthesis. The evaluation experiments indicate that sounds synthesized by WaveNet...

  5. [5]

    Statistical parametric speech synthesis,

    H. Zen, K. Tokuda, and A. Black, “Statistical parametric speech synthesis,” Speech Communication, vol. 51, no. 11, pp. 1039–1064, 2009

  6. [6]

    An overview of voice con- version systems,

    S. H. Mohammadi and A. Kain, “An overview of voice con- version systems,” Speech Communication, vol. 88, pp. 65–82, 2017

  7. [7]

    Deep learning techniques for music generation – a survey,

    J. P . Briot, G. Hadjeres, and F. Pachet, “Deep learning techniques for music generation – a survey,” arXiv preprint arXiv:1709.01620, 2017

  8. [8]

    Soun d synthesis for impact sounds in video games,

    D. B. Lloyd, N. Raghuvanshi, and N. K. Govindaraju, “Soun d synthesis for impact sounds in video games,” Proc. Sympo- sium on Interactive 3D Graphics and Games. ACM , pp. 55– 61, 20111

Show all 25 references
  1. [9]

    Visual to sound: Generating natural sound for videos in the wild,

    Y . Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg, “Visual to sound: Generating natural sound for videos in the wild,” Proc. IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pp. 3550–3558, 2018

  2. [10]

    Scaper: A library for soundscape synthesis and aug- mentation,

    J. Salamon, D. MacConnell, M. Cartwright, P . Li, and J. P . Bello, “Scaper: A library for soundscape synthesis and aug- mentation,” Proc. IEEE W orkshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp. 344–348, 2017

  3. [11]

    State of the art in sound texture synthesis,

    D. Schwarz, “State of the art in sound texture synthesis, ” Proc. Digital Audio Effects (DAFx), pp. 221–232, 2011

  4. [12]

    Seed: Resynthesi z- ing environmental sounds from examples,

    G. Bernardes, L. Aly, and M. E. Davies, “Seed: Resynthesi z- ing environmental sounds from examples,” Proc. the Sound and Music Computing Conference, pp. 55–62, 2016

  5. [13]

    WaveNet: A generative model for raw au- dio,

    A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw au- dio,” arXiv preprint, arXiv:1609.03499, 2016

  6. [14]

    Acoustic scene generation with conditional sam - pleRNN,

    Q. Kong, Y . Xu, T. Iqbal, Y . Cao, W. Wang, and M. D. Plumbley, “Acoustic scene generation with conditional sam - pleRNN,” Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 925–929, 2019

  7. [15]

    Introduction to acoustic event and scene ana lysis,

    K. Imoto, “Introduction to acoustic event and scene ana lysis,” Acoustical Science and Technology , vol. 39, no. 3, pp. 182– 188, 2018

  8. [16]

    SampleRNN: An unconditional end-to-end neural audio generation model,

    S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. So telo, A. Courville, and Y . Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” Proc. Interna- tional Conference for Learning Representations (ICLR), pp. 1–11, 2017

  9. [17]

    Perceptual evaluation of speech quality (PESQ): An ob - jective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,

    “Perceptual evaluation of speech quality (PESQ): An ob - jective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” ITU-T Recommendation P .862, 2001

  10. [18]

    Perceptual objective listening quality assessment,

    “Perceptual objective listening quality assessment, ” ITU-T Recommendation P .863, 2011

  11. [19]

    Method for objective measurements of perceived audio qual- ity,

    “Method for objective measurements of perceived audio qual- ity,” ITU-R Recommendation BS.1387-1, 2001

  12. [20]

    Generating sound words from audio signals of acoustic events with sequence-to-sequenc e model,

    S. Ikawa and K. Kashino, “Generating sound words from audio signals of acoustic events with sequence-to-sequenc e model,” Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 346–350, 2018

  13. [21]

    A u- dio style transfer,

    E. Grinstein, N. Q. K. Duong, A. Ozerov, and P . P´ erez, “A u- dio style transfer,” Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 586– 590, 2018

  14. [22]

    Time domain neural audio style transfer,

    P . K. Mital, “Time domain neural audio style transfer,” arXiv preprint, p. arXiv:1711.11160, 2017

  15. [23]

    Image style tran s- fer using convolutional neural networks,

    L. A. Gatys, A. S. Ecker, and M. Bethge, “Image style tran s- fer using convolutional neural networks,” Proc. IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR), pp. 2414–2423, 2016

  16. [24]

    Acoustic al sound database in real environments for sound scene under- standing and hands-free speech recognition,

    S. Nakamura, K. Hiyane, F. Asano, and T. Endo, “Acoustic al sound database in real environments for sound scene under- standing and hands-free speech recognition,” Proc. Language Resources and Evaluation Conference (LREC), pp. 965–968, 2000

  17. [25]

    https://www.ksuke.net/demos

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.