REVIEW 3 major objections 5 minor 25 references
Overview of Tasks and Investigation of Subjective Evaluation Methods in Environmental Sound Synthesis and Conversion
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Intelligibility alone cannot judge synthesized environmental sounds: a whistle identified correctly can still sound unnatural, so the paper recommends always pairing intelligibility tests with distinguishability and naturalness tests.
desk verdict Useful taxonomy of environmental sound synthesis tasks, but the empirical case for multi-dimensional evaluation rests on one class with very few samples and no direct significance test. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object carrying the argument is a conditional WaveNet, an autoregressive generative model that produces raw audio waveforms conditioned on a one-hot sound-event label, because it supplies the synthesized sounds whose quality is under test. The measurement machinery is the three-experiment subjective protocol: Experiment I forces listeners to choose a sound-event label for each sound, Experiment II is an AB preference test asking which of a real and synthesized pair sounds more real, and Experiment III is a five-scale mean opinion score for naturalness. The argument works by comparing performance across these three lenses for the same sound classes; the whistle's high identification score alongside low naturalness is what makes intelligibility look insufficient on its own.
What would settle it
Run the same three subjective tests across several generative models and a larger, more diverse set of sound-event classes. If every class that is correctly identified by listeners also receives a naturalness score close to that of the real recording, or if the whistle is the only class showing the dissociation, then the claim that intelligibility alone is unsatisfactory would not generalize. A cheap first check is spectral analysis: if the whistle's low naturalness comes from missing fine spectral structure, then classes with strong tonal identity but weak noise texture should show the same split, while broadband noisy classes should not.
Extended reading notes
Core claim
The central claim is that intelligibility is not a sufficient subjective measure for environmental sound synthesis: a synthesized sound can be correctly labeled by listeners yet still be perceived as clearly less natural than the real thing. From three listener experiments on sound event synthesis with a conditional WaveNet, the paper reports that average recognition F-scores were 86.22 percent for real sounds and 76.30 percent for synthesized sounds, while listeners identified real sounds as real only 82.71 percent of the time in an AB test, and mean opinion scores for naturalness varied strongly by category. The decisive case is the whistle: listeners named real and synthesized whistles with comparably high accuracy, yet the synthesized whistle's naturalness score was much lower than the real whistle's. The authors infer that a synthesis method that gets the category right but sounds wrong can pass an intelligibility-only evaluation, and they recommend evaluating environmental sound synthesis with intelligibility plus distinguishability and/or naturalness.
Load-bearing premise
The recommendation rests on the assumption that the split between good intelligibility and poor naturalness seen for one sound class, one generative model, one ten-class database, and 24 listeners is a general fact about environmental sound synthesis rather than a quirk of that particular setup.
Editorial extensions
If this is right
- Future comparisons of environmental sound synthesis methods should report at least one perceptual metric beyond label intelligibility, or a method that reproduces only category-level cues can appear artificially competitive.
- For applications that use synthesized sound directly, such as film and game production and virtual reality, naturalness and indistinguishability from real recordings become the deciding quality measures, whereas data augmentation for detection may continue to rely primarily on intelligibility.
- The same three-test protocol transfers naturally to sound scene synthesis and to sound event and scene conversion, where no standard subjective evaluation currently exists.
- Until an objective metric for environmental sound quality is validated, subjective multi-test evaluation remains necessary; the speech and audio objective metrics cited in the paper do not cover this domain.
- The failure case for intelligibility-only evaluation is not marginal: a category such as whistle can pass identification almost perfectly while failing naturalness by a large margin.
Reading between the lines
- Extending the paper's logic, an automatic classifier-based intelligibility score would likely show the same blind spot, over-rating tonal classes such as whistles; a robust objective metric for environmental sound would probably need two components, classifiability plus a distributional naturalness measure.
- The spectrogram evidence suggests a testable hypothesis: the intelligibility-naturalness gap should be largest for sounds whose identity lives in narrow spectral bands while their naturalness depends on fine stochastic texture, and smallest for broadband noisy sounds.
- The paper stops short of proposing a single combined subjective score; a practical next step would be to define a gated reporting convention, such as reporting intelligibility only for sounds that first pass a naturalness threshold, so results across papers become comparable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reviews the emerging area of environmental sound synthesis and conversion, organizing it into four task families: synthesis from sound event/scene labels, synthesis from onomatopoeic words, conversion between environmental sounds, and synthesis/conversion from multimedia inputs. It then reports a subjective evaluation study of a conditional WaveNet system for sound event synthesis using ten classes from RWCP-SSD. Three experiments were conducted with 24 listeners: a forced-choice intelligibility test (classification recall), an AB distinguishability test between real and synthesized sounds, and a five-scale MOS naturalness test. The intelligibility results show an average F-score drop from 86.22% (real) to 76.30% (synthesized); the AB test shows listeners identify real sounds with 82.71% accuracy; and the MOS results show generally lower naturalness for synthesized sounds, with particularly large gaps for the electric shaver, trash box, and whistle. On the basis of the whistle class, where intelligibility is high but naturalness is much lower for synthesized sounds, the paper proposes that environmental sound synthesis should be evaluated not only by intelligibility but also by distinguishability and/or naturalness.
Significance. If the central claim is accepted, the paper provides a useful practical guideline for a field that currently lacks standardized evaluation protocols. The review portion is valuable in itself: it gives the first compact taxonomy of environmental sound synthesis and conversion tasks, links them to applications such as data augmentation and media production, and identifies the absence of established subjective and objective evaluation methods. The paper also ships demo audio, which supports reproducibility of the perceptual judgments. The proposed three-dimensional evaluation (intelligibility, distinguishability, naturalness) is plausible and aligns with practices in speech and music synthesis. However, the empirical foundation for the core methodological recommendation is narrow: the dissociation between intelligibility and naturalness rests on a single sound class, a small number of stimuli, and no inferential statistics. The strength of the recommendation therefore currently exceeds the strength of the evidence.
major comments (3)
- [Sec. 3.2, Fig. 9, Table 2] The central recommendation is supported by exactly one empirical dissociation: the whistle class shows high intelligibility (93.3% recall for synthesized sounds in Fig. 6 and 95.8% for real sounds in Fig. 5) but a large apparent drop in naturalness MOS in Fig. 9. This comparison rests on only two whistle samples per condition (Table 2), 24 listeners per cell, and no significance test or confidence interval for the whistle MOS gap. If that gap is sampling noise or an artifact of the single synthesized whistle file used, the paper contains no remaining direct evidence that intelligibility alone is insufficient for evaluating environmental sound synthesis. Please report per-condition confidence intervals or a paired test for the whistle class, or explicitly reframe the proposal as a preliminary observation rather than a general methodological conclusion.
- [Sec. 3.2, Experiment II] Experiment II shows that listeners identify real versus synthesized sounds with 82.71% accuracy, which indicates that synthesized sounds are not indistinguishable from real sounds. However, this result is not connected to the intelligibility scores per class, so it does not by itself show that intelligibility testing is insufficient. To support the insufficiency claim, the paper needs a per-class comparison showing high intelligibility coexisting with poor distinguishability or naturalness; Experiment II as reported does not provide that. Please present the AB results alongside the intelligibility results, or analyze the joint per-class outcomes, so that the distinguishability dimension can actually be compared with intelligibility.
- [Sec. 3.1, Table 2] The experiments use only 24 listeners, ten labels, and two to five samples per label per condition, and all comparisons are reported as point estimates without uncertainty quantification. With this sample size, the absence of a large difference for classes such as coffee grinder, clock, and maracas in Fig. 9 cannot be interpreted as equivalence between real and synthesized sounds. Please provide uncertainty quantification for at least the main comparisons, and ideally use more stimuli per class, before drawing general conclusions about which evaluation dimensions are needed for environmental sound synthesis.
minor comments (5)
- [Sec. 3.2] The sentence 'From the results of experiment I, it considered that this subjective test is particularly helpful' should read 'it is considered that'; the current phrasing appears to be a typo.
- [References] Reference [4] lists the year as '20111' instead of '2011'.
- [Fig. 8] The caption 'Recognition rate of real sounds' is ambiguous for a figure reporting the AB test; please clarify that it is the percentage of trials in which listeners correctly identified the real sample.
- [Sec. 2.1] The claim that 'there is no literature giving an overview of the problem definitions and evaluation methods' would be safer as 'to the best of our knowledge'; as written, it overstates certainty about the literature.
- [Sec. 3.1] For the objective evaluation methods listed (PESQ, POLQA, PEAQ), the paper does not mention whether any correlation with subjective naturalness or intelligibility has been shown for environmental sounds; one sentence of context would help readers judge their applicability.
Circularity Check
No significant circularity: the evaluation-method recommendation rests on external subjective measurements, not on fitted or self-referential quantities.
full rationale
The paper's central claim in Sec. 3.2 is that environmental sound synthesis should be evaluated by intelligibility plus distinguishability and/or naturalness. This claim is supported by independently measured listener responses (classification accuracy, AB preference, MOS) on real and WaveNet-synthesized sounds. No parameter is fitted and then renamed as a prediction; the intelligibility/naturalness dissociation (Figs. 5, 6, 9) is an empirical observation, not a quantity defined in terms of the model's training objective. The paper does cite one author's prior overview [11], but that citation is background context and is not load-bearing for the experimental conclusion. The WaveNet configuration is external tooling, and the subjective evaluation results are measured data rather than outputs of a derivation. Consequently, no circular step can be exhibited, and the paper is self-contained with respect to its evaluation-methods proposal.
Assumptions & free parameters
assumptions (4)
- domain assumption RWCP-SSD sound event labels are accurate and representative of the 10 chosen environmental sound classes.
- domain assumption The 24 listeners (13 male, 11 female) are a reliable and generalizable panel for judging environmental sound naturalness.
- domain assumption A 5-point MOS scale yields interval-level data suitable for averaging and comparison with a 95% confidence interval.
- domain assumption Conditional WaveNet is a representative statistical generative model for environmental sound synthesis, so its performance indicates the current state of the art.
Cite this review
Pith. "Pith review of Overview of Tasks and Investigation of Subjective Evaluation Methods in Environmental Sound Synthesis and Conversion." pith.science (2026). https://pith.science/paper/DSU2OSZU
@misc{pith2026190810055,
author = {Pith},
title = {Pith review of: Overview of Tasks and Investigation of Subjective Evaluation Methods in Environmental Sound Synthesis and Conversion},
year = {2026},
howpublished = {\url{https://pith.science/paper/DSU2OSZU}},
note = {Machine review of arXiv:1908.10055}
}
read the original abstract
Synthesizing and converting environmental sounds have the potential for many applications such as supporting movie and game production, data augmentation for sound event detection and scene classification. Conventional works on synthesizing and converting environmental sounds are based on a physical modeling or concatenative approach. However, there are a limited number of works that have addressed environmental sound synthesis and conversion with statistical generative models; thus, this research area is not yet well organized. In this paper, we review problem definitions, applications, and evaluation methods of environmental sound synthesis and conversion. We then report on environmental sound synthesis using sound event labels, in which we focus on the current performance of statistical environmental sound synthesis and investigate how we should conduct subjective experiments on environmental sound synthesis.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION Sound synthesis and conversion are techniques for generating a nat- ural sound using a statistical model that associates input information with the generated sound. Sound synthesis and conversion me thods with the aim of generating speech or music have been widely de - veloped [1, 2, 3]. Recently, some researchers have also deve loped methods...
-
[2]
PROBLEM DEFINITIONS OF ENVIRONMENTAL SOUND SYNTHESIS AND CONVERSION In this section, we review applications, problem definition s, and evaluation methods of environmental sound synthesis and co nver- sion, specifically environmental sound synthesis using eve nt or scene labels (Sec. 2.1), environmental sound synthesis usi ng ono- toNtoNtoNtoNtoN zaaaaaabror...
-
[3]
INVESTIGATION OF SUBJECTIVE EV ALUATION METHOD 3.1. Experimental Conditions In this section, by evaluating SES using sound event labels based on the conditional WaveNet [9], we discuss the current perform ance of environmental sound synthesis and how we should conduct a Predicted label Coffee grinder Cup Clock Whistle Maracas Drum Shaver Trash box Tearing...
-
[4]
CONCLUSION In this paper, we presented the problem definitions of sound e vent synthesis, sound scene synthesis, and sound event and scene con- version. We then discussed the current performance of sound event synthesis and subjective evaluation methods of environmental sound synthesis. The evaluation experiments indicate that sounds synthesized by WaveNet...
-
[5]
Statistical parametric speech synthesis,
H. Zen, K. Tokuda, and A. Black, “Statistical parametric speech synthesis,” Speech Communication, vol. 51, no. 11, pp. 1039–1064, 2009
work page 2009
-
[6]
An overview of voice con- version systems,
S. H. Mohammadi and A. Kain, “An overview of voice con- version systems,” Speech Communication, vol. 88, pp. 65–82, 2017
work page 2017
-
[7]
Deep learning techniques for music generation – a survey,
J. P . Briot, G. Hadjeres, and F. Pachet, “Deep learning techniques for music generation – a survey,” arXiv preprint arXiv:1709.01620, 2017
arXiv 2017
-
[8]
Soun d synthesis for impact sounds in video games,
D. B. Lloyd, N. Raghuvanshi, and N. K. Govindaraju, “Soun d synthesis for impact sounds in video games,” Proc. Sympo- sium on Interactive 3D Graphics and Games. ACM , pp. 55– 61, 20111
Show all 25 references
-
[9]
Visual to sound: Generating natural sound for videos in the wild,
Y . Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg, “Visual to sound: Generating natural sound for videos in the wild,” Proc. IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pp. 3550–3558, 2018
2018
-
[10]
Scaper: A library for soundscape synthesis and aug- mentation,
J. Salamon, D. MacConnell, M. Cartwright, P . Li, and J. P . Bello, “Scaper: A library for soundscape synthesis and aug- mentation,” Proc. IEEE W orkshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp. 344–348, 2017
2017
-
[11]
State of the art in sound texture synthesis,
D. Schwarz, “State of the art in sound texture synthesis, ” Proc. Digital Audio Effects (DAFx), pp. 221–232, 2011
2011
-
[12]
Seed: Resynthesi z- ing environmental sounds from examples,
G. Bernardes, L. Aly, and M. E. Davies, “Seed: Resynthesi z- ing environmental sounds from examples,” Proc. the Sound and Music Computing Conference, pp. 55–62, 2016
2016
-
[13]
WaveNet: A generative model for raw au- dio,
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw au- dio,” arXiv preprint, arXiv:1609.03499, 2016
2016 arXiv
-
[14]
Acoustic scene generation with conditional sam - pleRNN,
Q. Kong, Y . Xu, T. Iqbal, Y . Cao, W. Wang, and M. D. Plumbley, “Acoustic scene generation with conditional sam - pleRNN,” Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 925–929, 2019
2019
-
[15]
Introduction to acoustic event and scene ana lysis,
K. Imoto, “Introduction to acoustic event and scene ana lysis,” Acoustical Science and Technology , vol. 39, no. 3, pp. 182– 188, 2018
2018
-
[16]
SampleRNN: An unconditional end-to-end neural audio generation model,
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. So telo, A. Courville, and Y . Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” Proc. Interna- tional Conference for Learning Representations (ICLR), pp. 1–11, 2017
2017
-
[17]
Perceptual evaluation of speech quality (PESQ): An ob - jective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,
“Perceptual evaluation of speech quality (PESQ): An ob - jective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” ITU-T Recommendation P .862, 2001
2001
-
[18]
Perceptual objective listening quality assessment,
“Perceptual objective listening quality assessment, ” ITU-T Recommendation P .863, 2011
2011
-
[19]
Method for objective measurements of perceived audio qual- ity,
“Method for objective measurements of perceived audio qual- ity,” ITU-R Recommendation BS.1387-1, 2001
2001
-
[20]
Generating sound words from audio signals of acoustic events with sequence-to-sequenc e model,
S. Ikawa and K. Kashino, “Generating sound words from audio signals of acoustic events with sequence-to-sequenc e model,” Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 346–350, 2018
2018
-
[21]
A u- dio style transfer,
E. Grinstein, N. Q. K. Duong, A. Ozerov, and P . P´ erez, “A u- dio style transfer,” Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 586– 590, 2018
2018
-
[22]
Time domain neural audio style transfer,
P . K. Mital, “Time domain neural audio style transfer,” arXiv preprint, p. arXiv:1711.11160, 2017
2017 arXiv
-
[23]
Image style tran s- fer using convolutional neural networks,
L. A. Gatys, A. S. Ecker, and M. Bethge, “Image style tran s- fer using convolutional neural networks,” Proc. IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR), pp. 2414–2423, 2016
2016
-
[24]
Acoustic al sound database in real environments for sound scene under- standing and hands-free speech recognition,
S. Nakamura, K. Hiyane, F. Asano, and T. Endo, “Acoustic al sound database in real environments for sound scene under- standing and hands-free speech recognition,” Proc. Language Resources and Evaluation Conference (LREC), pp. 965–968, 2000
2000
-
[25]
https://www.ksuke.net/demos
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.