REVIEW 4 cited by
Foley Sound Synthesis at the DCASE 2023 Challenge
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The addition of Foley sound effects during post-production is a common technique used to enhance the perceived acoustic properties of multimedia content. Traditionally, Foley sound has been produced by human Foley artists, which involves manual recording and mixing of sound. However, recent advances in sound synthesis and generative models have generated interest in machine-assisted or automatic Foley synthesis techniques. To promote further research in this area, we have organized a challenge in DCASE 2023: Task 7 - Foley Sound Synthesis. Our challenge aims to provide a standardized evaluation framework that is both rigorous and efficient, allowing for the evaluation of different Foley synthesis systems. We received 17 submissions, and performed both objective and subjective evaluation to rank them according to three criteria: audio quality, fit-to-category, and diversity. Through this challenge, we hope to encourage active participation from the research community and advance the state-of-the-art in automatic Foley synthesis. In this technical report, we provide a detailed overview of the Foley sound synthesis challenge, including task definition, dataset, baseline, evaluation scheme and criteria, challenge result, and discussion.
Forward citations
Cited by 4 Pith papers
-
Doppelganger: Sound Effects and Their Synthetic Twins
Instance-pair training matches synthetic sound-effect twins to their real sources on unseen events (~80% R@1), while class supervision degrades below the frozen baseline and the mapping stays generator-specific.
-
Enforcing Speech Content Privacy in Environmental Sound Recordings using Segment-wise Waveform Reversal
Segment-wise waveform reversal, combined with voice activity detection and source separation, reduces speech intelligibility in environmental audio (97.9% WER) while preserving source detectability (2.7% SCAD drop) an...
-
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
AudioBERTScore, a training-free metric combining max-norm and p-norm similarity of audio embeddings, correlates more strongly with human subjective scores for text-to-audio synthesis than conventional metrics.
-
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Dual-path attention over time-frequency representations lets a 3.26M-parameter waveform diffusion model match the quality of 50M+ parameter Foley generators.
Discussion (0). Continue with ORCID to comment.