REVIEW 5 cited by
Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper explores whether considering alternative domain-specific embeddings to calculate the Fr\'echet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings from VGGish, PANNs, MS-CLAP, L-CLAP, and MERT, which are tailored for either music or environmental sound evaluation. The FAD scores were calculated for sounds from the DCASE 2023 Task 7 dataset. Using perceptual data from the same task, we find that PANNs-WGM-LogMel produces the best correlation between FAD scores and perceptual ratings of both audio quality and perceived fit with a Spearman correlation higher than 0.5. We also find that music-specific embeddings resulted in significantly lower results. Interestingly, VGGish, the embedding used for the original Fr\'echet calculation, yielded a correlation below 0.1. These results underscore the critical importance of the choice of embedding for the FAD metric design.
Forward citations
Cited by 5 Pith papers
-
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
Smooth-Foley uses frame-level visual features and label-guided temporal conditions to generate continuous, synchronized audio for videos with moving or ambiguous sound sources.
-
CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
CoDiCodec unifies continuous and discrete audio compression in one consistency-trained autoencoder, using FSQ-dropout to serve both continuous ~11 Hz embeddings and 2.38 kbps discrete tokens.
-
The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation
Word-based style descriptors generated by an LLM can shift MusicGen outputs toward a target artist's sound almost as much as using the artist's name, defining a name-free gap.
-
Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation
Injecting random noise into input embeddings during training lets purely autoregressive models generate continuous audio embeddings without quality degradation over long sequences.
-
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
FMD applies the Frechet distance to CLaMP music embeddings to quantify distributional similarity between generated and reference symbolic music.
Discussion (0). Continue with ORCID to comment.