Pith. sign in

REVIEW 2 cited by

Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17508 v1 pith:4KQ52OF3 submitted 2024-03-26 cs.SD eess.AS

classification cs.SDeess.AS
keywords audiocorrelationechetembeddingembeddingsenvironmentalperceptualdistance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper explores whether considering alternative domain-specific embeddings to calculate the Fr\'echet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings from VGGish, PANNs, MS-CLAP, L-CLAP, and MERT, which are tailored for either music or environmental sound evaluation. The FAD scores were calculated for sounds from the DCASE 2023 Task 7 dataset. Using perceptual data from the same task, we find that PANNs-WGM-LogMel produces the best correlation between FAD scores and perceptual ratings of both audio quality and perceived fit with a Spearman correlation higher than 0.5. We also find that music-specific embeddings resulted in significantly lower results. Interestingly, VGGish, the embedding used for the original Fr\'echet calculation, yielded a correlation below 0.1. These results underscore the critical importance of the choice of embedding for the FAD metric design.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio

    cs.SD 2025-09 conditional novelty 5.0 of 10

    CoDiCodec unifies continuous and discrete audio compression in one consistency-trained autoencoder, using FSQ-dropout to serve both continuous ~11 Hz embeddings and 2.38 kbps discrete tokens.

  2. The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Word-based style descriptors generated by an LLM can shift MusicGen outputs toward a target artist's sound almost as much as using the artist's name, defining a name-free gap.

Pith tools