Pith. sign in

REVIEW 2 cited by

Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17508 v1 pith:4KQ52OF3 submitted 2024-03-26 cs.SD eess.AS

Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

classification cs.SD eess.AS
keywords audiocorrelationechetembeddingembeddingsenvironmentalperceptualdistance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This paper explores whether considering alternative domain-specific embeddings to calculate the Fr\'echet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings from VGGish, PANNs, MS-CLAP, L-CLAP, and MERT, which are tailored for either music or environmental sound evaluation. The FAD scores were calculated for sounds from the DCASE 2023 Task 7 dataset. Using perceptual data from the same task, we find that PANNs-WGM-LogMel produces the best correlation between FAD scores and perceptual ratings of both audio quality and perceived fit with a Spearman correlation higher than 0.5. We also find that music-specific embeddings resulted in significantly lower results. Interestingly, VGGish, the embedding used for the original Fr\'echet calculation, yielded a correlation below 0.1. These results underscore the critical importance of the choice of embedding for the FAD metric design.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio

    cs.SD 2025-09 conditional novelty 5.0

    CoDiCodec unifies continuous and discrete audio compression in one consistency-trained autoencoder, using FSQ-dropout to serve both continuous ~11 Hz embeddings and 2.38 kbps discrete tokens.

  2. The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation

    cs.SD 2025-08 conditional novelty 5.0

    Word-based style descriptors generated by an LLM can shift MusicGen outputs toward a target artist's sound almost as much as using the artist's name, defining a name-free gap.