REVIEW 8 cited by
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Although being widely adopted for evaluating generated audio signals, the Fr\'echet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and high computational complexity. As an alternative, we introduce the Kernel Audio Distance (KAD), a novel, distribution-free, unbiased, and computationally efficient metric based on Maximum Mean Discrepancy (MMD). Through analysis and empirical validation, we demonstrate KAD's advantages: (1) faster convergence with smaller sample sizes, enabling reliable evaluation with limited data; (2) lower computational cost, with scalable GPU acceleration; and (3) stronger alignment with human perceptual judgments. By leveraging advanced embeddings and characteristic kernels, KAD captures nuanced differences between real and generated audio. Open-sourced in the kadtk toolkit, KAD provides an efficient, reliable, and perceptually aligned benchmark for evaluating generative audio models.
Forward citations
Cited by 8 Pith papers
-
InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion
InvFlowFD measures music quality by inverting audio through a flow matching model and computing the distance of the inverted latents to the model's Gaussian prior.
-
Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
Kernel distances on CLaMP3/Aria embeddings score expressive MIDI performances about as well as human listeners and catch contextual corruptions invisible to attribute statistics.
-
RIME: Enabling Large-Scale Agentic Music Post-Production
RIME generates 3,000 synthetic music post-production edit triples and shows that current multimodal LLM agents can recover edit structure but often fail to set effect parameters correctly.
-
HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs
A training-only change to residual vector quantization orders codebook stages from bass to treble, yielding better audio quality and more predictable bitrate scaling with identical inference.
-
Fr\'echet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Four-step SR-FD fine-tuning cuts VoxCPM2's Seed-TTS English WER from 2.23% to 1.41%, beating the ten-step baseline at 1.74% with no inference-time cost.
-
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
Diff-TONE selects the prompt-swap timestep using a distilled latent instrument classifier, improving content preservation over fixed timestep baselines in text-to-music diffusion editing.
-
Can Large Language Models Predict Audio Effects Parameters from Natural Language?
LLMs can predict equalizer and reverb parameters from natural language descriptions, and adding DSP features, DSP function code, and few-shot examples improves the predictions.
-
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Across five text-to-music models, aesthetic predictor scores, pairwise human preferences, and reference-based distribution metrics produce inconsistent rankings, so the choice of evaluation metric changes the winner.
Discussion (0). Continue with ORCID to comment.