Polyphonia improves zero-shot stem-specific timbre transfer in polyphonic music by 15.5% target alignment via acoustic-informed attention calibration that uses probabilistic priors to set coarse boundaries.
arXiv preprint arXiv:2502.15602 , year=
5 Pith papers cite this work. Polarity classification is still indexing.
abstract
Although being widely adopted for evaluating generated audio signals, the Fr\'echet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and high computational complexity. As an alternative, we introduce the Kernel Audio Distance (KAD), a novel, distribution-free, unbiased, and computationally efficient metric based on Maximum Mean Discrepancy (MMD). Through analysis and empirical validation, we demonstrate KAD's advantages: (1) faster convergence with smaller sample sizes, enabling reliable evaluation with limited data; (2) lower computational cost, with scalable GPU acceleration; and (3) stronger alignment with human perceptual judgments. By leveraging advanced embeddings and characteristic kernels, KAD captures nuanced differences between real and generated audio. Open-sourced in the kadtk toolkit, KAD provides an efficient, reliable, and perceptually aligned benchmark for evaluating generative audio models.
citation-role summary
citation-polarity summary
years
2026 5roles
background 1polarities
background 1representative citing papers
Four-step SR-FD fine-tuning cuts VoxCPM2's Seed-TTS English WER from 2.23% to 1.41%, beating the ten-step baseline at 1.74% with no inference-time cost.
PianoKontext renders expressive piano performances from deadpan scores using flow matching in Music2Latent latent space with DTW alignment for paired training data.
OTAD replaces FAD's frozen embedding cost and Gaussian coupling with a learned Riemannian adapter and Sinkhorn OT, reporting higher MOS Spearman correlation on DCASE 2023 and per-sample diagnostics with AUROC at least 0.86.
The paper defines three desiderata (responsiveness, smoothness, symmetry) and empirically compares FAD, intensity vectors, and acoustic maps on controlled FOA scenes of increasing complexity.
citing papers explorer
-
Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
Polyphonia improves zero-shot stem-specific timbre transfer in polyphonic music by 15.5% target alignment via acoustic-informed attention calibration that uses probabilistic priors to set coarse boundaries.
-
Fr\'echet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Four-step SR-FD fine-tuning cuts VoxCPM2's Seed-TTS English WER from 2.23% to 1.41%, beating the ten-step baseline at 1.74% with no inference-time cost.
-
PianoKontext: Expressive Performance Rendering from Deadpan Context
PianoKontext renders expressive piano performances from deadpan scores using flow matching in Music2Latent latent space with DTW alignment for paired training data.
-
Optimal Transport Audio Distance with Learned Riemannian Ground Metrics
OTAD replaces FAD's frozen embedding cost and Gaussian coupling with a learned Riemannian adapter and Sinkhorn OT, reporting higher MOS Spearman correlation on DCASE 2023 and per-sample diagnostics with AUROC at least 0.86.
-
Sensitivity Analysis of Generative Spatial Audio Metrics: A Study on Responsiveness, Smoothness, and Symmetry
The paper defines three desiderata (responsiveness, smoothness, symmetry) and empirically compares FAD, intensity vectors, and acoustic maps on controlled FOA scenes of increasing complexity.