A single-stage model performs zero-shot voice conversion from silent lip video and target face images, with no acoustic input at inference.
CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Mutual information (MI) minimization has gained considerable interests in various machine learning tasks. However, estimating and minimizing MI in high-dimensional spaces remains a challenging problem, especially when only samples, rather than distribution forms, are accessible. Previous works mainly focus on MI lower bound approximation, which is not applicable to MI minimization problems. In this paper, we propose a novel Contrastive Log-ratio Upper Bound (CLUB) of mutual information. We provide a theoretical analysis of the properties of CLUB and its variational approximation. Based on this upper bound, we introduce a MI minimization training scheme and further accelerate it with a negative sampling strategy. Simulation studies on Gaussian distributions show the reliable estimation ability of CLUB. Real-world MI minimization experiments, including domain adaptation and information bottleneck, demonstrate the effectiveness of the proposed method. The code is at https://github.com/Linear95/CLUB.
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MuteSwap: Visual-informed Silent Video Identity Conversion
A single-stage model performs zero-shot voice conversion from silent lip video and target face images, with no acoustic input at inference.