Skip-modal generation translates images to speech without paired image-speech data by using text as a shared information bottleneck between disjoint image-text and text-speech datasets.
One-sided unsupervised do- main mapping
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck
Skip-modal generation translates images to speech without paired image-speech data by using text as a shared information bottleneck between disjoint image-text and text-speech datasets.