A VLM-guided retrieval pipeline sonifies images with matching audio, and a PixArt-alpha diffusion model trained on these synthetic pairs is competitive with state-of-the-art audio-to-image generators.
Sonicdiffusion: Audio-driven image generation and editing with pretrained diffusion models, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation
A VLM-guided retrieval pipeline sonifies images with matching audio, and a PixArt-alpha diffusion model trained on these synthetic pairs is competitive with state-of-the-art audio-to-image generators.