Pith. sign in

REVIEW 2 cited by

Images that Sound: Composing Images and Sounds on a Single Canvas

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12221 v3 pith:BIDMVHYE submitted 2024-05-20 cs.CV cs.LGcs.MMcs.SDeess.AS

classification cs.CVcs.LGcs.MMcs.SDeess.AS
keywords imagesspectrogramssoundaudiomodelsnaturalvisualdesired
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these visual spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPS as a Control Signal for Image Generation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.

  2. StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces

    cs.CV 2025-01 conditional novelty 6.0 of 10

    StochSync generates images on arbitrary surfaces such as spheres and meshes by alternating non-overlapping denoised views, maximum stochasticity, and multi-step clean-image prediction from a pretrained diffusion model.

Pith tools