Pith. sign in

REVIEW 1 cited by

Generating Realistic Images from In-the-wild Sounds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.02405 v1 pith:W3LLJADO submitted 2023-09-05 cs.CV cs.SDeess.AS

classification cs.CVcs.SDeess.AS
keywords imagessoundsoundsaudiogenerateproposewildattention
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous studies have focused on generating images from sound in limited categories or music. In this paper, we propose a novel approach to generate images from in-the-wild sounds. First, we convert sound into text using audio captioning. Second, we propose audio attention and sentence attention to represent the rich characteristics of sound and visualize the sound. Lastly, we propose a direct sound optimization with CLIPscore and AudioCLIP and generate images with a diffusion-based model. In experiments, it shows that our model is able to generate high quality images from wild sounds and outperforms baselines in both quantitative and qualitative evaluations on wild audio datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation

    cs.MM 2025-07 conditional novelty 6.0 of 10

    CatchPhrase improves audio-to-image generation by enriching weak class labels with LLM- and audio-caption-based prompts, filtering and retrieving the best prompt per clip, and training a mapping adapter with contrasti...

Pith tools