Pith. sign in

REVIEW 1 cited by

Tr\"aumerAI: Dreaming Music with StyleGAN

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.04680 v1 pith:QR7A3XX6 submitted 2021-02-09 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords musicaudiovideoaumeraiembeddingembeddingsexamplesgoal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The goal of this paper to generate a visually appealing video that responds to music with a neural network so that each frame of the video reflects the musical characteristics of the corresponding audio clip. To achieve the goal, we propose a neural music visualizer directly mapping deep music embeddings to style embeddings of StyleGAN, named Tr\"aumerAI, which consists of a music auto-tagging model using short-chunk CNN and StyleGAN2 pre-trained on WikiArt dataset. Rather than establishing an objective metric between musical and visual semantics, we manually labeled the pairs in a subjective manner. An annotator listened to 100 music clips of 10 seconds long and selected an image that suits the music among the 200 StyleGAN-generated examples. Based on the collected data, we trained a simple transfer function that converts an audio embedding to a style embedding. The generated examples show that the mapping between audio and video makes a certain level of intra-segment similarity and inter-segment dissimilarity.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio-Sync Video Generation with Multi-Stream Temporal Control

    cs.CV 2025-06 reject novelty 6.0 of 10

    MTV splits audio into speech, effects, and music to separately drive lip sync, event timing, and visual mood in video generation, trained on a new 392K-clip dataset.

Pith tools