Pith. sign in

REVIEW 1 cited by

Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04906 v1 pith:JMQ4OR4Y submitted 2024-10-07 cs.MM cs.CVcs.SDeess.AS

classification cs.MMcs.CVcs.SDeess.AS
keywords mathcalmusictextitartworksdigitizedmodelsgeneratemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-music models are limited to simple images, lacking the capability to generate music from complex digitized artworks. To address this gap, we introduce $\mathcal{A}\textit{rt2}\mathcal{M}\textit{us}$, a novel model designed to create music from digitized artworks or text inputs. $\mathcal{A}\textit{rt2}\mathcal{M}\textit{us}$ extends the AudioLDM~2 architecture, a text-to-audio model, and employs our newly curated datasets, created via ImageBind, which pair digitized artworks with music. Experimental results demonstrate that $\mathcal{A}\textit{rt2}\mathcal{M}\textit{us}$ can generate music that resonates with the input stimuli. These findings suggest promising applications in multimedia art, interactive installations, and AI-driven creative tools.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation

    cs.HC 2024-12 conditional novelty 5.0 of 10

    Vivar uses barycentric interpolation in a pre-trained CLIP embedding space to generate AR visualizations of multi-modal sensor data, with caching that speeds generation 11x.

Pith tools