Pith. sign in

Paper Citation Record · LEDGER

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation

As of 11 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2501.05413.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05413 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:16:26.768918Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved19
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e0ce949-7453-44a2-8144-0e6fc08acd99 · outbound

This paper cites Sonicdiffusion: Audio-driven image generation and editing with pretrained diffusion models, 2024.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sonicdiffusion: Audio-driven image generation and editing with pretrained diffusion models, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.298515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.579879Z digest=sha256:b551886340aecd8a04e34ff1a016cda3acb29fffe3c465fa11d95581015bdcb2

Observation 51c6e82d-cf96-4e41-9b47-89d8818e3a38 · outbound

This paper cites Large Scale GAN Training for High Fidelity Natural Image Synthesis.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Large Scale GAN Training for High Fidelity Natural Image Synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.584206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.584206Z digest=sha256:a148a61850346fca696a56c5147c69bbb1254adfc6f39e677d205a9cb2f7b7ab

Observation de48d517-d021-4755-91c2-f2e8b0dc579f · outbound

This paper cites VGGSound: a large-scale audio-visual dataset.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation VGGSound: a large-scale audio-visual dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.287495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.588353Z digest=sha256:36b08bd0dd464ae0b0fc5f6983b2f029cc2eda3ead500d48b919460ec1c4e700

Observation 1a91f36a-f9cf-4161-b5bc-8dbfdfbe2970 · outbound

This paper cites Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.592151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.592151Z digest=sha256:c97922635dc4fbab4848c91a9850d2e6a6cc19c0ae35001ceeffc14f17fe7a1a

Observation c9cc4ab3-7e59-48eb-857f-d456d8816b5b · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.596612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.596612Z digest=sha256:e7234f9a10ef9483303e7ea272f0074045eb84fb6dfea103506ea97bb382d62c

Observation 2915c9fe-2edb-41d3-a317-b3061d508179 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Diffusion models beat gans on image synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.601743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.601743Z digest=sha256:e0b098dd78bf98cc7e570d6e505bab95612a6954bbe84fc1cd32c2a970cb0bd5

Observation 8acffa26-d286-4799-a992-cc41a11010d0 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.605417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.605417Z digest=sha256:b5403e4635fbb7b63017e80f080f1fbffb8d4f089fe81339d9a444dfc00d95f3

Observation 91a535d9-07c5-4ee7-b8d6-43d15d7cb075 · outbound

This paper cites FSD50K: an open dataset of human- labeled sound events.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation FSD50K: an open dataset of human- labeled sound events

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.256510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.610037Z digest=sha256:d3df4f7e1af47ccf746bfc8eb4bb1a7c8dec6b02965a47d5af6aaac9d52d5462

Observation bfcb7571-cf0e-41be-b170-ceecbdb3ad08 · outbound

This paper cites Gemmeke, Daniel P.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Gemmeke, Daniel P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.246523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.613883Z digest=sha256:6b9db1f28ce34ae1ae782f71cd30322d4aa4b8da0324d0f48925fc53acff8ba2

Observation 915a2ef4-f076-42fe-bb3f-acaca5ad40ea · outbound

This paper cites ImageBind: One embedding space to bind them all.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation ImageBind: One embedding space to bind them all

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.236239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.618571Z digest=sha256:07155ab94b409a2b2c12bdb6d8cbc3c4d06269469ff834dff0ca8120aa6318c0

Observation c3773c92-a913-485b-915a-94ded1620be6 · outbound

This paper cites AST: Audio Spectrogram Transformer.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation AST: Audio Spectrogram Transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.225985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.622284Z digest=sha256:40f851bee3c312e0738c986703d322e18ea570d9ce322ac46bcc8df19446d7d2

Observation e83ae7af-5b91-449b-b725-2a8c9f8258da · outbound

This paper cites Generative adversarial nets.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Generative adversarial nets

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.215721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.626061Z digest=sha256:638512b32ff2b4a3783dfd46dd5e475bbecf753d5e7d9affa7f25b57c5c852c3

Observation 461deba7-c29c-479d-95ec-c62415e66ff7 · outbound

This paper cites Grimm and M.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Grimm and M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.204770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.629469Z digest=sha256:c4fe15cbc10982bac45f6b845d60e47abfac8e6f54c33a1dd448992a47e48164

Observation 8bc2b167-cdc3-4560-86a5-35d896fdeca8 · outbound

This paper cites AudioCLIP: Extending clip to image, text and au- dio.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation AudioCLIP: Extending clip to image, text and au- dio

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.194146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.633300Z digest=sha256:0180729b19a2790103e41eecbc508ac486af3625c5fd888bdee7afd1fdf30c5a

Observation d42f5e3a-e96b-4aa5-98b4-791e12af9d4c · outbound

This paper cites CLIPScore: a reference-free evaluation met- ric for image captioning.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation CLIPScore: a reference-free evaluation met- ric for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.176715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.636556Z digest=sha256:59688dcfd3e10629b4b14a1ec2d61d63f7624e3690dbd4c4def22677c5b0cdac

Observation 80962813-c05f-40a3-a72b-bbf6a2ca8afd · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.165470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.639871Z digest=sha256:fe38d47a10b36f9d63725b9270bee4faffee61805dd86aa92dbaa7c5e41b62f8

Observation eaffa1b6-079a-4d9d-b482-5dfe65cf768a · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Denoising Diffusion Probabilistic Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.643318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.643318Z digest=sha256:4c6f64d794ba339f9dbdfe87cc7718bece36b4112cb9dea94913374b57b44026

Observation 247e85c9-55b0-4a2e-a0fe-3461467556ee · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.153728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.649514Z digest=sha256:257eb165ed77650d94a4567c252d06b0f67e95af9d415c15757143e90fbd3201

Observation 1ff94be0-c2b8-4387-adef-1e50f4cd0396 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation A style-based generator architecture for generative adversarial networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.652448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.652448Z digest=sha256:363a1f668ea02653ea754a0f54a9bc0ab9afa40d4435ae4fc5a7f89e4084689e

Observation 6fe3fc34-2531-4a42-8bb6-012c2581de34 · outbound

This paper cites Segment Anything.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Segment Anything

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.655699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.655699Z digest=sha256:06aad18b69f734ead622d9cbe891352b0a3c29042c4b4e62f357c2e366f860da

Observation 3e05c5c7-05f4-44fa-85c1-1db526b22584 · outbound

This paper cites Sound-guided se- mantic video generation.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sound-guided se- mantic video generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.134051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.660376Z digest=sha256:7d21e497d1a8afef0304c234022deeb32e2d2ad2e96f9cb1ff31841b8964bd25

Observation ebb52430-1dd1-4f58-a47b-3b241faa0c1d · outbound

This paper cites Sound-guided semantic image manipulation.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sound-guided semantic image manipulation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.123884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.664064Z digest=sha256:283765e6b92e2755d331eeee6108acfb3b49ff94f922d6e823f7b946ce790040

Observation f5feafd1-08b8-49ea-86db-abf5d3569b32 · outbound

This paper cites Learning visual styles from audio-visual associations.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Learning visual styles from audio-visual associations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.113292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.667587Z digest=sha256:bfbd0a7b907b3d709511a1369f89adecff55ad82fb796e25c9f3ee451cdb0bb8

Observation bdf015c1-8c24-4cd0-ae44-77b5f3a67339 · outbound

This paper cites Visual instruction tuning, 2023.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Visual instruction tuning, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.091893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.674856Z digest=sha256:bd332c899fac3018e7b85192a5ee073687c01c52d263efb06cb6759d86360650

Observation ec476676-e1f0-4e16-93e0-cf1d1fb3ceff · outbound

This paper cites Decoupled weight de- cay regularization.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Decoupled weight de- cay regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.678345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.678345Z digest=sha256:b5bbdd27df1e83347bb45a63a9c1866e73cb3b6038a7218255c1d290ffa4c3f3

Observation fca77de4-d030-4318-94a9-aeb9f283cd2c · outbound

This paper cites Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.682091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.682091Z digest=sha256:4ecb0313be7b2ec7f448620bcfcbc6f0db9f9e8d2291d6d79dead1286d2f002c

Observation 902eb259-a633-4b7f-b3dd-d26530a70052 · outbound

This paper cites Visually indicated sounds.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Visually indicated sounds

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.066622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.686045Z digest=sha256:9ac8215fdd5b72f6d1a5cb0f9198cbd9a0356d2d2f92b0fc6fac407963a2d74a

Observation c8c983b4-a2a2-4ae2-95d9-161313f51227 · outbound

This paper cites Jour- neydb: A benchmark for generative image understanding,.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Jour- neydb: A benchmark for generative image understanding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.056429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.689570Z digest=sha256:5050c7b13e47d72b65e042324ff60e0348fdc935c472dcf0e1c1168739b43861

Observation 0789e7e5-735c-4ddf-892c-10e692251907 · outbound

This paper cites Scalable Diffusion Models with Transformers.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Scalable Diffusion Models with Transformers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.693041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.693041Z digest=sha256:8059268adc90848327ac0436143bd327ca02784bba3fa897cc28acc6ade5ee8e

Observation d0462784-de8e-46a8-9b5f-8bf867faed95 · outbound

This paper cites Glue- gen: Plug and play multi-modal encoders for x-to-image generation.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Glue- gen: Plug and play multi-modal encoders for x-to-image generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.046632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.696421Z digest=sha256:849bde0c0c179c96b601fac531797c156ba65b2637ced02573027daae2f9b298

Observation 6a786f6a-2f84-48f6-98ef-ce5c4f9443e9 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Learn- ing transferable visual models from natural language super- vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.699258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.699258Z digest=sha256:3b7d4021803f700a1d962e38f9a00c4b3f8ded3ec0f6c9c36e30f8f6c35f45f2

Observation 70fbd628-ff8d-42d7-9f89-849993071eb1 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.701912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.701912Z digest=sha256:2f99d0f19931b9d8d6d52a91dcc9fe6fce534b167bf89f4232a1f25623d21224

Observation 37e19563-1eb6-465d-97da-377aac6caff7 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.704878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.704878Z digest=sha256:55754894313b778fc35c86598ef0df0087ca2b8d625edab39c4b1908d040d406

Observation 805f24c7-3211-4cea-b0d2-456b7215765b · outbound

This paper cites Generative ad- versarial text to image synthesis.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Generative ad- versarial text to image synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.708154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.708154Z digest=sha256:02cc0f0407422382657996a58e1eccf1a0b08d48e9fc82895e9faa5de869e1fd

Observation 276d7aaf-12b1-4957-8208-92e4dbe5220b · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation High-resolution image syn- thesis with latent diffusion models, 2021

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.016208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.711135Z digest=sha256:f177786db69b2c80728e07ea8532e114808fffadef37ac94a29aaa38041704cb

Observation e3f9f8d5-a9f7-47fe-a678-057557bc89be · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:27.004837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.713972Z digest=sha256:410a92004fd45742ff7c6f92c6d66e070d0be01f808e8c228c57b5a4e66b0a4e

Observation f978e68f-69cd-46f1-a0e4-5f31d37a70fa · outbound

This paper cites Sound to visual scene genera- tion by audio-to-visual latent alignment, 2023.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sound to visual scene genera- tion by audio-to-visual latent alignment, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.992716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.716787Z digest=sha256:912f1a8e5c344cb570662360144ec07b2812455ec73959d4268088ada49c20ea

Observation eb488697-96ad-4800-a45b-791b85430115 · outbound

This paper cites Llama: Open and efficient foundation lan- guage models, 2023.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Llama: Open and efficient foundation lan- guage models, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.719960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.719960Z digest=sha256:4697e8784f82efc341945d4e457d77fd8c01f155dd40bb1b18a8efaab7cefedd

Observation d9f5eef9-6e34-4d87-aca7-28147d11c386 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2023.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Cogvlm: Visual expert for pretrained language models, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.973331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.723245Z digest=sha256:ce98d62a6b8e561e6ee37232e4000dc048ecc9629544332ecbf0ebcedcdb131b

Observation b9069101-be2a-426c-bdd8-6ff662bafb61 · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Wav2clip: Learning robust audio repre- sentations from clip

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.963731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.726406Z digest=sha256:4f4d4fefe13169c872566b203efa25a4408559f4ba115f56c367bdb82df76684

Observation bbfce54c-fc8b-49b7-917e-9dc2226e48f3 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.952789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.730079Z digest=sha256:0077a570bef1c46be322daa6d92f2da5b395970ba6050ba67f96d3053108da4b

Observation 4c888a2a-a2f4-4517-8a20-b6fd97faa436 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.942917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.733770Z digest=sha256:23a704c1a67adcd87f212681251c0b35167ae0eba19c4e059d145444dc6f5b93

Observation dcc7da23-d922-44ce-8387-f7246c722354 · outbound

This paper cites AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:26.744477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:26.744477Z digest=sha256:587036bf5f8db0f3b86864753890af40f684f90280005d03b330162b2bd49630

Observation 32fafa2e-50aa-4b0f-8552-400fb4a89113 · outbound

This paper cites A survey on segment anything model (sam): Vision foundation model meets prompt engineering, 2023.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation A survey on segment anything model (sam): Vision foundation model meets prompt engineering, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.932575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.748798Z digest=sha256:c06ae91f7306eb34ad65e34d6dabd88677915f32f4b48a6aeb5fc2fad1d6bf3b

Observation bd876b26-52e0-4aa5-aafb-19fcd4333f1e · outbound

This paper cites Stack- gan: Text to photo-realistic image synthesis with stacked generative adversarial networks.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Stack- gan: Text to photo-realistic image synthesis with stacked generative adversarial networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.921773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.751981Z digest=sha256:1807c0335755c861001c92fe27eb9200be1ee1b1c391fed014db518389ce9752

Observation 813ef01c-aff7-444a-9d34-372b856a9da0 · outbound

This paper cites an unresolved cited work.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:16:26.910331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.755315Z digest=sha256:6d90e713fb8a225f9aaad693f2cbd5f718ad065f7089e6aa97e772c9bbbdf12e

Observation aa260c62-1724-44fc-98e4-cf7a5503758a · outbound

This paper cites an unresolved cited work.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:16:26.896762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.759653Z digest=sha256:c28e619d86ed4ff1fdb8a7e244aa76e058106943ef157da0cfa335e935e531ba

Observation dda650b2-e859-48cc-b24c-9d80f3a81fb5 · outbound

This paper cites , Building - The sound of footsteps on the pavement.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation , Building - The sound of footsteps on the pavement

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.885533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.764660Z digest=sha256:87f9e2ee21735114060ba0fa31a5b5d1f53f515557813eb242a25b1c5af382c2

Observation 50a0f645-8c4c-4fae-88e0-8ff552216d47 · outbound

This paper cites Pre-trained.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Pre-trained

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:16:26.873719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.768918Z digest=sha256:12fc7d4dc1964a5fe100b4d29a70e206791fef8090445028ff9375ccc6ceba1a

Observation dfa80e39-a3e8-4550-b529-2d59254383e0 · outbound

This paper cites an unresolved cited work.

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Unresolved cited work

Reference 252

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T21:16:27.101607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:16:26.671353Z digest=sha256:7e536087167604e6d33dde34c6f895f6604220e1aad3ad17d7c31d89ba5368b6

Pith citing papers

No inbound Pith citation observations are available.