Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:16:26.768918Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2501.05413.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:16:26.768918Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5e0ce949-7453-44a2-8144-0e6fc08acd99 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sonicdiffusion: Audio-driven image generation and editing with pretrained diffusion models, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51c6e82d-cf96-4e41-9b47-89d8818e3a38 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Large Scale GAN Training for High Fidelity Natural Image Synthesis
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de48d517-d021-4755-91c2-f2e8b0dc579f · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation VGGSound: a large-scale audio-visual dataset
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a91f36a-f9cf-4161-b5bc-8dbfdfbe2970 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cc4ab3-7e59-48eb-857f-d456d8816b5b · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2915c9fe-2edb-41d3-a317-b3061d508179 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Diffusion models beat gans on image synthesis
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8acffa26-d286-4799-a992-cc41a11010d0 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation An image is worth 16x16 words: Transformers for image recognition at scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a535d9-07c5-4ee7-b8d6-43d15d7cb075 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation FSD50K: an open dataset of human- labeled sound events
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bfcb7571-cf0e-41be-b170-ceecbdb3ad08 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Gemmeke, Daniel P
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 915a2ef4-f076-42fe-bb3f-acaca5ad40ea · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation ImageBind: One embedding space to bind them all
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c3773c92-a913-485b-915a-94ded1620be6 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation AST: Audio Spectrogram Transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e83ae7af-5b91-449b-b725-2a8c9f8258da · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Generative adversarial nets
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 461deba7-c29c-479d-95ec-c62415e66ff7 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Grimm and M
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8bc2b167-cdc3-4560-86a5-35d896fdeca8 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation AudioCLIP: Extending clip to image, text and au- dio
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d42f5e3a-e96b-4aa5-98b4-791e12af9d4c · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation CLIPScore: a reference-free evaluation met- ric for image captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 80962813-c05f-40a3-a72b-bbf6a2ca8afd · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eaffa1b6-079a-4d9d-b482-5dfe65cf768a · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Denoising Diffusion Probabilistic Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247e85c9-55b0-4a2e-a0fe-3461467556ee · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ff94be0-c2b8-4387-adef-1e50f4cd0396 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation A style-based generator architecture for generative adversarial networks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe3fc34-2531-4a42-8bb6-012c2581de34 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Segment Anything
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e05c5c7-05f4-44fa-85c1-1db526b22584 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sound-guided se- mantic video generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ebb52430-1dd1-4f58-a47b-3b241faa0c1d · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sound-guided semantic image manipulation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f5feafd1-08b8-49ea-86db-abf5d3569b32 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Learning visual styles from audio-visual associations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bdf015c1-8c24-4cd0-ae44-77b5f3a67339 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Visual instruction tuning, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec476676-e1f0-4e16-93e0-cf1d1fb3ceff · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Decoupled weight de- cay regularization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca77de4-d030-4318-94a9-aeb9f283cd2c · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 902eb259-a633-4b7f-b3dd-d26530a70052 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Visually indicated sounds
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8c983b4-a2a2-4ae2-95d9-161313f51227 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Jour- neydb: A benchmark for generative image understanding,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0789e7e5-735c-4ddf-892c-10e692251907 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Scalable Diffusion Models with Transformers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0462784-de8e-46a8-9b5f-8bf867faed95 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Glue- gen: Plug and play multi-modal encoders for x-to-image generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a786f6a-2f84-48f6-98ef-ce5c4f9443e9 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Learn- ing transferable visual models from natural language super- vision
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70fbd628-ff8d-42d7-9f89-849993071eb1 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e19563-1eb6-465d-97da-377aac6caff7 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 805f24c7-3211-4cea-b0d2-456b7215765b · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Generative ad- versarial text to image synthesis
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276d7aaf-12b1-4957-8208-92e4dbe5220b · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation High-resolution image syn- thesis with latent diffusion models, 2021
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e3f9f8d5-a9f7-47fe-a678-057557bc89be · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f978e68f-69cd-46f1-a0e4-5f31d37a70fa · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Sound to visual scene genera- tion by audio-to-visual latent alignment, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb488697-96ad-4800-a45b-791b85430115 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Llama: Open and efficient foundation lan- guage models, 2023
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f5eef9-6e34-4d87-aca7-28147d11c386 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Cogvlm: Visual expert for pretrained language models, 2023
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9069101-be2a-426c-bdd8-6ff662bafb61 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Wav2clip: Learning robust audio repre- sentations from clip
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bbfce54c-fc8b-49b7-917e-9dc2226e48f3 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4c888a2a-a2f4-4517-8a20-b6fd97faa436 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Attngan: Fine- grained text to image generation with attentional generative adversarial networks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dcc7da23-d922-44ce-8387-f7246c722354 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32fafa2e-50aa-4b0f-8552-400fb4a89113 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation A survey on segment anything model (sam): Vision foundation model meets prompt engineering, 2023
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd876b26-52e0-4aa5-aafb-19fcd4333f1e · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Stack- gan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 813ef01c-aff7-444a-9d34-372b856a9da0 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa260c62-1724-44fc-98e4-cf7a5503758a · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dda650b2-e859-48cc-b24c-9d80f3a81fb5 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation , Building - The sound of footsteps on the pavement
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 50a0f645-8c4c-4fae-88e0-8ff552216d47 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Pre-trained
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dfa80e39-a3e8-4550-b529-2d59254383e0 · outbound
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation Unresolved cited work
Reference 252
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.