Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.381968Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.09529.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.381968Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5499e55c-4442-485c-a0ce-bd274fcb657b · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1016fa-d14f-4487-a833-e02b546b800a · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d3592b-0ec5-49dc-8047-bca52ce43250 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c8da77-0574-4eaa-bfe3-c6507eb04917 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f347c07-c602-4f2c-9057-c1e130551bbe · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7f64e7b-04f9-4c96-b155-f6ad8df9f481 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d258208-eb36-4d96-9d87-f42ae61d879b · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b954811-bcdf-4f68-98ca-badcb11de8cf · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f7bfc2-eef1-4023-80a4-976ee2762f86 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81131255-de76-44bc-adcd-a2201c8416e7 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 407959c6-c6ae-48f5-8ce6-4c796e047570 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen2-Audio Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 445908df-85c5-4b6e-8f1c-177512540198 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a65de6-7316-4f2d-bf81-582b7ea5dc66 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 170ce7ee-8829-44a9-b2ec-0311c63bab55 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab50835-eb3a-444c-92f0-63f2e4bfdd6c · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdfeec41-52cb-4c17-ab5e-f51637f8aae6 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 085e0dd4-a64a-4cbc-9805-7e14955564b2 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b158ec18-97bb-4b2a-bbd7-8c095c677568 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51445caf-5a85-41d0-afb7-0efef22bf28f · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a2c0ef-b839-4460-b0bb-414505746313 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0e26a85-064d-4ffb-bf78-a034169f0d44 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 47d65d10-c76a-4f9e-8db1-8db6158f11fc · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0a176863-2ca2-4c67-8b38-1b3a937649ec · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed524262-779f-452a-aa52-7855b4c1be89 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48d333c0-3f15-4fb8-93cb-f3b4b4fe15c9 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 529399a4-06e1-420c-977a-77791d751c37 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5881b4de-0b83-4946-9684-cdc6ca2e0973 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0580e7e2-d925-4381-b93b-d9ef06abf2a6 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3114ed37-9df1-4fda-b656-cc91ef055d32 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c02742d9-3e0c-45fc-b9a8-24c9ffe28e52 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996eb128-6663-429c-a08d-09ed058974a5 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 285cf977-bf1b-4acb-b96e-aabff348d07e · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c544954d-47fa-487f-901c-17aba8c82d8d · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5435a5d-54f9-42ea-a39f-2c32235422e8 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b61f7b0-894b-45bb-817c-5084ad65406b · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b94af2fb-3928-47b3-abdf-9f2d7a1a054e · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eaed4ba-c82c-49fc-9570-f147d7200303 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be5bf654-8cbe-4ad3-9153-98eccb2f164f · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49af4d98-5205-44f9-89f9-27b18638e561 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ffc7f8-406c-4ebc-8961-f11d36e33b6d · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework TransNet V2: An effective deep network architecture for fast shot transition detection
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ef91ca-ce95-4f76-b966-c5dd5866864d · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ac2da736-ac44-45da-bf7c-b87ffbb70170 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06941a77-2e5c-418b-8df4-ee7edef33f36 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 445351ab-91db-4c6e-b5e2-ffdb4e3c8056 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a761130c-9057-4301-afa8-f83b212722c5 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe144ba-39ec-4501-b027-5122f5c9f762 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0db9894-a99d-4c82-8c84-6694902529e0 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aef01029-75b3-4962-9c29-7563eb6c0a20 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a51c55-39ff-4bbf-bb94-bdbf3ca4cc5a · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3c2a70-445c-49fa-a886-e9efa78515e5 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6173af98-3b2e-4a96-9062-41097d0cc30c · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33e7086b-aa82-4375-83bf-94946de8c5eb · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ceabb6b-f3d7-4e72-bc19-07aedac7bb3a · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ac456d1-b78b-402b-b432-68c7559eafbe · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44e608b-7bd0-49d1-8b02-933e5ccad5ec · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework silent frames
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9bd2df92-e91e-4ca8-aa53-5fa0a9f34f8e · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework the sound of a vehicle driving
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f7eef15a-5a49-4d27-a3b0-31cd6b8af678 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dbdff0f5-4c34-4ab9-b1ad-5ccd20273336 · outbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.