Pith. sign in

Paper Citation Record · LEDGER

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

As of 12 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.09529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09529 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.381968Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5499e55c-4442-485c-a0ce-bd274fcb657b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.106949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.106949Z digest=sha256:51ada1402f12db9a390f92fd1b31fc4d04602e35c5be128bb1b4f35441ad4a27

Observation 7e1016fa-d14f-4487-a833-e02b546b800a · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.117045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.117045Z digest=sha256:ae4cee0f8382e5ebbe32951be5382e3a78499ad4135869ff245db9f01efa1b68

Observation 54d3592b-0ec5-49dc-8047-bca52ce43250 · outbound

This paper cites Qwen3-VL Technical Report.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.121620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.121620Z digest=sha256:24ee463447d50d5eaa01ef92310c8ff2089475a8ee31d76be178e90a366d06ec

Observation 77c8da77-0574-4eaa-bfe3-c6507eb04917 · outbound

This paper cites SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.126404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.126404Z digest=sha256:cfe276325708d159178ad42bdac35a3550bf90b2559a2a2deede5bea53b4ed73

Observation 1f347c07-c602-4f2c-9057-c1e130551bbe · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.403358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.131345Z digest=sha256:5aa480500ac4cedb9ab8388173511c26cc08b78932749d9286dfe7498c28445d

Observation c7f64e7b-04f9-4c96-b155-f6ad8df9f481 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.391346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.136526Z digest=sha256:111941458196fa212f29002a719a2d47f6cf7fd2e9b54838be46c7457decb796

Observation 7d258208-eb36-4d96-9d87-f42ae61d879b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.141840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.141840Z digest=sha256:5650695f2ff5fdc196de2bac35716bd59426ca7651562d9310371e9f7e720182

Observation 9b954811-bcdf-4f68-98ca-badcb11de8cf · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.147229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.147229Z digest=sha256:9b9511339f676143b577ba95970a292b387659f1caa51d7713276d4f23a563b0

Observation 81f7bfc2-eef1-4023-80a4-976ee2762f86 · outbound

This paper cites FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.152062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.152062Z digest=sha256:a6758a02c87264af07a671da404afb6e29bb2333604f053a616771bcaa2da755

Observation 81131255-de76-44bc-adcd-a2201c8416e7 · outbound

This paper cites Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.156810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.156810Z digest=sha256:fb1e13752b277ed7f90616f001cac60e1e2ce9ccf5250d74edb45891220c7fe7

Observation 407959c6-c6ae-48f5-8ce6-4c796e047570 · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen2-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.161186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.161186Z digest=sha256:36738494ab7ab4bb2a0f960b68c9541e0bba55e07986638088f3aa33ae18831c

Observation 445908df-85c5-4b6e-8f1c-177512540198 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.165702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.165702Z digest=sha256:c4fddc81237c86d4be35c8dbff89c448e5b1cefa5cbeda4031c945d578ad775b

Observation 44a65de6-7316-4f2d-bf81-582b7ea5dc66 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.371299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.169825Z digest=sha256:dbb03341f96f9dbe9f1d41101f880c4554b008148a76045cf37145499da38b07

Observation 170ce7ee-8829-44a9-b2ec-0311c63bab55 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.174885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.174885Z digest=sha256:434dab3b97d4b74940e4ffd377088e2081a6816b946f0c6f41d58b4a9812164b

Observation 7ab50835-eb3a-444c-92f0-63f2e4bfdd6c · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.178845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.178845Z digest=sha256:dcff70be2023f2798cf00dc5ed5577d3796079b037b3ac21fb189609c072d826

Observation fdfeec41-52cb-4c17-ab5e-f51637f8aae6 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.340192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.182233Z digest=sha256:a6f8cc6dbb080b0bea49392e87ad2a9b0c5854a623bf0f79414b8de6b0e8daec

Observation 085e0dd4-a64a-4cbc-9805-7e14955564b2 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.185572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.185572Z digest=sha256:7e70d682c0b12995c154d84cef9e111ec62225afc4bd77ba58ed7ed6b785e83c

Observation b158ec18-97bb-4b2a-bbd7-8c095c677568 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.191958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.191958Z digest=sha256:2a3cef04c53f245a715c46c6c95637e52f6dae6b3609c1fb59e0fc1c5b90cd96

Observation 51445caf-5a85-41d0-afb7-0efef22bf28f · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.195746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.195746Z digest=sha256:1f8b6eb243a653fdaa63198cfcba116921738d44ec9b4d0544175bce110debe1

Observation 05a2c0ef-b839-4460-b0bb-414505746313 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.310410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.202252Z digest=sha256:20e55b0ee441202b2dfa80fefd1b202a27fc0588cd8fd1d163c48229686fae31

Observation c0e26a85-064d-4ffb-bf78-a034169f0d44 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.296132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.223442Z digest=sha256:1e2181cc93c0ac0d1cc0d8c645860acbedafb79e18125bacf5a2258d6289d5f4

Observation 47d65d10-c76a-4f9e-8db1-8db6158f11fc · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.283319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.227907Z digest=sha256:e6fcdb710d9a671273505fcf413741f5dc7d358b4631d176fb88923c4de3da4e

Observation 0a176863-2ca2-4c67-8b38-1b3a937649ec · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.233176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.233176Z digest=sha256:77607bf46a1069d37bc8d48ce8b63c5d4faaae2c95b9e223c0b3596a99866390

Observation ed524262-779f-452a-aa52-7855b4c1be89 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.261879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.237280Z digest=sha256:0276913ea833f9dc804bee7879bf18a2daa166369a1f3f11a7df5ea298b30015

Observation 48d333c0-3f15-4fb8-93cb-f3b4b4fe15c9 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-11T15:41:34.422290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.241666Z digest=sha256:ad25c1013468627479a55703c8857ee9f80a458a707ae2523c9f735a9de232c3

Observation 529399a4-06e1-420c-977a-77791d751c37 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.248327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.246374Z digest=sha256:78e7b6181c853f8475173082caeff068b2b41316d5f58a66d2ed80dc0242bad3

Observation 5881b4de-0b83-4946-9684-cdc6ca2e0973 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.251401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.251401Z digest=sha256:97b606b01accb6122f42399bbe010a2f02546ffddc67ae03b65510b6efddf569

Observation 0580e7e2-d925-4381-b93b-d9ef06abf2a6 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.256085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.256085Z digest=sha256:f9cd71442a9c9eb98100039f6cd6ebc18e17fac1ab980f022a48a377001e4f7c

Observation 3114ed37-9df1-4fda-b656-cc91ef055d32 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.227361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.260375Z digest=sha256:d4b67afca7e398293b5939b1bb7e91a37340cb29348cd46baf111674139854ae

Observation c02742d9-3e0c-45fc-b9a8-24c9ffe28e52 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.264355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.264355Z digest=sha256:784b293cf7a7ba67e07a3926bad3c34215020446763fe64163722518a437a585

Observation 996eb128-6663-429c-a08d-09ed058974a5 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.203076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.268402Z digest=sha256:6e0856b071fc7174e012a5a833460605cb0329175736deb065ac2a225dc46dd6

Observation 285cf977-bf1b-4acb-b96e-aabff348d07e · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.272510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.272510Z digest=sha256:9c28c345cfe2a45fd002ca7860ffba86a4492db146bb0dad1ae8d3de041d13fb

Observation c544954d-47fa-487f-901c-17aba8c82d8d · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.276452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.276452Z digest=sha256:17e37989e086fc18d6b10dfb0326688ffad6258c22f1767b400c258c444773c9

Observation b5435a5d-54f9-42ea-a39f-2c32235422e8 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.280600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.280600Z digest=sha256:fd16fd841fa85f9e3e28a10aba397e0e521fe5a8eed0309b1beb5aa5e7fc325b

Observation 8b61f7b0-894b-45bb-817c-5084ad65406b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.174848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.284566Z digest=sha256:1bfe70ed5ba34781d1e4acc3f8131c13f638101c363bcf71edea9d46316a7850

Observation b94af2fb-3928-47b3-abdf-9f2d7a1a054e · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.288493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.288493Z digest=sha256:44b0e8fc19f3109dc0c6c191142c7b59d55166102426d400079b9af5be95d1d9

Observation 8eaed4ba-c82c-49fc-9570-f147d7200303 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.152737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.292361Z digest=sha256:ce92963fb6191963448ac6f2a5010ee543ef1da509e2274a1bc5b3b996abbb0c

Observation be5bf654-8cbe-4ad3-9153-98eccb2f164f · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.297516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.297516Z digest=sha256:3e31d7cac24a9ddd59c4392de0c0cc0ffc4d8260a25f0fab7753fe4f521c3383

Observation 49af4d98-5205-44f9-89f9-27b18638e561 · outbound

This paper cites InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.302159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.302159Z digest=sha256:686ed4379de4a86feef5c0edeb54bc7043a28021d12f49a25d21451f7247b1cd

Observation 72ffc7f8-406c-4ebc-8961-f11d36e33b6d · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.307150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.307150Z digest=sha256:90797cd0c730805d699e2d88ab68af6a2b10c1c86793093e474cb2d82535ddd5

Observation b3ef91ca-ce95-4f76-b966-c5dd5866864d · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.137389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.311858Z digest=sha256:8210a383b8af324b8bc6ed6c51b36c216408fedd121040585ccec3c2b932c78c

Observation ac2da736-ac44-45da-bf7c-b87ffbb70170 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.102644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.321229Z digest=sha256:3e3540cdb91f1f399f69eaf6b065a874f04211cf166f09936c0634e934677956

Observation 06941a77-2e5c-418b-8df4-ee7edef33f36 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.325382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.325382Z digest=sha256:1893c046ea0f1ce5466711d08de5e190bcd6a322e044f5ca090d6451d94f8544

Observation 445351ab-91db-4c6e-b5e2-ffdb4e3c8056 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.076228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.329856Z digest=sha256:42ea7d7602d0f318da86536112f83e10f2071af5be136e318c76378e66734165

Observation a761130c-9057-4301-afa8-f83b212722c5 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.334245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.334245Z digest=sha256:ec1a370388f0e1ea1c88dff7ea537761976000de500add86cecbb7edec803ecf

Observation 8fe144ba-39ec-4501-b027-5122f5c9f762 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.050662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.338582Z digest=sha256:a75acc079ddfa84fd2a501a58cbd8cca9f84cfe31983ce36db87e19e7912e538

Observation c0db9894-a99d-4c82-8c84-6694902529e0 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.343280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.343280Z digest=sha256:383a35ae56a56fc49fab427f6d25963b3e1e1c15dbc4aa4743bc0a9cbbd260e7

Observation aef01029-75b3-4962-9c29-7563eb6c0a20 · outbound

This paper cites AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.347508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.347508Z digest=sha256:e7405ef4490836ae3f00eb4fbc6d81087b46c5ea0958738140147064d5d32dfa

Observation e1a51c55-39ff-4bbf-bb94-bdbf3ca4cc5a · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.352108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.352108Z digest=sha256:041dd1757707754e24d749762ab8802a51ad541c8a60bd180ad02bd0cdd1e2ce

Observation 1e3c2a70-445c-49fa-a886-e9efa78515e5 · outbound

This paper cites Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.355920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.355920Z digest=sha256:8c3e07ec1a0e2159d62bfc7e078a24dcf5df8f07ce65f60f3de575d2559ab14e

Observation 6173af98-3b2e-4a96-9062-41097d0cc30c · outbound

This paper cites FoleySpace: Vision-Aligned Binaural Spatial Audio Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:41:34.554124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.359873Z digest=sha256:81b9001831f6eb5466104544f75f9f811e6ac315d332670bb08b09299f7c5b6a

Observation 33e7086b-aa82-4375-83bf-94946de8c5eb · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.033078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.363774Z digest=sha256:9de014a1285dfaa23e8d60f0ffe69668d538ddd9f0443dd6b0e40a3af6aeafc8

Observation 8ceabb6b-f3d7-4e72-bc19-07aedac7bb3a · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 53

Resolution
verified exact
raw_fallback, observed 2026-08-11T15:41:34.532298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.367982Z digest=sha256:5c32c9b3d90c64a340fdd8a99ed74d9758c5907dd9089ffd45a9dfcb86b4e55e

Observation 8ac456d1-b78b-402b-b432-68c7559eafbe · outbound

This paper cites ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.372388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.372388Z digest=sha256:9ae86d1b834da4cebf38a39b2b24bc57d329e3397796fec6b2badfee9107d955

Observation d44e608b-7bd0-49d1-8b02-933e5ccad5ec · outbound

This paper cites silent frames.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework silent frames

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:35.015532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.377167Z digest=sha256:36b821765144b9793111a2a97e3869db38c8fb7a60eb518fa0eeff9235507586

Observation 9bd2df92-e91e-4ca8-aa53-5fa0a9f34f8e · outbound

This paper cites the sound of a vehicle driving.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework the sound of a vehicle driving

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:34.992918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.381968Z digest=sha256:39a5768de78637164aba5b6efe65a4a065f145b08aed18143af344e0153d6c4f

Observation f7eef15a-5a49-4d27-a3b0-31cd6b8af678 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:35.119557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.316499Z digest=sha256:a3c89a96ed59bbeed17a15471b2a2fa9963b8df1b0470dbcc1c4a8b74d4dcf0e

Observation dbdff0f5-4c34-4ab9-b1ad-5ccd20273336 · outbound

This paper cites In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.112150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.112150Z digest=sha256:ac56e95064847ff242041a4300d7ac6c7614e86cc3ab61939f89a459825b85cf

Pith citing papers

No inbound Pith citation observations are available.