Pith. sign in

Paper Citation Record · LEDGER

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

As of 12 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.09529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09529 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.381968Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5499e55c-4442-485c-a0ce-bd274fcb657b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.106949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.106949Z digest=sha256:39de47f2507885e0b55e94dfb72879eeb198e9a738aff2f6fc8a7eeb15809e3d

Observation 7e1016fa-d14f-4487-a833-e02b546b800a · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.117045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.117045Z digest=sha256:b36f226be3ecf133befe2d9e36c960b05a1e03e2159cfc0f1cd73e40d98aea5a

Observation 54d3592b-0ec5-49dc-8047-bca52ce43250 · outbound

This paper cites Qwen3-VL Technical Report.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.121620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.121620Z digest=sha256:1e245b5cf6a58c42360d550e94dc96a7471876f4465a4e6f9a1c2be44595d8c9

Observation 77c8da77-0574-4eaa-bfe3-c6507eb04917 · outbound

This paper cites SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.126404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.126404Z digest=sha256:e5434ab511b6274b778fbc1bfdcd86658fded6d77969265abfedd415604c8dad

Observation 1f347c07-c602-4f2c-9057-c1e130551bbe · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.403358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.131345Z digest=sha256:91caca1a326a3ac79045d191db5612901c029f0d4a112442fa5e296da351e8bf

Observation c7f64e7b-04f9-4c96-b155-f6ad8df9f481 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.391346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.136526Z digest=sha256:4195291200c83a5dd1e4d8ae2c6ce46c7f0c197d67e545657ff6f3a8f1820389

Observation 7d258208-eb36-4d96-9d87-f42ae61d879b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.141840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.141840Z digest=sha256:25c921a5ceb86a7886de88a0a94e79bf42c493e2ce14904a63e7a058596aa92d

Observation 9b954811-bcdf-4f68-98ca-badcb11de8cf · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.147229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.147229Z digest=sha256:30d7e02ffdb2a24b803ecbea7c848981683fc7dbd9cc23d9e16857edf57cf9e1

Observation 81f7bfc2-eef1-4023-80a4-976ee2762f86 · outbound

This paper cites FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.152062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.152062Z digest=sha256:9ecf64277cbc20bfc5eb9e2f06255522ce62a33840449c6802275855542da3f2

Observation 81131255-de76-44bc-adcd-a2201c8416e7 · outbound

This paper cites Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.156810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.156810Z digest=sha256:04fb30ea6303bd0f3fedbb802d1eb9e101409bfe73ce350ce02795efc5ec5d7b

Observation 407959c6-c6ae-48f5-8ce6-4c796e047570 · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen2-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.161186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.161186Z digest=sha256:b09a802a0ed1a31605453cbce7bb57e6139c19ddbd70b4cf5155dba09f5ebb9d

Observation 445908df-85c5-4b6e-8f1c-177512540198 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.165702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.165702Z digest=sha256:54e8a636a03c336305d6580ae4ed0c08646bd038b8f09c84cae9051703e751db

Observation 44a65de6-7316-4f2d-bf81-582b7ea5dc66 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.371299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.169825Z digest=sha256:5ab281416be0a1c9492915a6ed86e20bc1a49fc32c23143853e8278efb629437

Observation 170ce7ee-8829-44a9-b2ec-0311c63bab55 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.174885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.174885Z digest=sha256:9a463ca2384ff64ec183a169bf29b4e0824ab6969e116d79b3b1a750b27a3b9a

Observation 7ab50835-eb3a-444c-92f0-63f2e4bfdd6c · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.178845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.178845Z digest=sha256:f2c88e8b567fb1efc32d2104a55c845ffd50f6f1142c71bbc1e9e9b633619157

Observation fdfeec41-52cb-4c17-ab5e-f51637f8aae6 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.340192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.182233Z digest=sha256:568b1d00b16fd0f23da4c9beb969d931ae47347ed6aa40ebbef6ecbd82cceb91

Observation 085e0dd4-a64a-4cbc-9805-7e14955564b2 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.185572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.185572Z digest=sha256:9d649353242c43c3a31925b19b81e4b4734f45404211bf28cb9812283a5c13a4

Observation b158ec18-97bb-4b2a-bbd7-8c095c677568 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.191958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.191958Z digest=sha256:c3f9e81ad580276b5e564598773627d3ccc3450ceade386cb5355215d3cd879d

Observation 51445caf-5a85-41d0-afb7-0efef22bf28f · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.195746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.195746Z digest=sha256:ce305169ea18dacf0dbe538089f3ec6027462437aeef20d365642d0b2ed0c279

Observation 05a2c0ef-b839-4460-b0bb-414505746313 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.310410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.202252Z digest=sha256:76874e87959bd4e6da0f699cd48b553e8da6f4da515b7d0cbf4b5103f6f6ab6b

Observation c0e26a85-064d-4ffb-bf78-a034169f0d44 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.296132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.223442Z digest=sha256:91007d7eec0bac6ea0a649c276c70fd02994fb865ba802117e3f2f254ca17c6f

Observation 47d65d10-c76a-4f9e-8db1-8db6158f11fc · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.283319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.227907Z digest=sha256:72e6a6f53fd6618f7d817f52f9162877027621e01d20b3b5290d91594b785fde

Observation 0a176863-2ca2-4c67-8b38-1b3a937649ec · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.233176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.233176Z digest=sha256:626ae7f4cdcc1dbd9dc6fa79c271738ffe82ed83d390cd90dd41eba686d956ff

Observation ed524262-779f-452a-aa52-7855b4c1be89 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.261879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.237280Z digest=sha256:25fac18eee1f59676a40926dce4779c28d43d750fc4acf7bfb81e9eff979a035

Observation 48d333c0-3f15-4fb8-93cb-f3b4b4fe15c9 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-11T15:41:34.422290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.241666Z digest=sha256:c792fe03ed753d07dca7ab3fc423bc70512a2e76783c2fedafb60cd97028ca4f

Observation 529399a4-06e1-420c-977a-77791d751c37 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.248327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.246374Z digest=sha256:42e16d2114b25e14b7b47147491b5d53db084b7371445cf9279aee03f91ad047

Observation 5881b4de-0b83-4946-9684-cdc6ca2e0973 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.251401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.251401Z digest=sha256:61e7a283cad6b2d0cb716a8a2d9336421dd66f2614b55e9052fd0191500f0e59

Observation 0580e7e2-d925-4381-b93b-d9ef06abf2a6 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.256085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.256085Z digest=sha256:ea79dc83989d6cf49b401f71cf306e93278bc4f52d8562009d4ad45a0045635b

Observation 3114ed37-9df1-4fda-b656-cc91ef055d32 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.227361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.260375Z digest=sha256:fd5cc499cfd21cae629c0f2b5ee2d232ec549c190590b7ec1da2f38d56423445

Observation c02742d9-3e0c-45fc-b9a8-24c9ffe28e52 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.264355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.264355Z digest=sha256:217282ab51aa2f5918d4f15195fc95fb218919b73e23d0cf282541f78639df30

Observation 996eb128-6663-429c-a08d-09ed058974a5 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.203076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.268402Z digest=sha256:9e2cea1430d739e32006f96811f6a71dedd3ad02b56011d06a00aa589c718386

Observation 285cf977-bf1b-4acb-b96e-aabff348d07e · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.272510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.272510Z digest=sha256:ddddf7ac83f7ff2f3364977074fb85ce92cecde4796c3b8ded1d1453b2b4224e

Observation c544954d-47fa-487f-901c-17aba8c82d8d · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.276452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.276452Z digest=sha256:e11b07834d445648c369c834ce1ee3e677e3085ad48ca429904f1c2d88a2d1ee

Observation b5435a5d-54f9-42ea-a39f-2c32235422e8 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.280600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.280600Z digest=sha256:26e533f1e96f189689dd5a33e6151b6285774baa72f5d451d058a422238cd8e7

Observation 8b61f7b0-894b-45bb-817c-5084ad65406b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.174848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.284566Z digest=sha256:65bee513352100159c97094b8316e28dedc9b464fa22bf97eda673b207c6c4a0

Observation b94af2fb-3928-47b3-abdf-9f2d7a1a054e · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.288493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.288493Z digest=sha256:1b53226a9d6432f336e9b4e780063255666b4d884382db582e80346688ed9135

Observation 8eaed4ba-c82c-49fc-9570-f147d7200303 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.152737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.292361Z digest=sha256:e63bafac2e9632790e54c5fbf53ceffb06817d4da626e248496d60c284881a46

Observation be5bf654-8cbe-4ad3-9153-98eccb2f164f · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.297516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.297516Z digest=sha256:bcbf85581cb64c94204fc7b8152b54cb811760c2eb153b139e487a3342239ee5

Observation 49af4d98-5205-44f9-89f9-27b18638e561 · outbound

This paper cites InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.302159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.302159Z digest=sha256:1c0898d94150c7b11a11866398e73a8f3bb9ce3d83179a293f5e2d492e809418

Observation 72ffc7f8-406c-4ebc-8961-f11d36e33b6d · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.307150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.307150Z digest=sha256:a5e655344bf706e4b0b81ffa1363cdcf7eca6285f57d77409640341d879caa49

Observation b3ef91ca-ce95-4f76-b966-c5dd5866864d · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.137389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.311858Z digest=sha256:4d2762fdff308f47635f24fcc3c17a3ab672f36f6e3b34dc974c24799c8201f4

Observation ac2da736-ac44-45da-bf7c-b87ffbb70170 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.102644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.321229Z digest=sha256:4103d788ac85be41bd287ecad348536e8fc18c38815f9116827fdfc734d12ade

Observation 06941a77-2e5c-418b-8df4-ee7edef33f36 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.325382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.325382Z digest=sha256:b4ef29f8d41bbbb67426b67dc35d4d97b23830536b9b770ef4c017f1f9ef8792

Observation 445351ab-91db-4c6e-b5e2-ffdb4e3c8056 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.076228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.329856Z digest=sha256:343db5063e3540ca4fa3b72c97a5b609dee0fa70d76ffab95a38b45d66a7dbcf

Observation a761130c-9057-4301-afa8-f83b212722c5 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.334245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.334245Z digest=sha256:8dda3105dbcad43e576a4f55ee2d29de57bfe6f40a9658e289a0e9c8f8ecbd11

Observation 8fe144ba-39ec-4501-b027-5122f5c9f762 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.050662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.338582Z digest=sha256:af5683669731eefde11f5ee8ed7a96ac82b3988e14b0f7ebd51474959784bab6

Observation c0db9894-a99d-4c82-8c84-6694902529e0 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.343280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.343280Z digest=sha256:95e6280aced79b87714a864d7f8400b9ce96ea54a62c68943cb256cb2c017f9e

Observation aef01029-75b3-4962-9c29-7563eb6c0a20 · outbound

This paper cites AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.347508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.347508Z digest=sha256:fc230f4ca365d3bc2b014a61183cefb263ddad04085547b3c49b17013ffa956f

Observation e1a51c55-39ff-4bbf-bb94-bdbf3ca4cc5a · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.352108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.352108Z digest=sha256:0d4182d679b26f7c40beb33b0b437ae18ee40232e940aba3fd15dd62ec4d7688

Observation 1e3c2a70-445c-49fa-a886-e9efa78515e5 · outbound

This paper cites Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.355920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.355920Z digest=sha256:872822825c89b95668481bfa1f416083e6e243c00b4946b8b5cc665fccf0a8ce

Observation 6173af98-3b2e-4a96-9062-41097d0cc30c · outbound

This paper cites FoleySpace: Vision-Aligned Binaural Spatial Audio Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:41:34.554124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.359873Z digest=sha256:a848e81b5fbe56ebc8e84c2796674edc0022005a6833905cfdb6c2d8e4484bf0

Observation 33e7086b-aa82-4375-83bf-94946de8c5eb · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.033078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.363774Z digest=sha256:4d5ee1cb2b80d27212715a3bccb89f7838a499ad58d0f3f4236f87c1970c29d3

Observation 8ceabb6b-f3d7-4e72-bc19-07aedac7bb3a · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 53

Resolution
verified exact
raw_fallback, observed 2026-08-11T15:41:34.532298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.367982Z digest=sha256:66c050a9f7df0336b2b0cbd5387666a2a9f6c3452df1fe294f0f7145634a84f6

Observation 8ac456d1-b78b-402b-b432-68c7559eafbe · outbound

This paper cites ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.372388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.372388Z digest=sha256:3d20fb17c126b04bbc041f987ffdcb1fe18a8202e7372bc25f42371c13187a0e

Observation d44e608b-7bd0-49d1-8b02-933e5ccad5ec · outbound

This paper cites silent frames.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework silent frames

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:35.015532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.377167Z digest=sha256:af0ac9600750651a963f830d49148767c60968cbac40522775fe5cb888cf4e1f

Observation 9bd2df92-e91e-4ca8-aa53-5fa0a9f34f8e · outbound

This paper cites the sound of a vehicle driving.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework the sound of a vehicle driving

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:34.992918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.381968Z digest=sha256:3e5bff2d1b7cecb35482f6c9f906141a12774d11ad752a27788f2f134d1e9487

Observation f7eef15a-5a49-4d27-a3b0-31cd6b8af678 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:35.119557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:34.316499Z digest=sha256:2fcc0966d45b160d1039583df2ae6f3b5b40be78095873e1d668a9d25d5fd77b

Observation dbdff0f5-4c34-4ab9-b1ad-5ccd20273336 · outbound

This paper cites In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.112150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.112150Z digest=sha256:f3025ccb819e59005e11bf595ba708ca8849d4450cd22a2bafe4a06c24d27039

Pith citing papers

No inbound Pith citation observations are available.