Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:31.849774Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2508.20379.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:31.849774Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3920ab5b-1b92-456c-babc-9857e328da00 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc71368-bcf9-480b-aff4-9589c8e14b6e · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2live: Text-driven layered image and video editing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1b863f0-6252-49cf-8419-27a579b8d80f · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47e2c4cb-2364-43fa-8d26-fad3ea7f3196 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align your latents: High-resolution video synthesis with latent diffusion models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1a46e1-9ea0-42b4-9a49-88786c7f2087 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ledits++: Limitless image editing using text-to-image models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8038bc95-56ab-4c82-a49a-426e5953eb8f · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Vggsound: A large-scale audio-visual dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 15255ef9-bc95-4305-ae2d-92fa964d1eab · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffEdit: Diffusion-based semantic image editing with mask guidance
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9ba256-3b8c-407b-9a4f-47b4036a63ea · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Con- ditional generation of audio from video via foley analogies
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 70e339d7-4417-4a3b-8b9f-20a65d53e31f · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Structure and content-guided video synthesis with diffusion models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1a0ee9a-b796-402c-a5d3-19025e5b4ed8 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts TokenFlow: Consistent Diffusion Features for Consistent Video Editing
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 017fc962-2b1e-4a23-b0a5-9fed696c05ea · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Imagebind: One embedding space to bind them all
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a38220d8-905a-45b9-a58f-de8655d3d5f0 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts AtomoVideo: High Fidelity Image-to-Video Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3099f05-64d3-4e41-bf32-662f496e67f9 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1de15e2-4ec3-43b2-9a43-f8d8bd6fb07e · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6c938676-cf99-4b22-b3d5-29caf78618bd · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c1f186-073e-48e5-a586-4874717b2784 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4435c43-05ee-4b21-9678-726f78a5d31f · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c81fa2-f69a-4611-8626-a787974b7d05 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2video-zero: Text-to- image diffusion models are zero-shot video generators
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d69db800-7323-4db2-a592-e0ed9afcac39 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ebd06e6-9de3-4d2c-8045-5942e4032e18 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Multi-concept customization of text-to-image diffusion
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b2b2f6-8177-49e4-9f46-f74b2ee39ef4 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Sound-guided semantic image manipulation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 44e2fce0-b5e8-460a-89a0-53bb27120f97 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Soundini: Sound-Guided Diffusion for Natural Video Editing
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b58a244-9615-45e9-87b1-0f98f98c73d5 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Generating real- istic images from in-the-wild sounds
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dbb0e571-384b-4b43-9fc2-58e48bf14f32 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Learning visual styles from audio-visual associations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d0ef518f-84ba-4c12-bfae-ff480a868876 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff760cde-a948-48cb-b605-60bd5298135c · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6bfd0e8-3685-41e7-ab70-c02f60c6012f · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0ef9092-7b6a-45a5-a957-900a8070585c · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2af4216c-7cf4-4b16-8b8f-ab8664569220 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 79bccee3-4715-4b5f-baf5-07bff6f8656c · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Videofusion: Decomposed diffusion models for high-quality video generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75333614-6bef-46d9-9f16-2e05dbf39fcb · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81065768-7ebe-4cfd-8d4a-63f846f602dc · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15750d2-18e5-4fd5-9900-ea7ac919ef72 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f717f9cb-22d1-47f5-97fb-41227fd3e1b8 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Null-text inversion for editing real images using guided diffusion models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 705322b1-435b-438a-83ee-9b6922827448 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Conditional image-to-video generation with latent flow diffusion models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d3e2c24-9ba4-496b-8b26-cda1590b0b48 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c37a2fe5-dc43-484b-a07a-55102189a072 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts The 2017 DAVIS Challenge on Video Object Segmentation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed40fd61-e4ed-44a3-9e05-3eaaa2dddf45 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Grad-tts: A diffusion probabilistic model for text-to-speech
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8de396-afe3-4b6b-93c1-4d7b4d040fc0 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Fatezero: Fusing attentions for zero-shot text-based video editing
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a225ef8d-4329-4cee-8f46-ac398c6ce10a · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d343c0ec-5f2a-4eb8-9491-4e2c3ba36eb8 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts High-resolution image synthesis with latent diffusion models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc1637b2-aefb-40aa-a17b-b07656bcda18 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Photorealistic text-to-image diffusion models with deep language understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6562b558-0a3e-403b-9f40-55481a3dceab · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0277b3f2-5a4f-4538-8813-931e2d73d77c · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Denoising Diffusion Implicit Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68203758-5db6-4225-9d4d-4f63e58a8ec0 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8002fae2-0e9c-43c9-b91c-38834b9edba0 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Any-to-Any Generation via Composable Diffusion
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8390fd36-fdb4-4726-9136-619a2a8f4708 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Splicing vit features for semantic appearance transfer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99c62098-9257-42b2-8a37-27ae0cebcfaa · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Plug-and-play diffusion features for text-driven image-to-image translation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7fff4f-95cf-4e2e-8053-f3d43caeab9a · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Audit: Audio editing by following instructions with latent diffusion models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 509d7ec8-fc05-4c4a-a237-72817d40639c · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 12fef8f2-af64-4a1a-b178-bf8f1f08b6cc · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts CVPR 2023 Text Guided Video Editing Competition
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b239d8f0-2453-48f3-a5b9-0c93d6930919 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ar-diffusion: Auto-regressive diffusion model for text generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6617b490-5b00-4988-a50e-b1cc63fb1f43 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3de73937-bc91-4434-b7a8-97c435a9ecdc · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align, Adapt and Inject: Sound-guided Unified Image Generation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91f0718d-73b0-4c14-a3c9-8a09db9e19bf · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts The unreasonable effectiveness of deep features as a perceptual metric
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40491051-3946-400f-9b6c-7283985938a4 · outbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts Uni-controlnet: All-in-one control to text-to-image diffusion models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.