Pith. sign in

Paper Citation Record · LEDGER

Audio-Guided Visual Editing with Complex Multi-Modal Prompts

As of 10 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2508.20379.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20379 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:31.849774Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3920ab5b-1b92-456c-babc-9857e328da00 · outbound

This paper cites GPT-4 Technical Report.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.658133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.658133Z digest=sha256:a182b0f062cfe7d34c55b93af701012d59e603d295a29df74285270c8ed0b7c8

Observation adc71368-bcf9-480b-aff4-9589c8e14b6e · outbound

This paper cites Text2live: Text-driven layered image and video editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2live: Text-driven layered image and video editing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.469198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.662463Z digest=sha256:cfd05db94ec630d88599cded5c10541cf9e41bac573c579d4be132c0fbbe9196

Observation c1b863f0-6252-49cf-8419-27a579b8d80f · outbound

This paper cites SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:32.202014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.667945Z digest=sha256:52d5d04b35bda4a934ce9169707069e962802b96e0d230282ae92deaaefcb154

Observation 47e2c4cb-2364-43fa-8d26-fad3ea7f3196 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align your latents: High-resolution video synthesis with latent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.671833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.671833Z digest=sha256:87407f32439df685fb041fa5bc345414fce27efaa4a31ed79bcf1a9b39422071

Observation cb1a46e1-9ea0-42b4-9a49-88786c7f2087 · outbound

This paper cites Ledits++: Limitless image editing using text-to-image models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ledits++: Limitless image editing using text-to-image models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.454984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.674911Z digest=sha256:723e49bdcf60f89b99cd9a370add2f90020107f6b26cc634b0843997149e0ce3

Observation 8038bc95-56ab-4c82-a49a-426e5953eb8f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Vggsound: A large-scale audio-visual dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.444636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.678420Z digest=sha256:901b0a51ead6422bcc01fbbdd5a251485092f0ba6bcc014f708a2d7bc3a62c95

Observation 15255ef9-bc95-4305-ae2d-92fa964d1eab · outbound

This paper cites DiffEdit: Diffusion-based semantic image editing with mask guidance.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffEdit: Diffusion-based semantic image editing with mask guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.682377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.682377Z digest=sha256:60316fe2a06e4975e5f06a2d259c64a0e359bb5ecf65e3cb0dbdfca64809c2d0

Observation 3a9ba256-3b8c-407b-9a4f-47b4036a63ea · outbound

This paper cites Con- ditional generation of audio from video via foley analogies.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Con- ditional generation of audio from video via foley analogies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.434962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.685950Z digest=sha256:da8c67a2008960b27ff200464395f9600d4ba3f0ee6dbe2bec9d0ed743218bb1

Observation 70e339d7-4417-4a3b-8b9f-20a65d53e31f · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Structure and content-guided video synthesis with diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.425812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.689192Z digest=sha256:647ebdbc46c38791ee80d170dd2faa3cc2a5cb57870c283a870e0740461a5bc1

Observation c1a0ee9a-b796-402c-a5d3-19025e5b4ed8 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.692936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.692936Z digest=sha256:66b49f60e6059a71922f5c981f6fd3607c85ffe34c82ae40dd3ca0e3e8acb439

Observation 017fc962-2b1e-4a23-b0a5-9fed696c05ea · outbound

This paper cites Imagebind: One embedding space to bind them all.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Imagebind: One embedding space to bind them all

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.416553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.696749Z digest=sha256:457d11b8129f04ae0946cc73a5f2eca64384ab8aeb375398f45a26f42989f777

Observation a38220d8-905a-45b9-a58f-de8655d3d5f0 · outbound

This paper cites AtomoVideo: High Fidelity Image-to-Video Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AtomoVideo: High Fidelity Image-to-Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.700664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.700664Z digest=sha256:b69a96a8988909b7a6293af776ce56fa8a96a40786985b2ed707a27def567bf5

Observation b3099f05-64d3-4e41-bf32-662f496e67f9 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.704083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.704083Z digest=sha256:79c2f44f66f00500a07cc00c4a2ae7c01543a6192d33c192e7b7d7c297c2ef60

Observation c1de15e2-4ec3-43b2-9a43-f8d8bd6fb07e · outbound

This paper cites FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:32.154458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.708990Z digest=sha256:1f4beec4cb031b3d123b4592299e3178d3833801b358b547d42bc587872b052a

Observation 6c938676-cf99-4b22-b3d5-29caf78618bd · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.713176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.713176Z digest=sha256:e705604eb2d76a174fb2cb35a0947c6c94cf54cfd2db0d66b6ccb8e44d3ae864

Observation d9c1f186-073e-48e5-a586-4874717b2784 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.716262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.716262Z digest=sha256:3ebf022f07220beb1d686f73d8a34f61ff67bba776d4ea94c357e20b311456b3

Observation c4435c43-05ee-4b21-9678-726f78a5d31f · outbound

This paper cites Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.719815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.719815Z digest=sha256:78fb98a03a9f40c28483c47c27c25afdb7c644b3c97f461cd82798d018fab6d3

Observation 38c81fa2-f69a-4611-8626-a787974b7d05 · outbound

This paper cites Text2video-zero: Text-to- image diffusion models are zero-shot video generators.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2video-zero: Text-to- image diffusion models are zero-shot video generators

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.407619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.722751Z digest=sha256:2dfcfd2460f75c8f2b4ccdadd11bdba1f32155f67c3494bbda5a5cb3c952a4c6

Observation d69db800-7323-4db2-a592-e0ed9afcac39 · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.725752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.725752Z digest=sha256:38cdc180f53d781a37f12ebea502b75287daa0b9ffbf2b81cc92baae717d9c70

Observation 6ebd06e6-9de3-4d2c-8045-5942e4032e18 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Multi-concept customization of text-to-image diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.728658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.728658Z digest=sha256:ca1de99331199e70e5c966d3ff9b5cb824dd3846b385606bd44cff156410aab6

Observation e6b2b2f6-8177-49e4-9f46-f74b2ee39ef4 · outbound

This paper cites Sound-guided semantic image manipulation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Sound-guided semantic image manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.392936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.731687Z digest=sha256:96179cd9037331cca351be647be7c3a92f2f3db843a59648aec3373dd9a4d32e

Observation 44e2fce0-b5e8-460a-89a0-53bb27120f97 · outbound

This paper cites Soundini: Sound-Guided Diffusion for Natural Video Editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Soundini: Sound-Guided Diffusion for Natural Video Editing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.734970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.734970Z digest=sha256:253380bbcfdec0750e7a60b40f427df144be86611cc6bb021a4a1c967f0c243f

Observation 2b58a244-9615-45e9-87b1-0f98f98c73d5 · outbound

This paper cites Generating real- istic images from in-the-wild sounds.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Generating real- istic images from in-the-wild sounds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.383588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.737985Z digest=sha256:f8b66ef995a40a86fa69e66616887639e30ec1040e6992e93157e8fcf8ac2b9c

Observation dbb0e571-384b-4b43-9fc2-58e48bf14f32 · outbound

This paper cites Learning visual styles from audio-visual associations.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Learning visual styles from audio-visual associations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.374499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.741316Z digest=sha256:0c5bf7198cb19d91dd25fa141ae35658110616079e6ac39454e9cf23c33dceb8

Observation d0ef518f-84ba-4c12-bfae-ff480a868876 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.365099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.744280Z digest=sha256:e0c9327837ad240ae09199f721f486892ba7f243cd4780a4f6608f175c680e32

Observation ff760cde-a948-48cb-b605-60bd5298135c · outbound

This paper cites Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.747156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.747156Z digest=sha256:7bfc9224ef1167b1a2314847adc431f10156b249f5e825d4a1cf7efc6da3365b

Observation d6bfd0e8-3685-41e7-ab70-c02f60c6012f · outbound

This paper cites Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.750629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.750629Z digest=sha256:28d4d2d932384b7b530458824e69a26abd7a703b92b911e6dcd5fb48dfffe828

Observation b0ef9092-7b6a-45a5-a957-900a8070585c · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.753877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.753877Z digest=sha256:2b273ca11d0813ce217817fd58101b2a24f17fd2f9fd179bdc2ff58808cbf848

Observation 2af4216c-7cf4-4b16-8b8f-ab8664569220 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.355091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.757006Z digest=sha256:2ebba1acde9dc6392e3b4ee7e2d62f23a2e1e63a4a4df248172a3d61c006cc7f

Observation 79bccee3-4715-4b5f-baf5-07bff6f8656c · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Videofusion: Decomposed diffusion models for high-quality video generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.345669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.760023Z digest=sha256:0fc46437e65dbf5226ac333c054922661a66b3743ca6c3e484dea53c60567510

Observation 75333614-6bef-46d9-9f16-2e05dbf39fcb · outbound

This paper cites Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.763240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.763240Z digest=sha256:dfb34b5b60543584590e99420a2b5917224a3e3237a4f8e2da5ecbc366385831

Observation 81065768-7ebe-4cfd-8d4a-63f846f602dc · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.766335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.766335Z digest=sha256:c5cc62be0bab35cfc112d4e1ed1fbc51fb22aaf5c6e4f15bbfa6632733a51959

Observation b15750d2-18e5-4fd5-9900-ea7ac919ef72 · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.336618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.769728Z digest=sha256:6847abbb5e040d12b6c5f4aed85720456c0e77fc3de4b82a8808198d2eb0311e

Observation f717f9cb-22d1-47f5-97fb-41227fd3e1b8 · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Null-text inversion for editing real images using guided diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.327476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.773612Z digest=sha256:c70ff2f98f582f71ed94a1493df876c5e8fece47d37e9849b5811ec655a89a10

Observation 705322b1-435b-438a-83ee-9b6922827448 · outbound

This paper cites Conditional image-to-video generation with latent flow diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Conditional image-to-video generation with latent flow diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.318131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.776779Z digest=sha256:de0baa7657ed18c6ddb1c501d66140471a6fdd59e175dbef5f2407ab98fdbc67

Observation 0d3e2c24-9ba4-496b-8b26-cda1590b0b48 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.780021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.780021Z digest=sha256:bc288b33b1af4b287ba1f2b54d5854aeb26f8d91be526d0da40b0ea13f7a15db

Observation c37a2fe5-dc43-484b-a07a-55102189a072 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts The 2017 DAVIS Challenge on Video Object Segmentation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.784026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.784026Z digest=sha256:b3a3bd5a11402fc037ed1b72afef1d412b32720ee9cf4180293592605a651f1f

Observation ed40fd61-e4ed-44a3-9e05-3eaaa2dddf45 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Grad-tts: A diffusion probabilistic model for text-to-speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.788198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.788198Z digest=sha256:f71217a5b99451d3fa1f4deed02d00c8938c7d486f5888f7d72ff3528b9c2bc8

Observation 8d8de396-afe3-4b6b-93c1-4d7b4d040fc0 · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Fatezero: Fusing attentions for zero-shot text-based video editing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.301659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.791205Z digest=sha256:5e931e86ee1c8e031183f7eb156ead6123f9d0125832e412fc78d2504d85ab6d

Observation a225ef8d-4329-4cee-8f46-ac398c6ce10a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.794473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.794473Z digest=sha256:795ccefd674751b22607360b1803628d3c458e5d610826a030f12ee8bcf54ae9

Observation d343c0ec-5f2a-4eb8-9491-4e2c3ba36eb8 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts High-resolution image synthesis with latent diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.292428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.797335Z digest=sha256:60d35e1e2c35e2643364a088b028a37325c0e55189b25833404df21ad8a830c2

Observation cc1637b2-aefb-40aa-a17b-b07656bcda18 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Photorealistic text-to-image diffusion models with deep language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.283106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.801313Z digest=sha256:2c81d3274e4574123a7a5d8a88ed00a1c4db5a9325f679149b032f0bc17c0814

Observation 6562b558-0a3e-403b-9f40-55481a3dceab · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.804516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.804516Z digest=sha256:b948d08bf9ebc683099403f66f680902542dbae5fe6576b5a0c732e0e2eda20f

Observation 0277b3f2-5a4f-4538-8813-931e2d73d77c · outbound

This paper cites Denoising Diffusion Implicit Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Denoising Diffusion Implicit Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.807657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.807657Z digest=sha256:7f1bc0faa1ddb312dbebabcd612d5f689c5d77f0fdc969cbb30d56878ffcd0c7

Observation 68203758-5db6-4225-9d4d-4f63e58a8ec0 · outbound

This paper cites CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.811712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.811712Z digest=sha256:1643f331ffa5bc42b041ee0616c293d42747a96aee2ccc29cbf692f1b26bbafa

Observation 8002fae2-0e9c-43c9-b91c-38834b9edba0 · outbound

This paper cites Any-to-Any Generation via Composable Diffusion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Any-to-Any Generation via Composable Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.814837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.814837Z digest=sha256:95869f47a5a97bf7f8b3638fbf8d7e8052dde2c9f683fb16a43a6b7123bd7e26

Observation 8390fd36-fdb4-4726-9136-619a2a8f4708 · outbound

This paper cites Splicing vit features for semantic appearance transfer.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Splicing vit features for semantic appearance transfer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.273359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.818212Z digest=sha256:da4ac0bb892dfe0a7d140725fd051bf2f3b3074f0f8f35d9ce1c3cfed0b0df6c

Observation 99c62098-9257-42b2-8a37-27ae0cebcfaa · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Plug-and-play diffusion features for text-driven image-to-image translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.821082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.821082Z digest=sha256:517ef7f4aeaaa0c9164f65367d9a3411f74fb554b3a66b13fef231d7ce2946a7

Observation fa7fff4f-95cf-4e2e-8053-f3d43caeab9a · outbound

This paper cites Audit: Audio editing by following instructions with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Audit: Audio editing by following instructions with latent diffusion models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.258243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.824068Z digest=sha256:79a1c7c6193000c79ec31c3ffc0a08ca9305af0d8b22cfefd70cd7937be202af

Observation 509d7ec8-fc05-4c4a-a237-72817d40639c · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.248945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.827209Z digest=sha256:5a3e1034b4109240852a61c43eaadb9686a844c2d2baea18fe59f7d132a2754d

Observation 12fef8f2-af64-4a1a-b178-bf8f1f08b6cc · outbound

This paper cites CVPR 2023 Text Guided Video Editing Competition.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CVPR 2023 Text Guided Video Editing Competition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.830361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.830361Z digest=sha256:3ccbcdb17ffde90dcfebf7c67f19559217048905a4c6d4e109e80e3bcd057db7

Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.833708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.833708Z digest=sha256:98fe7dd9c775e1a86edca384aec48ec3d80c7b53283a67aae064500fa94ee577

Observation b239d8f0-2453-48f3-a5b9-0c93d6930919 · outbound

This paper cites Ar-diffusion: Auto-regressive diffusion model for text generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ar-diffusion: Auto-regressive diffusion model for text generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.239486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.837124Z digest=sha256:980415d32d57624a53df4adc745244b8c3bb61ee4c65f8a9b07e0cff1710cc31

Observation 6617b490-5b00-4988-a50e-b1cc63fb1f43 · outbound

This paper cites Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:31.902187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.840253Z digest=sha256:64c482c0d08564a8a8dce84889122c3a3376f8999b04133e7bf55451e7e57990

Observation 3de73937-bc91-4434-b7a8-97c435a9ecdc · outbound

This paper cites Align, Adapt and Inject: Sound-guided Unified Image Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align, Adapt and Inject: Sound-guided Unified Image Generation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:31.886034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.843363Z digest=sha256:f194912e8b34bae5858ae6c4f0f16de73a3e9de334fa773f0f920cde7743c1be

Observation 91f0718d-73b0-4c14-a3c9-8a09db9e19bf · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts The unreasonable effectiveness of deep features as a perceptual metric

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.230376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.846869Z digest=sha256:d08abdb7468ef8a971ce16c6c6146f6640c35d668bf7cc4b131eedb0fef70a2d

Observation 40491051-3946-400f-9b6c-7283985938a4 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Uni-controlnet: All-in-one control to text-to-image diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.221511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:10:31.849774Z digest=sha256:fb71425da1704a11fd054648713524c0edba3782959bba2361cf2e429e5d7023

Pith citing papers

No inbound Pith citation observations are available.