Pith. sign in

Paper Citation Record · LEDGER

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

As of 22 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.20995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20995 v4

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:45:08.841789Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2efef86-b747-425b-86c4-8cc021f95666 · outbound

This paper cites The Foley Grail: The Art of Performing Sound for Film, Games, and Animation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance The Foley Grail: The Art of Performing Sound for Film, Games, and Animation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.351577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.880574Z digest=sha256:bed7cfc62d9922221e6904da9d1d39ca39e9eb9f03368ff6a76b2f252861732d

Observation bf501168-58c8-491e-88bf-c0087732d0c7 · outbound

This paper cites Qwen2.5-VL Technical Report.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.943204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.943204Z digest=sha256:c559ffdc9934b0738a11eb8eab0a2d007da1bf0ea1409d9397478648de7c5aba

Observation 138ac0b1-663d-4891-ae36-377b6e6f1b38 · outbound

This paper cites Erasedraw: Learning to insert objects by erasing them from images.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Erasedraw: Learning to insert objects by erasing them from images

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.340230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.024306Z digest=sha256:431494781c187c9ced1fe49daa03d7ebafe8f4153a63fff0008e33fe4e78b347

Observation 9cc5593b-09cc-4f8e-9f1e-3ef031053a27 · outbound

This paper cites Action2sound: Ambient-aware generation of action sounds from egocentric videos.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Action2sound: Ambient-aware generation of action sounds from egocentric videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.327905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.113307Z digest=sha256:d12fb8ceb7eb941b068d355692811c9a665c0a55fe12aa9d98eeb3fc4e6fb3c7

Observation f19cb447-815d-43e3-b6e9-057ba58a388e · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.316890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.234622Z digest=sha256:1fd04dd6def30b72b96c0ea87b776a9b73b9e55cd056e16edbaad45ba4a1afb6

Observation 50d1729f-30ce-4e2c-9a55-874ca54bf8db · outbound

This paper cites Generating visually aligned sound from videos.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Generating visually aligned sound from videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.305399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.353789Z digest=sha256:e042ca70c4844293763204290cb65823186dc084163d897d7fa8d564d37035d9

Observation 83671cc3-1f8c-46ab-adbb-f180111af7f3 · outbound

This paper cites Video-Guided Foley Sound Generation with Multimodal Controls.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Video-Guided Foley Sound Generation with Multimodal Controls

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:05.422297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:05.422297Z digest=sha256:142601434b720bcf0a032477efdc0c99a2318a1bce008efa36842cb6aa3363e3

Observation 465ce467-2a71-45bb-94bb-5bd783ab6dd0 · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.293192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.510328Z digest=sha256:30a8456bb8e0e39b7395a33cfedfbd5123ce6e7efc1bf77250ccaba5dcffea42

Observation f6f79a6e-9c2b-497d-9ccb-aa341c3f8f5e · outbound

This paper cites Clotho: An audio captioning dataset.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Clotho: An audio captioning dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.281452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.606265Z digest=sha256:f1c9938874503790d5adb543a4d299f851700efb16621d0a8697dbad6d8aa35c

Observation c48835b7-ca7e-4973-9b99-a865742c2709 · outbound

This paper cites Compositional visual generation with energy based models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Compositional visual generation with energy based models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.269734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.722732Z digest=sha256:e81ab8069e50dae62a738ccfe64d9ee8c991bcac605b7345fe7b897dbd22a8ca

Observation 8fcf95f3-9723-4d9e-8778-5fdb8415bf20 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Scaling rectified flow transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.258064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.808806Z digest=sha256:78487313e938ad2ef11f8cfb92c7bcefde65a81c79d115e688b0c2c1885a357b

Observation 5a745c0a-b954-408e-bd54-993eacde0289 · outbound

This paper cites Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.245835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.904797Z digest=sha256:fde440b44ddfbdbc467c58f7a269d6e5ea3b42a5d4b2e94e2f7e0aca62c67f29

Observation e2701f7c-992a-43d0-bba1-362698e35d76 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audio set: An ontology and human-labeled dataset for audio events

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.233666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.968972Z digest=sha256:434a59a61bfdf0ef632dc358299e16204e7efb37f8d90950f6bc5adf5d8246c4

Observation 8ab04e99-e18d-484c-9a57-60fc2b2b7602 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Imagebind: One embedding space to bind them all

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.222069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.021697Z digest=sha256:6a2854027359a7760509167f207711efdbc1fd73dea9bc4a2e49d5bff8397c1b

Observation d2359448-9249-44de-8680-c523e1dad7d8 · outbound

This paper cites Instructme: An instruction guided music edit and remix framework with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Instructme: An instruction guided music edit and remix framework with latent diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.209188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.075730Z digest=sha256:39ff88568de7e308abf0ea933c2f1cf0c8a520b66d4b7e500707c7ea671336bc

Observation 784a3cab-fd12-4243-b191-259de7ea9baf · outbound

This paper cites Classifier-free diffusion guidance.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Classifier-free diffusion guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.198403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.151487Z digest=sha256:120ea054991c95f56c157563e6d23abc33f1eb372dfadf5a150c7ae44745b149

Observation 791d89c6-44a6-45a4-9721-f24578894e16 · outbound

This paper cites Taming visually guided sound generation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Taming visually guided sound generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.186932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.203790Z digest=sha256:44fbe950ca5c0e550b5b68f5741e5c75aa348ce2bdefa771b974efabd762132d

Observation 3f00ebf4-1299-48dc-809f-7d36449e9505 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Synchformer: Efficient synchronization from sparse cues

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.174622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.275002Z digest=sha256:5a473edcb36b89d465602f6fd62feb28bf0c7ccb9c168fafbcb45829ab9564e7

Observation 76d83d63-5816-4a20-bcb7-456bbf0e3d3b · outbound

This paper cites Audioeditor: A training-free diffusion-based audio editing framework.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audioeditor: A training-free diffusion-based audio editing framework

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.163649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.323102Z digest=sha256:dd8cf719e6f4eac2a61b012d2a65b42423892dfcea7b5e64f70a67e9c8e8cc25

Observation a6d67126-3422-4ee6-b220-50bc93246462 · outbound

This paper cites Simultaneous music separation and generation using multi-track latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Simultaneous music separation and generation using multi-track latent diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.151760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.376824Z digest=sha256:870e0ed55b9dd466160c6df340effca7dc07e22bd087038f42d901162ef53fca

Observation 543f135c-bde9-4a08-907e-32953996f07c · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Analyzing and improving the training dynamics of diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.140505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.432024Z digest=sha256:cdc6982b131954590aa96b18798a08dc309556bce69765560d3f4197c0428ea4

Observation f0c3b862-20cb-440b-8185-4c5cc2213f44 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audiocaps: Generating captions for audios in the wild

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.128711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.488469Z digest=sha256:c05098fbbe1e1e60c02eb9e5bf2124739c27d4e5ba0fc8d0c0b292111ef05738

Observation ee2e0e48-327f-416e-93d1-9878a97a85f3 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Panns: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.116398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.539100Z digest=sha256:9297d6c5461a0c90a4435ec4609e96a57b8746979f43faeb726fcfe9885d584f

Observation 2285d8a5-0e8d-4396-94bf-2d20c51f43b8 · outbound

This paper cites Vintage: Joint video and text conditioning for holistic audio generation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Vintage: Joint video and text conditioning for holistic audio generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.104521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.608604Z digest=sha256:3372963793bdf5a6b58708164be5c74a378e98136ae99e432cc74f1450df5ee8

Observation 3b64cfc8-9b49-4187-b718-f8eddc76d437 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Gonzalez, Hao Zhang, and Ion Stoica

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.091741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.675323Z digest=sha256:4b84d5a0a8f3074dafa7aacbc53a5752644862f0d00185c962a494ed229705a1

Observation 8d5aed84-561b-4ff3-b3da-80fc48acb6f0 · outbound

This paper cites Flow matching for generative modeling.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Flow matching for generative modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.080943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.722269Z digest=sha256:632f7945b57a44fa86f974d23dc9856dd940ec8fd064e40ab78b6ca32a131743

Observation 72f18168-e4f9-43bd-bd23-35e8754133ee · outbound

This paper cites Compositional visual generation with composable diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Compositional visual generation with composable diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.069743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.789909Z digest=sha256:1677a881f887d2affc119cc29282a687894d692b831664d39a3b7e69a3177a42

Observation fbac445a-9263-4075-8796-713d4175d007 · outbound

This paper cites Tell what you hear from what you see-video to audio generation through text.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Tell what you hear from what you see-video to audio generation through text

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.058915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.849457Z digest=sha256:05366dc5f83ee56b3aaba2c9934dcebc078ddf878e58cfd10808740e95c1db77

Observation 5c27eb4f-97b7-4233-b673-0f8f00d8fd96 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.047856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.916636Z digest=sha256:814a0a53907b5956cf635d4da851133f6b8a231014a1d6a9255f41c106beb176

Observation 79b5cffe-d872-4a6b-a594-afe71466b990 · outbound

This paper cites Multi-source diffusion models for simultaneous music generation and separation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Multi-source diffusion models for simultaneous music generation and separation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.037038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.972516Z digest=sha256:18d35f27013e56caab52064e64ccfb7a5725828fc2e4cc04793a8bbda9afa154

Observation 1c51d7c4-9601-4119-a8e5-7b6e238e7efa · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.024863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.046130Z digest=sha256:f647fb10be1e5c733e7c203d4132dec570616055a27a5061008bc9837203f435

Observation 20fca138-03ed-415c-958c-0173fb57af3a · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audio-visual scene analysis with self-supervised multisensory features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.012958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.096950Z digest=sha256:6db9cf95e6672d121eca429bb8fcbf59da372364c95e39cf06963170a4035b35

Observation 0f77170d-dce8-464e-84cd-b1b0bb67c57f · outbound

This paper cites Stemgen: A music generation model that listens.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Stemgen: A music generation model that listens

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.001536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.138164Z digest=sha256:9930121cfce0eaea78de0c759363e73c08d133c38fd52c2f7890a41495deced9

Observation 186907ed-1502-45d5-82a0-79d9c0153b15 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Movie Gen: A Cast of Media Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:07.206934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:07.206934Z digest=sha256:de419e755b53f28be4dd0efb959298b13fc3a7a0954e69d2ac2901d1195d225c

Observation 463dd510-0502-4450-9692-436cc26a425b · outbound

This paper cites Generalized multi-source inference for text conditioned music diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Generalized multi-source inference for text conditioned music diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.990280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.274186Z digest=sha256:d9ead1d738a75a041bfbfd5d0bc08bdee23761823e54666477d5ea86250a2f9e

Observation b96150f7-b8f9-4748-a4de-37565fc32655 · outbound

This paper cites Improved techniques for training gans.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Improved techniques for training gans

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.978310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.347645Z digest=sha256:6d5a36a7a2a437f7d29f235fffc700cde8e28103156acdb775de1170a9d50647

Observation e65fd7f1-2874-4163-a217-7e3bb9e57f15 · outbound

This paper cites Smartmask: context aware high-fidelity mask generation for fine-grained object insertion and layout control.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Smartmask: context aware high-fidelity mask generation for fine-grained object insertion and layout control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.960730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.473056Z digest=sha256:9dc5b059c796b4e4f637f57bcc07a1d4a4e13338e5779c5eeb79ae35ce6f18c7

Observation a4f55f75-e5ed-4ed4-bb7c-d95b5b065c55 · outbound

This paper cites Visually guided sound source separation with audio-visual predictive coding.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Visually guided sound source separation with audio-visual predictive coding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.832404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.585707Z digest=sha256:47111ee5faed4af16a13d576cc6ead4d013969f36198e60fc2758f318a16f411

Observation 4e31c034-6b9a-4f01-93ab-ae7ff2779e8c · outbound

This paper cites sd3.5, 2024.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance sd3.5, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.502321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.682202Z digest=sha256:1260f8dc7396932deed44efc9445014637fe5f36aee56cd7b1c4978a8e3d3293

Observation 37145041-e787-421d-b6f3-a1e17e010e53 · outbound

This paper cites Steinmetz and Joshua D.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Steinmetz and Joshua D

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.318638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.779691Z digest=sha256:a0d818f686bec8ea6ce406596481412b277f1e6bbe3659d3c77cd500051a7397

Observation 03c38e60-10be-4a19-ae61-8e74395b6276 · outbound

This paper cites Add-it: Training-free object insertion in images with pretrained diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Add-it: Training-free object insertion in images with pretrained diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.045041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.883778Z digest=sha256:1b999e1626a8b221f0c05be96032ac23c796cf6d82d64d45e8d80385d4be8915

Observation 0bc16ca5-f94f-4c75-ba2a-1b4004b179b2 · outbound

This paper cites Liu, Kevin J.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Liu, Kevin J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.900034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.011468Z digest=sha256:34d5f391a59f7895c8c63a0cd4eb0332d6bc7a4e6654999ef88605b880132d44

Observation 9777209a-2c83-403a-82c9-35f0b3cdcafe · outbound

This paper cites Temporally aligned audio for video with autoregression.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Temporally aligned audio for video with autoregression

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.758908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.098102Z digest=sha256:5b01ee77b02a3ea85b67a386e145d60ba281ab5cd12e95c6a38805a47e4273cc

Observation f6e8b40c-660b-4d33-9552-06e095b8af5c · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.554435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.178693Z digest=sha256:014303027f49a12da3d69c56cda5f156007bd3ef19fb895a1ce454bf2254bd1b

Observation 5ba41fe0-c0b5-41dc-88ec-625b962f1694 · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Frieren: Efficient video-to-audio generation network with rectified flow matching

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.395064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.285868Z digest=sha256:cdfce237fc8f27bd5c68d35324f25c714fe3ff1d42097938f906c95260a7c539

Observation e0dde940-a687-4083-9f75-0920aa92541e · outbound

This paper cites Audit: Audio editing by following instructions with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audit: Audio editing by following instructions with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.226759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.365573Z digest=sha256:809b3da6e5891d9cc9835f9de1653e0e76c941c473ef0e9157d67242101bbb57

Observation d4d331b0-7589-46d7-a134-6f921ba9eae9 · outbound

This paper cites Stable diffusion 2.0 and the importance of negative prompts for good results, 2022.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Stable diffusion 2.0 and the importance of negative prompts for good results, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.978231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.474046Z digest=sha256:4a9f61deee6aaa404951079122604391e5cb0598b21033cd0d2e6ce467964b5d

Observation fbf2a7b1-9d47-4e00-a5ab-eed923cef885 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.732295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.535917Z digest=sha256:b0ad365f4dcbe7a197098c715687c64fe707a9ebbb76e08f66b83bd46a7649b9

Observation 0102a54a-4938-4296-8fd9-88cd98f822af · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.528185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.546056Z digest=sha256:1995fb5845a795a3d633172d7233fc3014e42883f724ac5cc0f8988df4933e83

Observation 6d026ab3-f7b1-4d29-9af6-1d3023bc86fd · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.304523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.621106Z digest=sha256:927f53b59b7546eccc7da33d944a137d5d1dab3bf3ae7f9c5f0af4d3ca0ccfc9

Observation f86c4cf0-2d82-4547-8a55-783902287549 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:08.730625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:08.730625Z digest=sha256:a9bad53fa5a005be643aa6fbc57e6ff37e56a27954e15955e8f6dd764b2b1ee5

Observation b754503b-876d-4787-b826-07e6136fc464 · outbound

This paper cites Visually guided sound source separation using cascaded opponent filter network.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Visually guided sound source separation using cascaded opponent filter network

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.086947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.841789Z digest=sha256:0998e36474f0c85016880154fa173ceb2d72e5024b6f0765229b385dc28e3020

Pith citing papers

No inbound Pith citation observations are available.